Information processing device and information processing method
The dialogue system addresses the challenge of managing user-passenger communication by classifying user input words into concept levels and adjusting language based on passenger presence, ensuring appropriate and privacy-protected communication.
Patent Information
- Application Number
- JP2021120381
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-21
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2041-07-21
AI Technical Summary
Conventional dialogue systems fail to appropriately manage communication between a user and a passenger, often prohibiting necessary information exchange due to categorization of topics, which can lead to privacy concerns and inefficient communication.
The system extracts and classifies words from user input into multiple concept levels, adjusting the language information based on the presence of a passenger within a given distance, using higher concept levels when a passenger is present to ensure privacy protection.
This approach enables appropriate communication while protecting user privacy, by suppressing unnecessary information and enhancing user satisfaction and affinity with the agent.
Smart Images

Figure 0007676253000001 
Figure 0007676253000002 
Figure 0007676253000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device and an information processing method. [Background technology]
[0002] A dialogue system is known in which topic information is classified and stored into a first category that may be used when a passenger is present, and a second category that is prohibited from being used when a passenger is present, and the content of an utterance is determined from topic information belonging to the first or second category when there is no passenger, and the content of an utterance is determined from topic information belonging to the first category when there is a passenger (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2019-101805 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the above conventional dialogue system, even if there is information to be conveyed to a user such as a driver, if there is a passenger and the topic is in the second category, the information is prohibited from being spoken, and even if the information should not be conveyed depending on the relationship between the user and the passenger, it is spoken if the topic is in the first category. Therefore, it is desired to have a so-called agent function that can carry out appropriate communication while protecting the user's privacy.
[0005] An object of the present invention is to provide an information processing device and an information processing method that enable appropriate communication while protecting the privacy of a user. [Means for solving the problem]
[0006] The present invention solves the above problem by extracting words contained in input data acquired from a user, classifying and storing the words into a plurality of concept levels including at least a superordinate concept level and a subordinate concept level that is relatively subordinate to the superordinate concept level, generating output data by including words classified into any of the concept levels in linguistic information, and outputting the output data by an agent function.In this case, it is determined whether or not a person other than the user exists within a predetermined distance, and when it is determined that a person exists within the predetermined distance, the output data is generated by including words classified into a relatively superior concept level in the linguistic information, compared to when it is determined that no other person exists within the predetermined distance. Effect of the Invention
[0007] According to the present invention, it is possible to carry out appropriate communication while protecting the user's privacy. In particular, when providing information using the agent function, the provision of unnecessary information is suppressed, so that it is expected that the user will be more likely to have a favorable impression and a sense of familiarity with the agent. [Brief description of the drawings]
[0008] [Figure 1] 1 is a block diagram showing an embodiment of a vehicle information processing device according to a first embodiment of the present invention. [Diagram 2] 2A and 2B are diagrams showing an example of the interior of a vehicle where the agent of FIG. 1 is installed. [Diagram 3] 2 is a diagram for explaining information processing of a word executed by the vehicle information processing device of FIG. 1; [Figure 4] 4 is a diagram showing an example of the configuration of a database in which the words that have been subjected to the information processing shown in FIG. 3 are stored. [Diagram 5] 4 is a flowchart showing an example of an information processing procedure of a word executed by the vehicle information processing device of FIG. 1; [Figure 6] 4 is a flowchart showing an example of an information processing procedure for output data executed by the vehicle information processing device of FIG. 1; [Figure 7] 7(a) is a diagram showing an example of a scene when step S23 in FIG. 6 is reached, and (b) is a diagram showing an example of a scene when step S26 in FIG. 6 is reached. [Figure 8] FIG. 11 is a block diagram showing an embodiment of a vehicle information processing device according to a second embodiment of the present invention. [Figure 9] 9(a) to 9(c) are diagrams for explaining information processing of the tolerance level executed by the vehicle information processing device of FIG. [Figure 10] 10 is a diagram showing an example of the configuration of a database in which the tolerance level having been subjected to the information processing shown in FIG. 9 is stored. FIG. [Figure 11] 9 is a flowchart showing an example of an information processing procedure of an allowable level executed by the vehicle information processing device of FIG. 8; [Figure 12] 9 is a flowchart showing an example of an information processing procedure for output data executed by the vehicle information processing device of FIG. 8; [Figure 13] 12, (a) is a diagram showing an example of a scene when step S45 of FIG. 12 is reached, (b) is a diagram showing an example of a scene when step S46 of FIG. 12 is reached, and (c) is a diagram showing an example of a scene when step S47 of FIG. 12 is reached. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] <<First embodiment>> Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing an embodiment of a vehicle information processing device 1 according to a first embodiment of the present invention, and FIGS. 2(a) and 2(b) are diagrams of the interior of a vehicle showing an example of a location where an agent is installed. The vehicle information processing device 1 of this embodiment is a device that provides information to a user through an agent function by an anthropomorphized agent (hereinafter, also simply referred to as an agent 4), specifically, through a medium of voice, image, character robot action, and a combination of these, or that converses with the user. In this embodiment, the user refers to a person who uses the vehicle information processing device 1, and although an example applied to a driver of a vehicle will be described below, the user may also be a passenger other than the driver (hereinafter, the driver and the passenger will also be simply referred to as passengers). In addition, an example applied to a scene in which two-way communication is performed in which the agent 4 converses with the user will be described below, but the present invention is not limited thereto, and one-way communication in which the agent 4 presents information to the user may also be performed.
[0010] Here, an example in which the agent 4 is provided in a vehicle is shown, but the form and installation location of the agent 4 are not limited to this. The agent 4 may be any electronic device equipped with an agent function, for example, a portable speaker-type electronic device or an electronic device with a display. Furthermore, the functions related to the audio output and video output of the agent 4 described below may be installed in a mobile phone such as a smartphone. In addition, the personified agent is only one example, and the agent may not be imitated as a human, but may be an agent that displays a predetermined character, avatar, or icon. The agent 4 may be provided as a physical individual, or the shape of a human or character as the agent may be displayed as an image on a display.
[0011] 1, 2(a) and 2(b), the agent 4 of this embodiment is a human-like character robot 42 that is mounted on a base 41 so as to be able to appear and disappear by an actuator (not shown). When the agent 4 receives a control command from the output unit 16 and communicates with the passenger using the agent function, the agent 4 appears from the base 41 as shown in Fig. 2(b), and when communication with the passenger is finished, the character robot 42 is stored in the base 41 as shown in Fig. 2(a).
[0012] The agent 4 includes a speaker or other audio output unit for outputting voice and sound effects, and a display or other display unit for displaying images including text, and outputs communication information by providing voice, sound effects, text, and other images to the occupant along with the movement of the character robot 42. Note that, in this embodiment, the agent 4 is a three-dimensional object such as the character robot 42, but is not limited to this and may be a two-dimensional image displayed on a display.
[0013] Returning to Fig. 1, the vehicle information processing device 1 of this embodiment extracts words contained in voice information acquired from the driver (user), classifies the words into a plurality of concept levels such as superordinate concepts, intermediate concepts, and subordinate concepts, and stores them. Then, as output data of dialogue sentences uttered by the agent 4, output data is generated by including the words classified by the concept levels in the linguistic information. Based on the above, the vehicle information processing device 1 of this embodiment will be described.
[0014] The vehicle information processing device 1 of the present embodiment is configured with a controller 10 including a ROM (Read Only Memory) storing programs for executing various processes, a CPU (Central Processing Unit) as an operating circuit that functions as the vehicle information processing device 1 by executing the programs stored in the ROM, and a RAM (Random Access Memory) that functions as an accessible storage device, an input device 2, a sensor 3, and an agent 4 as an output device. These devices are connected, for example, by a CAN (Controller Area Network) or other in-vehicle LAN, and can transmit and receive information to and from each other. In terms of the functional configuration exhibited by the execution of the information processing program, the controller 10 includes an input data acquisition unit 11, a word processing unit 12, a language information database 13, an occupant determination unit 14, an output data generation unit 15, and an output unit 16, as shown in FIG. 1.
[0015] The input data acquisition unit 11 acquires voice information of the vehicle occupant via the input device 2 and stores it at least temporarily. The input device 2 is, for example, a microphone that is provided in the vehicle and allows voice input. The setting position of the input device 2 is not particularly limited, but it is preferable that the input device 2 is provided near the seat of the occupant. The voice information of the occupant temporarily stored in the input data acquisition unit 11 is output to the word processing unit 12.
[0016] The word processing unit 12 performs voice recognition processing on the occupant's voice information acquired from the input data acquisition unit 11, extracts words contained in the voice information, and classifies them. Specifically, it identifies the driver D's voice information from the occupant's voice information, and performs voice recognition processing on the driver D's voice information. A known technique can be applied to the identification processing of the occupant's voice information (whether it is the driver D or a passenger X other than the driver). In addition, the speaker may be identified from image data captured by an in-vehicle camera as a sensor 3 installed in the vehicle, and the voice information may be identified. In the voice recognition processing, the voice information is digitized to generate text data (character strings), and words contained in the text data are extracted and classified into multiple concept levels.
[0017] For example, as shown in FIG. 3, when voice information "I ate tonkotsu ramen in Atsugi last night" is acquired from driver D, the word processor 12 extracts one or more words such as "Atsugi" and "tonkotsu ramen" from text data generated by digitizing this voice information. Next, the extracted words are classified in association with a plurality of predefined concept levels using a method such as ontology. The concept levels are not particularly limited, but may be "superordinate concept", "middle concept", "subordinate concept", etc. Furthermore, a category is set for each attribute at each concept level.
[0018] Fig. 4 is a diagram showing an example of the configuration of a word database that stores words classified into concept levels. As shown in Fig. 4, for example, for an attribute of "place name", categories such as "region" and "prefecture" are set as the "superordinate concept", "city / ward / town / village" as the "middle concept", and "region" and "facility name" as the "subordinate concept". For an attribute of "gourmet", categories such as "type" and "style" are set as the "superordinate concept", "genre" as the "middle concept", and "food name" and "store name" as the "subordinate concept". Similarly, for an attribute of "date and time", categories such as "year / era" and "month" are set as the "superordinate concept", "day" as the "middle concept", and "day of the week" and "time" as the "subordinate concept".
[0019] Based on these classifications, if the word extracted from the voice information of the driver D is "Atsugi," the word processing unit 12 classifies it as a "middle level concept" since it is a "city, ward, town, or village" with the attribute "place name." Similarly, if the extracted word is "tonkotsu ramen," it classifies it as a "lower level concept" since it is a "dish name" with the attribute "gourmet." The words are stored in a word database together with the classified concept level. The word database is used when the output data generation unit 15 generates output data.
[0020] The language information database 13 stores words extracted from the voice information of the driver D and a word database that stores the conceptual level of the words. In addition to data related to these words, the language information database 13 stores language information such as typical response sentence examples, vocabulary such as words and phrases, and grammar information used by the output data generating unit 15 to generate output data.
[0021] The occupant determination unit 14 identifies the occupants of the vehicle based on the input signal from the sensor 3 and at least temporarily stores the occupants. Specifically, the occupant determination unit 14 identifies the driver D and determines whether or not there is another passenger X other than the driver D within a range of a predetermined distance. The predetermined distance is not particularly limited, but is, for example, a range in which communication with the agent 4 can be performed, and in the case where the agent 4 is provided in the vehicle as in this embodiment, it is within the cabin of the vehicle. The occupant determination unit 14 identifies the driver D based on image data captured by an in-vehicle camera as the sensor 3, for example, and then determines whether or not there is another passenger X (hereinafter, also simply referred to as a passenger) other than the driver D. In the scene shown in FIG. 3, the occupant determination unit 14 identifies a person sitting in the driver's seat as the driver D and identifies a person sitting in the passenger seat as the passenger X1. The result of the determination of whether or not there is a passenger X1 is used to identify the concept level of words used when the output data generation unit 15 generates output data.
[0022] The driver D and the passenger X1 may be identified using an input signal from a seating sensor provided inside the seat of the seat or a seat belt sensor. The driver D may be identified by, for example, storing the driver's identification data in a wireless key of the vehicle and automatically or semi-automatically reading the driver's identification data when the vehicle is unlocked or started using the wireless key. Furthermore, the occupant determination unit 14 may identify the attribute information of the passenger X1 using image data captured by an in-vehicle camera as the sensor 3 or audio information acquired from a microphone in the vehicle. The attribute information of the passenger X1 may include, for example, riding history information such as whether the passenger X1 has a history of riding in the past and how frequently the passenger X1 rides in the vehicle, as well as personal information indicating the relationship between the driver D and the passenger X1.
[0023] When the word processing unit 12 receives the determination result of whether or not the passenger X1 is present from the occupant determination unit 14, it transmits information on the conceptual level of words extracted from the voice information of the driver D together with the determination result to the output data generation unit 15. The output data generation unit 15, which has received this, identifies the conceptual level of words to be used in generating output data based on the determination result of whether or not the passenger X1 is present. Then, it generates text data of a dialogue sentence to be spoken by the agent 4 by including words corresponding to this identified conceptual level in language information such as typical examples of response sentences, vocabulary such as words and phrases, and grammatical information obtained from the language information database 13.
[0024] At this time, when there is a passenger X1 other than the driver D in the vehicle, the output data generating unit 15 generates dialogue text data using words classified at a relatively higher concept level compared to when there is no passenger X1. The dialogue content becomes more specific as it uses words with lower concepts, and in some cases is more likely to include privacy-related content such as facility names, person names, and disease names. In contrast, the dialogue content becomes more abstract as it uses words with higher concepts. Therefore, when there is no passenger X1 other than the driver D in the vehicle, the dialogue may include words with lower concepts, but when there is passenger X1, the dialogue text data is generated using words with higher concepts.
[0025] In the scene shown in FIG. 7(a), since a passenger X1 other than the driver D is present in the vehicle, the output data generating unit 15 generates text data of a dialogue such as "Was the Chinese food in Kanagawa Prefecture that we went to last night delicious?" using words classified as higher concepts, and has the agent 4 speak the dialogue. In contrast, in the scene shown in FIG. 7(b), since a passenger X1 other than the driver D is not present in the vehicle, the output data generating unit 15 generates text data of a dialogue such as "Was the tonkotsu ramen at □□-ken (restaurant name) that we went to last night delicious?" using words classified as lower concepts, and has the agent 4 speak the dialogue. In this way, when the passenger X1 is present, the agent 4 can appropriately communicate with the occupants while protecting the privacy of the driver D by using words classified as higher concepts compared to when the passenger X1 is not present.
[0026] Even if there is no passenger X1 other than the driver D, it is preferable to generate the text data of the dialogue using words classified at a concept level equivalent to the concept level of the words included in the voice information of the driver D. For example, if the words included in the voice information of the driver D are subordinate concepts, the words classified at the subordinate concepts are used.
[0027] As described above, the words of the lower concept are specific and may include contents related to privacy. If the words included in the voice information of the driver D are lower concepts, it can be determined that the driver D is willing to communicate with the agent 4 including these specific contents and contents related to privacy. In contrast, if the words included in the voice information of the driver D are higher concepts, it is possible that the driver D does not want to communicate including specific contents and contents related to privacy. Therefore, if the words included in the voice information of the driver D are higher concepts, but the agent 4 speaks using words of lower concepts, the driver D may have an impression that the agent 4 is familiar or bothersome. Therefore, when the words included in the voice information of the driver D are lower concepts, the words classified as lower concepts are included in the language information to generate text data of the dialogue. This prevents the agent 4 from including unnecessary information in his / her speech, which makes it easier for the driver D to have a favorable impression of the agent 4.
[0028] The output data generation unit 15 converts the text data of the dialogue into voice data by voice synthesis processing, and transmits this to the output unit 16 as output data for the dialogue spoken by the agent 4. A known technique can be applied to the voice synthesis processing.
[0029] When the output unit 16 receives output data from the output data generation unit 15, it outputs a control signal to the speaker or other audio output unit of the agent 4, the display or other display unit, and outputs communication information (here, audio information of a dialogue responding to the driver D) using the agent function of the agent 4.
[0030] Next, an information processing procedure of the vehicle information processing device 1 of this embodiment will be described. Fig. 5 is a flowchart showing an information processing procedure of input data from the driver D, which is executed by the vehicle information processing device 1 of Fig. 1. This is a flowchart showing an example of information processing of words included in the voice information of the driver D, which is executed by the word processing unit 12.
[0031] First, in step S1 of Fig. 5, when the ignition switch of the vehicle is turned ON, the following information processing is executed. In step S2, the occupant determination unit 14 identifies the driver D based on the information acquired from the sensor 3. Next, in step S3, the input data acquisition unit 11 acquires voice information of the driver D. In the following step S4, the word processing unit 12 extracts words from the voice information of the driver D. In the scene shown in Fig. 3, when voice information that "I ate tonkotsu ramen in Atsugi last night" is acquired from the driver D, words such as "Atsugi" and "tonkotsu ramen" are extracted from this voice information.
[0032] Next, in step S5, the word processor 12 classifies the extracted words according to their concept levels. "Atsugi" is classified as a "middle-level concept" because it is a "city, ward, town, or village" with an attribute of "place name." "Tonkotsu ramen" is classified as a "dish name" with an attribute of "gourmet," so it is classified as a "lower-level concept."
[0033] In step S6, the word processor 12 stores the word and the concept level into which the word is classified in the word database. Then, in step S7, when the vehicle ignition switch is turned off, the information processing of the input data from the driver D is terminated. In response to this, the processing from step S3 to step S6 is repeated until the vehicle ignition switch is turned off.
[0034] Fig. 6 is a flowchart showing an information processing procedure of output data executed by the vehicle information processing device 1 of Fig. 1. This is a flowchart showing an example of information processing of output data of a dialogue uttered by the agent 4, generated by the output data generating unit 15. Note that the information processing of the input data from the driver D shown in Fig. 5 and the information processing of the output data shown in Fig. 6 may be processed in parallel by the vehicle information processing device 1.
[0035] First, when an occupant gets into the vehicle, in step S21 of Fig. 6, the occupant determination unit 14 identifies the driver D based on information acquired from the sensor 3. In the following step S22, the occupant determination unit 14 determines whether or not there is a passenger X1 other than the driver D.
[0036] If it is determined in step S22 that there is a passenger X1 other than the driver D, the process proceeds to step S23. In step S23, the output data generating unit 15 generates text data of a dialogue such as "Was the Chinese food in Kanagawa Prefecture that we went to last night delicious?" by including the words classified as the superordinate concept in the linguistic information, as shown in Fig. 7(a).
[0037] On the other hand, if the result of the judgment in step S22 is that there is no passenger X1 other than the driver D, the process proceeds to step S26, and the output data generation unit 15 includes the words classified as subordinate concepts in the linguistic information and generates text data of a dialogue such as "Was the tonkotsu ramen at □□ken (restaurant name) that I went to last night delicious?" as shown in Figure 7(b).
[0038] In step S24, the output data generating unit 15 converts the text data of the dialogue into voice data by voice synthesis processing to generate output data. In the following step S25, the output data is output by the agent function of the agent 4 as voice information of the dialogue responding to the driver D.
[0039] As described above, according to the information processing device 1 and the information processing method of the present embodiment, input data from the driver D (user) is acquired, words included in the input data from the driver D (user) are extracted, the words are classified into a plurality of concept levels including at least a higher concept level and a lower concept level that is relatively lower than the higher concept level, and stored, output data is generated by including words classified into any of the concept levels in the linguistic information, and the output data is output by the agent function. At that time, it is determined whether or not another passenger X1 (other person) other than the driver D (user) is present within a predetermined distance, and when it is determined that another passenger X1 (other person) is present within the predetermined distance, output data is generated by including words classified into a relatively higher concept level in the linguistic information, compared to when it is determined that another passenger X1 (other person) is not present within the predetermined distance. This allows appropriate communication while protecting the privacy of the user.
[0040] Furthermore, according to the information processing device 1 and information processing method of this embodiment, the output data generation unit 15 generates output data as dialogue data responding to input data from the driver D (user), thereby enabling two-way communication between the user and the agent 4.
[0041] Furthermore, according to the information processing device 1 and the information processing method of this embodiment, the range of the specified distance is the interior of the vehicle in which the driver D (user) is riding, and the other person is a passenger X1 other than the driver D (user) who is riding in the vehicle, so that appropriate communication can be carried out while protecting the user's privacy during communication inside the vehicle.
[0042] Furthermore, according to the information processing device 1 and the information processing method of this embodiment, when there is no other passenger X1 (other person) within a predetermined distance and the conceptual level into which the words included in the input data from the driver D (user) are classified is a relatively lower conceptual level, the output data generating unit 15 generates output data by including the words of the relatively lower conceptual level in the linguistic information. This prevents unnecessary information from being included in the utterance when communication information is provided using the agent 4, making it easier for the user to have a favorable impression of the agent 4.
[0043] <<Second embodiment>> Next, a second embodiment of the present invention will be described with reference to Figs. 8 to 13. In the second embodiment, the controller 10 of the vehicle information processing device 1 is provided with an allowable level setting unit 17 and a passenger database 18. In this embodiment, the degree of intimacy between the driver D and the passenger X is determined, and an allowable level is set for each passenger X and stored in the passenger attribute database. When generating output data of a dialogue uttered by the agent 4, the output data is generated by including words of a concept level corresponding to the allowable level in the language information. Note that the configurations of the input data acquisition unit 11, the word processing unit 12, the language information database 13, the passenger determination unit 14, the output data generation unit 15, and the output unit 16 of the vehicle information processing device 1, the input device 2, the sensor 3, and the agent 4 are the same as those of the first embodiment shown in Fig. 1, and therefore the description of these blocks will be cited from the above-mentioned embodiment.
[0044] The tolerance level setting unit 17 associates information on the concept level of words included in the voice information of the driver D received from the word processing unit 12 with attribute information of the passenger X received from the passenger determination unit 14 to determine the degree of intimacy between the driver D and the passenger X, and sets the tolerance level of the passenger X based on the degree of intimacy. The attribute information of the passenger X includes riding history information such as whether the passenger X has ridden in the past and how frequently he / she rides, as well as personal information indicating the relationship between the driver D and the passenger X. The personal information indicating the relationship between the driver D and the passenger X is not particularly limited, but may be "family," "friend," "workplace," etc. The relationship between the driver D and the passenger X may be registered in advance by the driver D, or may be registered automatically or semi-automatically when the passenger X is recognized to have ridden for the first time by an in-vehicle camera. The tolerance level is the concept level of words that are permitted to be used when generating output data when the passenger X is present in the vehicle.
[0045] For example, in the scene shown in Fig. 9(a), when a driver D and a passenger X2 get into a vehicle and voice information is acquired from the driver D saying "I ate tonkotsu ramen at □□ken (store name) last night," the word processing unit 12 extracts the words "□□ken (store name)" and "tonkotsu ramen" by the above-mentioned voice recognition process and classifies these words as "subordinate concepts." The tolerance level setting unit 17 determines that the intimacy between the driver D and the passenger X2 is "high" because the driver D is conversing using words of subordinate concepts including specific contents, and sets the tolerance level of the passenger X2 to "subordinate concepts."
[0046] In the scene shown in FIG. 9(b), the driver D and the passenger X3 are in the vehicle, and the driver D has voice information saying, "I ate ramen in Yokohama last night." The word processor 12 executes voice processing, extracts the words "Yokohama" and "ramen," and classifies these words as "medium concepts." The tolerance level setting unit 17 determines that the intimacy between the driver D and the passenger X3 is "medium" because the driver D uses words of medium concepts, and sets the tolerance level of the passenger X3 to "medium concept." Similarly, in the scene shown in FIG. 9(c), when the driver D has voice information saying, "I ate Chinese food in Kanagawa Prefecture last night," the word processor 12 extracts the words "Kanagawa Prefecture" and "Chinese," and classifies these words as "higher concepts." Since the driver D is conversing in abstract terms using words with higher concepts, the tolerance level setting unit 17 determines that the intimacy between the driver D and the passenger X4 is "low" and sets the tolerance level for the passenger X4 to "higher concept."
[0047] In this way, the tolerance level set based on the intimacy between the driver D and the passenger X is stored in the passenger attribute database together with the attribute information of the passenger X. When the output data generating unit 15 generates output data, the output data is generated by including in the linguistic information words at a conceptual level corresponding to this tolerance level, so that the agent 4 can provide communication information taking into account the intimacy between the driver D and the passenger X.
[0048] FIG. 10 is a diagram showing an example of the configuration of a passenger attribute database that stores attribute information of passenger X and an acceptable level set based on the intimacy between driver D and passenger X. As shown in FIG. 10, for example, for an attribute of "family", the intimacy with driver D, the acceptable level, the riding frequency, etc. are stored for each specified person such as "father" and "mother". For example, for "mother", the intimacy is stored as "high", the acceptable level is stored as "subordinate concept", and the riding frequency is stored as "medium". Note that for a passenger who has boarded a vehicle for the first time, there is no record in the passenger attribute database, so new data may be created and stored, for example, with the intimacy as "low" and the acceptable level as "superordinate concept".
[0049] The passenger database 18 stores a passenger attribute database that stores attribute information of the passenger X and an acceptable level set based on the degree of intimacy between the driver D and the passenger X. In addition to this information, image data of the passenger X and information such as topics that are often used when the driver D talks with the passenger X (such as the results of professional baseball games and information about favorite stores) may be stored.
[0050] Returning to Fig. 8, when the occupant determination unit 14 determines that there is a passenger X other than the driver D based on image data captured by the in-vehicle camera as the sensor 3, it identifies the passenger X. Then, it transmits attribute information of the identified passenger X to the tolerance level setting unit 17. The tolerance level setting unit 17 searches the passenger database 18 based on the attribute information of the passenger X, acquires information on the tolerance level of the passenger X, and transmits it to the occupant determination unit 14. When the occupant determination unit 14 acquires the information on the tolerance level of the passenger X from the tolerance level setting unit 17, it transmits it to the word processing unit 12 together with the attribute information of the passenger X.
[0051] When the word processing unit 12 receives the attribute information of the passenger X and the tolerance level of the passenger X from the occupant determination unit 14, it transmits this information to the output data generation unit 15. The output data generation unit 15, which has received this information, generates text data of a dialogue sentence to be spoken by the agent 4, using words classified into a concept level corresponding to the tolerance level of the passenger X.
[0052] In the scene shown in FIG. 13(a), the tolerance level of the passenger X2 is "lower concept", so the output data generating unit 15 generates text data of a dialogue such as "Was the tonkotsu ramen at □□-ken (restaurant name) that I went to last night delicious?" using words classified as lower concepts, and the agent 4 speaks it. In the scene shown in FIG. 13(b), the tolerance level of the passenger X3 is "middle concept", so the output data generating unit 15 generates text data of a dialogue such as "Was the ramen at Yokohama that I went to last night delicious?" using words classified as middle concepts. Similarly, in the scene shown in FIG. 13(c), the tolerance level of the passenger X4 is "higher concept", so the output data generating unit 15 generates text data of a dialogue such as "Was the Chinese food at Kanagawa that I went to last night delicious?" using words classified as higher concepts.
[0053] In this way, output data is generated using words categorized into a concept level corresponding to the tolerance level of passenger X, which is set based on the intimacy between the driver D and passenger X, so that it is possible to provide appropriate communication information according to the intimacy between the driver D and passenger X. Note that the process of converting the text data of the dialogue into voice data and generating output data uttered by the agent 4 is similar to that of the first embodiment, and therefore the above description is used here.
[0054] The tolerance level set by the tolerance level setting unit 17 may be configured to be changed when the passenger X gets in the car again after being set once for the passenger X. This is because by changing the setting of the tolerance level as needed, appropriate output data corresponding to the change in the intimacy between the driver D and the passenger X can be generated. The change in the intimacy between the driver D and the passenger X is determined, for example, by comparing the information on the concept level of words included in the voice information of the driver D received from the word processing unit 12 with the tolerance level of the passenger X stored in the passenger attribute database. If there is a difference between the information on the concept level of words included in the voice information of the driver D and the tolerance level of the passenger X, the tolerance level setting unit 17 determines that the intimacy has changed. In addition to this, the change in the intimacy may be determined by identifying the frequency of riding of the passenger X and the concept level of words included in the voice information of the passenger X.
[0055] For example, in the case of "friend C" shown in FIG. 10, the concept level of the word included in the voice information of the driver D received from the word processing unit 12 is set to "subordinate concept", and the tolerance level is set to "medium concept". In such a case, the tolerance level setting unit 17 determines that the intimacy between the driver D and the friend C has changed from "medium" to "high", changes the tolerance level to "subordinate concept" (black arrow), and updates the passenger attribute database. In contrast, for "senior Y", the concept level of the word included in the voice information of the driver received from the word processing unit 12 is set to "subordinate concept", and the tolerance level is set to "subordinate concept", so it determines that the intimacy between the driver D and senior Y has changed from "high" to "medium", changes the tolerance level to "medium concept" (white arrow), and updates the passenger attribute database.
[0056] In addition, when changing the allowable level to a lower concept level, it is preferable that the concept level of the words included in the voice information of the driver D is a lower concept level than the currently set allowable level. As described above, the words classified as lower concepts are specific and include privacy-related content. When the words included in the voice information of the driver D are lower concepts, that is, when the driver is talking using words of lower concepts, it can be determined that the driver is willing to communicate with the passenger X and the agent 4, including these specific contents and privacy-related contents. On the other hand, for example, even if the passenger X is talking using words of lower concepts, when the driver D uses words of higher concepts, the communication information output by the agent 4 should also use words of higher concepts, which will be communication in line with the driver D's intention. In this way, when the words included in the voice information of the driver D are at a lower concept level, the provision of unnecessary information by the agent 4 is suppressed by changing the allowable level to a lower concept level. Furthermore, since communication is performed in line with the driver D's intention, the driver D is more likely to feel close to the agent 4.
[0057] Furthermore, the tolerance level setting unit 17 may set the tolerance level of the passenger X to "no limit" if a predetermined condition is satisfied. The predetermined condition is not particularly limited, but may be, for example, a case where the tolerance level is "subordinate concept" for a certain period of time and the riding frequency is "high". The tolerance level being "no limit" means that there is no limit to the concept level of the words used when generating the output data, and the output data generating unit 15 can generate the output data in the same manner as when only the driver D is riding. In the example shown in FIG. 10, since "Friend A" has a tolerance level of "subordinate concept" and a riding frequency of "high", the tolerance level is changed from "subordinate concept" to "no limit" (diagonal arrow). When the tolerance level of the passenger X is "no limit", the passenger determining unit 14 may be configured not to identify the passenger X with "no limit" as a passenger X (other person) other than the driver D. This facilitates communication between the driver D, the passenger X, and the agent 4.
[0058] Next, the information processing procedure of the vehicle information processing device 1 of this embodiment will be described. Note that the information processing of the input data from the driver D shown in FIG. 5 is similar to that of the first embodiment, and therefore the description in the above embodiment will be used. FIG. 11 is a flowchart showing an example of the information processing procedure for setting the tolerance level of the passenger X, which is executed by the vehicle information processing device 1 of FIG. 8. The following description will be given with reference to the scene in FIG. 9(a).
[0059] First, in step S31 of FIG. 11, when the ignition switch of the vehicle is turned ON, the following information processing is executed. In step S32, the occupant determination unit 14 identifies the driver D and a passenger X2 other than the driver D based on the information acquired from the sensor 3. Next, in step S33, the input data acquisition unit 11 acquires voice information of the driver D. In the following step S34, the word processing unit 12 extracts words from the voice information of the driver D. For example, as shown in FIG. 9(a), when voice information such as "I ate tonkotsu ramen at □□ken (store name) last night" is acquired from the driver D, words such as "□□ken (store name)" and "tonkotsu ramen" are extracted from the voice information.
[0060] Next, in step S35, the word processing unit 12 identifies the concept level of the extracted word. If the word is "□□ken (store name)" or "tonkotsu ramen", it is identified as a "subordinate concept". Then, in step S36, the tolerance level setting unit 17 determines that the intimacy between the driver D and the passenger X2 is "high" because the concept level of the word is a "subordinate concept", and sets the tolerance level to "subordinate concept". If a different tolerance level has already been stored for the passenger X2 in the passenger attribute database, the tolerance level is changed and stored.
[0061] When the ignition switch of the vehicle is turned OFF in step S37, the information processing for setting the tolerance level for passenger X2 is terminated. In response to this, the processing from step S33 to step S36 is repeated until the ignition switch of the vehicle is turned OFF. Note that the setting and changing of the tolerance level may be configured to be repeated at predetermined intervals while passenger X2 is in the vehicle, or the information processing may be configured to be executed once each time passenger X2 gets in the vehicle.
[0062] Fig. 12 is a flowchart showing an information processing procedure of output data executed by the vehicle information processing device 1 of Fig. 8. This is a flowchart showing an example of information processing of output data of a dialogue uttered by the agent 4, generated by the output data generating unit 15. Note that the information processing of the output data shown in Fig. 12, the information processing of the input data from the driver D shown in Fig. 5, and the information processing for setting the tolerance level of the passenger X shown in Fig. 11 may be processed in parallel by the vehicle information processing device 1.
[0063] First, when an occupant gets into the vehicle, in step S41 of Fig. 12, the occupant determination unit 14 identifies the driver D based on information acquired from the sensor 3. In the following step S42, the occupant determination unit 14 determines whether or not there is a passenger X other than the driver D. If it is determined in step S42 that there is a passenger X other than the driver D, the process proceeds to step S43. On the other hand, if it is determined that there is no passenger X other than the driver D, the process proceeds to step S45.
[0064] If it is determined in step S42 that there is a passenger X other than the driver D, in step S43, the occupant determination unit 14 identifies the passenger X and transmits attribute information of the passenger X to the tolerance level setting unit 17, and the process proceeds to step S44.
[0065] In step S44, the tolerance level setting unit 17 searches the passenger database 18 based on the attribute information of the passenger X, and acquires the tolerance level of the passenger X. If the tolerance level of the passenger X acquired in step S44 is a "subordinate concept", the process proceeds to step S45. If it is determined in step S42 that there is no passenger X other than the driver D, and if the tolerance level of the passenger X acquired in step S44 is a "subordinate concept", in step S45, the output data generating unit 15 generates text data of a dialogue such as "Was the tonkotsu ramen at □□ken (store name) that I went to last night delicious?" by including the words classified as subordinate concepts in the linguistic information, as shown in FIG. 13(a).
[0066] If the tolerance level of passenger X acquired in step S44 is a "middle concept", the process proceeds to step S46, where the words classified as the middle concept are included in the linguistic information to generate text data of a dialogue such as "Was the ramen in Yokohama that we went to last night delicious?" as shown in Fig. 13(b). If the tolerance level of passenger X acquired in step S44 is a "higher concept", the process proceeds to step S47, where the words classified as the higher concept are included in the linguistic information to generate text data of a dialogue such as "Was the Chinese food in Kanagawa that we went to last night delicious?" as shown in Fig. 13(c).
[0067] In step S48, the output data generating unit 15 converts the text data of the dialogue into voice data by voice synthesis processing to generate output data. In the following step S49, the output data is output by the agent function of the agent 4 as voice information of the dialogue responding to the driver D.
[0068] As described above, according to the information processing device 1 and the information processing method of the present embodiment, the attributes of the passenger X (other person) are identified and stored, and the conceptual level of words that are permitted to be used in generating output data is set as an allowable level for each passenger X (other person) based on the attributes of the passenger X (other person) and the conceptual level into which words included in the input data from the driver D (user) are classified. Then, when the passenger X (other person) is present within a predetermined distance and the attributes of the passenger X (other person) are stored in the passenger database 18, output data is generated by including words corresponding to the allowable level in the linguistic information. This allows the agent 4 to provide appropriate communication information according to the degree of intimacy between the user and the other person.
[0069] Furthermore, according to the information processing device 1 and the information processing method of the present embodiment, the tolerance level setting unit 17 changes the setting of the tolerance level for the specific passenger X (other person) based on words included in the input data from the driver D (user) and other predetermined conditions after the tolerance level for the specific passenger X (other person) is set. This allows the agent 4 to provide appropriate communication information corresponding to the change in the intimacy between the user and the other person.
[0070] Furthermore, according to the information processing device 1 and information processing method of this embodiment, the setting of the tolerance level is changed based on the conceptual level at which the words contained in the input data from the driver D (user) are classified, so that communication information that is in line with the user's intentions can be provided, making it easier for the user to feel close to the agent 4.
[0071] According to the information processing device 1 and the information processing method of the present embodiment, when the concept level into which the words included in the input data from the driver D (user) are classified is relatively lower than the tolerance level, the tolerance level setting unit 17 changes the tolerance level to the relatively lower concept level. This prevents unnecessary information from being provided when communication information is provided using the agent 4.
[0072] Furthermore, according to the information processing device 1 and information processing method of this embodiment, the occupant determination unit 14 does not identify a passenger X (other person) who satisfies an acceptable level or other predetermined conditions as a person other than the driver D (user), thereby facilitating communication among the user, the other person, and the agent 4.
[0073] The above-described embodiments are described for the purpose of facilitating understanding of the present invention, and are not described for the purpose of limiting the present invention. Therefore, the elements disclosed in the above embodiments are intended to include all design modifications and equivalents that fall within the technical scope of the present invention. [Explanation of symbols]
[0074] 1...Information processing device for vehicle 10…Controller 11...Input data acquisition section 12...Word processing section 13…Language information database 14…Occupant Judgment Section 15...Output data generation unit 16...Output section 17...Tolerance level setting section 18…Passenger database 2. Input device 3. Sensor 4. Agent D…Driver (user) X, X1, X2, X3, X4... Passengers (other people)
Claims
1. an input data acquisition unit that acquires input data from a user; a word processing unit that extracts words included in the input data from the user, classifies the words into a plurality of concept levels including at least a superordinate concept level and a subordinate concept level that is relatively subordinate to the superordinate concept level, and stores the result; an output data generation unit that generates output data by including words classified into any one of the concept levels in linguistic information; an output unit that outputs the generated output data to the user by an agent function; a determination unit that determines whether or not a person other than the user is present within a predetermined distance range, The output data generation unit When the judgment unit determines that the other person is present within the specified distance range, the information processing device generates the output data by including in the linguistic information the words classified at a relatively higher conceptual level compared to when the judgment unit determines that the other person is not present within the specified distance range.
2. The information processing apparatus according to claim 1 , wherein the output data generating unit generates the output data as dialogue data in response to input data from the user.
3. The information processing device according to claim 1 , wherein the range of the predetermined distance is within a vehicle in which the user is riding, and the other person is a passenger other than the user who is riding in the vehicle.
4. a third party attribute storage unit that stores the third party attributes; and an allowance level setting unit that sets, for each of the other people, the conceptual level of the word that is allowed to be used in generating the output data as an allowance level based on the attribute of the other people and the conceptual level into which the word included in the input data from the user is classified, When the other person is present within the range of the predetermined distance and the attribute of the other person is stored in the other person attribute storage unit, 4. The information processing apparatus according to claim 1, wherein the output data generating section generates the output data by including the word corresponding to the tolerance level in the language information.
5. The information processing device according to claim 4 , wherein the tolerance level setting unit, after the tolerance level is set for a specific other person, changes the setting of the tolerance level for the specific other person based on the word contained in the input data from the user and other specified conditions.
6. The information processing apparatus according to claim 5 , wherein the tolerance level setting unit changes a setting of the tolerance level based on the concept level into which the words included in the input data from the user are classified.
7. 7. The information processing device according to claim 5, wherein the tolerance level setting unit changes the tolerance level to the relatively lower conceptual level when the conceptual level into which the word included in the input data from the user is classified is a conceptual level relatively lower than the tolerance level.
8. 8. The information processing apparatus according to claim 4, wherein the determination unit does not identify the other person who satisfies the tolerance level or other predetermined conditions as a person other than the user.
9. An information processing device as described in any one of claims 1 to 8, wherein when the other person is not present within the specified distance and the conceptual level to which the word contained in the input data from the user is classified is a relatively lower conceptual level, the output data generation unit generates the output data by including the word of the relatively lower conceptual level in the linguistic information.
10. An information processing method for causing a processor to execute a process, comprising the steps of: Get input data from the user, extracting words included in the input data from the user, classifying the words into a plurality of concept levels including at least a superordinate concept level and a subordinate concept level that is relatively subordinate to the superordinate concept level, and storing the words; generating output data by including words classified into any one of the concept levels in linguistic information; In the information processing method, the output data is output to the user by an agent function, The processor, determining whether or not a person other than the user is present within a predetermined distance range; An information processing method for generating output data by including in the linguistic information the words classified at a relatively higher conceptual level when it is determined that the other person is present within the specified distance range, compared to when it is determined that the other person is not present within the specified distance range.
Citation Information
Patent Citations
Dialogue system
JP2019101805A
Information processing apparatus, program and information processing method
JP2019200669A