System
A system that analyzes user interests and location to provide engaging voice interactions synchronized with navigation instructions addresses driver drowsiness, ensuring safe and comfortable long-distance driving.
Patent Information
- Application Number
- JP2024116548
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Driver drowsiness during long-distance driving, especially when driving alone, is a significant safety concern that conventional methods like music or radio fail to adequately address, and there is a need for a system that provides engaging topics based on the driver's interests while ensuring safety by linking with car navigation devices.
A system that analyzes a user's past activities and interests, selects relevant topics based on location information, engages in voice interaction, and synchronizes voice utterances with car navigation to prevent drowsiness by pausing interactions during navigation instructions.
The system effectively prevents driver drowsiness by providing engaging conversations tailored to individual interests and ensures safe driving by synchronizing voice interactions with navigation instructions, enhancing the overall driving experience.
Smart Images

Figure 2026015074000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Driver drowsiness during long-distance driving, especially when driving alone, is a serious problem that affects safe driving. This problem often cannot be completely solved by conventional passive methods such as music or radio. Furthermore, to prevent drowsiness by chatting while driving, it is important to provide topics based on the driver's interests. Furthermore, it is necessary to ensure safety by linking with car navigation devices. Therefore, a system that can more effectively prevent driver drowsiness is required. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems with a system including: a means for analyzing a user's interests based on past activities and providing appropriate topics; a means for acquiring location information and selecting topics related to that information; a means for engaging in voice interaction with the user using the selected topics; and a means for controlling synchronization of voice utterances with a car navigation device. Specifically, the system collects and analyzes the user's past search history and purchase history to generate topics that interest the user and engage in conversation with the user based on those topics. Meanwhile, when the car navigation device issues voice instructions, the system automatically stops the interaction to ensure driving safety. This allows the driver to avoid drowsiness through interesting conversations and enjoy a safer and more comfortable drive.
[0006] "User" refers to a driver who uses the system.
[0007] "Past activity" refers to usage information such as a user's past search history and purchase history.
[0008] "Means of analyzing interests and providing appropriate topics" refers to a function that analyzes a user's interests and concerns based on past activity, and then selects and provides topics that are likely to interest the user.
[0009] "Means for obtaining location information and selecting topics related to that information" refers to a function for obtaining the user's current geographic location and selecting and providing interesting topics related to that location.
[0010] "Means for voice interaction" refers to a function that provides a selected topic to a user by voice and exchanges back and forth with the user in a conversational format.
[0011] "Means for controlling synchronization of voice utterances with the car navigation device" refers to a function in which the system automatically stops interaction with the user when the car navigation device gives a voice instruction, and resumes interaction after the instruction is completed. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0020] [First embodiment]
[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0033] System Overview
[0034] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to allow them to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities and provides appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, thereby ensuring driving safety.
[0035] Program processing overview and specific examples
[0036] Initial Setup
[0037] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[0038] 1. The server sends a prompt to the device asking the user for consent to data collection.
[0039] 2. When the user presses the consent button, the information is sent back to the server.
[0040] 3. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[0041] Example: If a user has made a lot of travel-related searches and purchases in the past, use that information to prioritize travel topics.
[0042] Obtaining location information
[0043] The device obtains its current location and sends that information to the server.
[0044] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[0045] 2. The device sends the acquired location information to the server.
[0046] Example: If the user is near a tourist attraction in Tokyo, the current location information is sent to the server that the user is near Tokyo Tower.
[0047] Topic selection
[0048] The server selects topics based on interests and location.
[0049] 1. The server matches the received location information with the user's interest list.
[0050] 2. The server picks up topics related to the user's interests and location information and sends them to the device.
[0051] Example: If the current location is near Tokyo Tower and the user is interested in architecture, select the topic "The history of the architecture of Tokyo Tower."
[0052] Start a conversation
[0053] The device will interact via voice based on the selected topic.
[0054] 1. The device uses speech synthesis technology to speak to the user based on the topic data received from the server.
[0055] Example: The device says, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[0056] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[0057] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0058] Linking with car navigation systems
[0059] The device synchronizes with the car navigation system's speech.
[0060] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0061] 2. When the car navigation system issues voice instructions, the device automatically interrupts interaction with the user.
[0062] Example: The device pauses voice conversation while the car navigation system says, "Turn right at the next traffic light."
[0063] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0064] The system takes into account the user's individual interests and current geographical environment to provide a safe and comfortable drive while preventing drowsiness while driving.
[0065] The processing flow will be explained below.
[0066] Step 1:
[0067] The server confirms the user's consent
[0068] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[0069] 2. The device displays a consent confirmation screen to the user.
[0070] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0071] Step 2:
[0072] The server collects and analyzes the user's past activity data.
[0073] 1. Once the server receives the user's consent, it will collect usage history data for LINE Yahoo! services (e.g., search history, purchase history).
[0074] 2. The server analyzes the collected data and identifies the user's areas of interest.
[0075] Example: If a user searches for and purchases a lot of automotive-related items, tag their interests as automotive-related.
[0076] 3. The server generates a priority list of interests based on the analysis results.
[0077] Step 3:
[0078] The device acquires location information using GPS
[0079] 1. The device periodically receives GPS signals and obtains its current location information.
[0080] 2. The device sends the acquired location information to the server.
[0081] Example: If your current location is a tourist spot in a city, send that information.
[0082] Step 4:
[0083] The server selects topics based on interests and location information
[0084] 1. The server matches the location information with the user's interest list.
[0085] 2. The server selects interesting topics related to the location.
[0086] Example: If the location is near Tokyo Tower and the user is interested in architecture, select topics related to the history of architecture.
[0087] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[0088] Step 5:
[0089] The device interacts with the user through voice
[0090] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0091] Example: Say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[0092] 2. When the user responds by voice, the device analyzes the voice.
[0093] 3. The device sends the analysis results to the server and requests the next conversation content.
[0094] Step 6:
[0095] The terminal monitors and controls the voice instructions of the car navigation device
[0096] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0097] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0098] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0099] For example: "Let's talk about the construction of Tokyo Tower."
[0100] This allows users to avoid drowsiness while driving and enjoy conversations about interesting topics. In addition, safe driving is ensured by linking with car navigation systems.
[0101] Example 1
[0102] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0103] Conventional long-distance driving assistance systems have had problems with drivers feeling drowsy, reducing driving safety. They also lacked entertainment and information provision while driving, making it impossible to provide topics that matched the driver's interests. Furthermore, if voice utterances were not properly synchronized with the car navigation system, the driver's attention could be distracted, potentially leading to an accident.
[0104] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0105] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, means for collecting data with the user's consent, and means for analyzing the user's responses and generating the next topic or response based on the analysis results. This allows for a safe and comfortable drive while preventing drowsiness while taking into account the user's individual interests and current geographical environment.
[0106] "Means of analyzing a user's interests and concerns based on their past activities and providing appropriate topics" refers to a system that collects past behavioral data such as a user's search history and purchase history, analyzes this data to identify the user's interests and concerns, and automatically selects and provides related topics based on that.
[0107] "Means for obtaining location information and selecting topics related to that information" refers to a mechanism in which the user's device obtains current geographical location information using GPS or other means, and selects interesting related topics based on this location information.
[0108] "Means for interacting with the user via voice using a selected topic" refers to a system that uses voice synthesis technology to provide information to the user via voice based on a selected topic, and allows the user to respond via voice, thereby achieving two-way communication.
[0109] The "means for controlling synchronization of voice utterances with the car navigation device" is a mechanism for starting or interrupting voice interaction with the user at appropriate timing in synchronization with voice instructions from the car navigation device.
[0110] "Means for collecting data with user consent" refers to a mechanism by which the system sends a prompt to the user requesting consent for data collection, obtains consent from the user, and collects the necessary data based on that consent.
[0111] "Means for analyzing the user's response and generating the next topic or response based on the analysis results" refers to a mechanism that analyzes the content of the user's voice response and generates the next topic or an appropriate response based on the analysis results.
[0112] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user through voice and further controls the synchronization of voice utterances with the car navigation system to ensure driving safety.
[0113] Initial Setup
[0114] The server confirms the user's consent and collects the necessary data. First, the server sends a prompt to the device to obtain the user's consent. This prompt includes a message such as "Please consent to data collection." When the user presses the consent button, the information is returned to the server, which stores it in a database. The server then collects the user's past search history, purchase history, etc., and uses this data to analyze areas of interest and generate a topic list. A natural language processing algorithm is preferably used in this process. Specific techniques that can be used include data analysis tools such as Python and R.
[0115] Obtaining location information
[0116] The device obtains its current location and sends that information to the server. The device (car navigation system or smartphone) periodically obtains its current location information using a GPS module. For example, it can be set to update its location information every 5 minutes. The obtained location information is sent to the server via an HTTP POST request, and the server stores this information in a database.
[0117] Topic selection
[0118] The server compares the received location information with the user's interest list and selects a topic based on that interest. The server uses database searches to compare the location information with the user's areas of interest. The selected topic is sent to the device in JSON format or similar. For example, if the current location is near Tokyo Tower and the user is interested in architecture, the topic "The architectural history of Tokyo Tower" will be selected.
[0119] Start a conversation
[0120] The device uses speech synthesis technology to speak to the user based on the selected topic. Specifically, it uses the Google Text-to-Speech API to say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" When the user responds verbally, the device analyzes this speech and converts it into text data using the Google Speech-to-Text API or similar. The analyzed data is sent to the server, which generates the next topic and response based on the analysis results. At this stage, a generative AI model (such as GPT-4) is used.
[0121] Linking with car navigation systems
[0122] The device constantly monitors voice instructions from the car navigation system. This is achieved through API integration between the car navigation system and the device. When the car navigation system issues a voice instruction, the device automatically suspends interaction with the user. For example, when the car navigation system says, "Turn right at the next traffic light," the voice conversation is paused. When the car navigation system finishes issuing the instruction, the device resumes the conversation with the user.
[0123] Specific examples and examples of prompts for generative AI models
[0124] As a concrete example, if a user asks, "How tall is Tokyo Tower?", the server will generate a response based on the analysis results and reply, "Tokyo Tower is approximately 333 meters tall."
[0125] An example of a prompt for a generative AI model is:
[0126] "Based on your past search history, what are some travel topics that have recently interested you?"
[0127] This system takes into account the user's individual interests and geographical environment, making it possible to provide a safe and comfortable drive while preventing drowsiness while driving.
[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0129] Step 1: Confirm user agreement
[0130] The server sends a prompt to the device to obtain the user's consent. This uses an HTTP request and includes the message "Please consent to data collection." The input is a consent confirmation message, and the output is a prompt display. Specifically, the server sends an HTTP request to the user's device, and the consent confirmation message is displayed on the device.
[0131] Step 2: Submit consent data
[0132] When the device receives the user's consent, it sends it back to the server. Specifically, it uses an HTTP POST request to send the user's consent data. The input is the user's consent, and the output is the transmission of the consent data to the server. Specifically, the device recognizes the user's "Agree" operation and sends the consent data to the server.
[0133] Step 3: Collect and analyze user data
[0134] The server collects the user's past search history and purchase history and analyzes their areas of interest based on this data. The collected data includes search history, purchase history, etc., and analyzes the text data using a natural language processing algorithm. The input is the user's past search history and purchase history, and the output is the user's areas of interest and a topic list. Specifically, the server retrieves the user's history from the database and inputs it into the interest analysis algorithm.
[0135] Step 4: Obtaining location information
[0136] Devices (car navigation systems and smartphones) periodically obtain their current location information using a GPS module. The input is GPS location data, and the output is sending the location information to a server. Specifically, the device updates its location information every five minutes and sends that data to the server via an HTTP POST request.
[0137] Step 5: Send location information
[0138] The device sends the acquired location information to the server. This information is sent to the server in JSON format or similar. The acquired location information is used as input, and the location information is sent to the server as output. Specifically, the device sends coordinate data from the GPS to the server.
[0139] Step 6: Matching and selecting topics
[0140] The server compares the received location information with the user's interest list and selects relevant topics. This process uses a database search. The location information and interest list are input, and the selected topics are output and sent to the device. Specifically, the server searches the database using an SQL query to extract relevant topics.
[0141] Step 7: Submit topic data
[0142] The server sends the selected topic to the terminal using a data format such as JSON. The selected topic is used as input, and the topic data is sent to the terminal as output.
[0143] Step 8: Uttering topic data
[0144] The device uses speech synthesis technology to speak to the user based on the topic data received from the server. The received topic data is used as input, and the voice speech is used as output. Specifically, using the Google Text-to-Speech API, the device speaks, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[0145] Step 9: Analyze user responses
[0146] The user responds by voice, and the device analyzes the content. The user's voice response is the input, and the analysis results are sent to the server as the output. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text and send it to the server.
[0147] Step 10: Generate response data
[0148] The server generates the next topic and response based on the analysis results and sends them to the device. The input is the analysis result of the user's response, and the output is the generated response. Specifically, a generative AI model (e.g., GPT-4) is used to generate an appropriate response.
[0149] Step 11: Monitor car navigation voice instructions
[0150] The device constantly monitors the voice instructions from the car navigation system, and when these are heard, the device pauses the interaction. The input is the voice instructions from the car navigation system, and the output is pausing the interaction. Specifically, the device monitors the voice instructions through the car navigation API, and pauses the interaction when it says "Turn right at the next traffic light."
[0151] Step 12: Restart the conversation after receiving instructions from the car navigation system
[0152] When the car navigation system's voice instructions end, the device resumes the conversation with the user. The input is the end of the car navigation system's voice instructions, and the output is the resumption of the interaction. Specifically, the device confirms that the car navigation system's instructions have ended, and resumes the voice interaction based on the previous topic and the next topic.
[0153] (Application example 1)
[0154] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0155] During long-distance driving, drivers can become bored or drowsy, increasing the risk of accidents. Even in self-driving vehicles, passengers can lose interest and become bored. It is essential to prevent this and ensure a safe and comfortable drive.
[0156] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0157] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, and means for synchronizing the interaction with safety instructions from the automated driving system and pausing the interaction under certain conditions, thereby enabling drivers and passengers to maintain their interest even during long-distance driving, ensuring a comfortable and safe drive.
[0158] "User's past activities" is a general term for behavioral data such as a user's past search history, purchase history, and browsing history.
[0159] "Interests and concerns" refers to the tendency or degree of interest a user has in a particular theme or topic.
[0160] "Location information" refers to information indicating a user's current location and past travel routes obtained using GPS or other location measurement technologies.
[0161] "Topics" refer to information and content of conversations provided to users, and are selected based on the user's interests and location information.
[0162] "Voice interaction" refers to communication between a user and a system through voice, including the user asking questions and responding by voice.
[0163] A "car navigation device" is a device that displays map information for a vehicle and provides route guidance to a destination.
[0164] "Synchronization of voice utterances" refers to the timely coordination of voice guidance from the car navigation device and voice interaction with the user.
[0165] An "automated driving system" is a system that enables a vehicle to drive automatically, controlling the vehicle without the need for driver intervention.
[0166] "Safety instructions" refers to guidance and warning information provided by the automated driving system to ensure operational safety.
[0167] "Means for temporarily suspending interaction under certain conditions" refers to a control method for temporarily halting voice dialogue with the user when a safety instruction is issued by a car navigation device or an automated driving system.
[0168] This invention provides a system for preventing drowsiness during long-distance driving and allowing drivers to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities, provides appropriate topics, acquires location information and selects topics related to that information, and interacts with the user through voice using the selected topics. It also includes a function for controlling the synchronization of voice utterances with a car navigation system and synchronizing interactions with safety instructions from an automated driving system.
[0169] System configuration
[0170] The system is realized with the following components:
[0171] 1. Server: Collects and analyzes data such as your past search and purchase history to identify your interests.
[0172] 2. Terminal: Includes the dashboard, GPS module, and speech synthesis and recognition system installed in the vehicle.
[0173] 3. Car navigation device: Displays map information for the vehicle and provides guidance on the route to the destination.
[0174] 4. Autonomous driving system: A system in which a vehicle drives itself.
[0175] 5. Speech synthesis and recognition engine: An engine for realizing voice interaction with the user. Specifically, Google Cloud Text-to-Speech and Speech-to-Text APIs are used.
[0176] Program processing overview
[0177] Below is a description of how each component works together to process and calculate data:
[0178] Initial Setup
[0179] The server verifies the user's consent and collects the necessary data. The specific steps are as follows:
[0180] 1. The device prompts the user to consent to data collection.
[0181] 2. The user consents and the information is returned to the server.
[0182] 3. The server collects and analyzes the user's past activity data, identifies areas of interest, and generates a list of topics.
[0183] Obtaining location information
[0184] The device obtains its current location and sends that information to the server. The specific steps are as follows:
[0185] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[0186] 2. The acquired location information is sent to the server.
[0187] Topic selection
[0188] The server selects topics based on interests and location, as follows:
[0189] 1. The server matches the received location information with the user's interest list.
[0190] 2. Pick up topics related to the user's interests and location information and send them to the device.
[0191] Start a conversation
[0192] The device will then interact with the user through voice based on the selected topic. The specific steps are as follows:
[0193] 1. The device uses speech synthesis technology to provide topics to the user.
[0194] 2. The user responds verbally, and the device analyzes the speech.
[0195] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0196] Cooperation with autonomous driving systems
[0197] The terminal synchronizes the safety instructions and interactions of the automated driving system. The specific steps are as follows:
[0198] 1. The device constantly monitors the voice instructions of the autonomous driving system.
[0199] 2. When an autonomous driving safety instruction is issued, the system pauses interaction with the user.
[0200] 3. When the instruction is completed, the interaction resumes.
[0201] Specific examples
[0202] For example, if a user has been interested in travel in the past and is currently near Tokyo Tower, the server will select a topic about the "history of Tokyo Tower's architecture." The device will then say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" If the autonomous driving system issues the instruction "Turn right at the next traffic light," the infotainment system will pause and resume the conversation after the instruction is completed.
[0203] Prompt Sentence Examples
[0204] Here are some example prompts to input to a generative AI model:
[0205] "Design an infotainment system for autonomous vehicles that can provide topical information related to the driver's current location and enable voice interaction based on the driver's interests."
[0206] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0207] Step 1:
[0208] The server sends a prompt to the terminal asking the user for consent to data collection.
[0209] Input: The server receives the data needed to generate the consent prompt (e.g., the consent message).
[0210] Data calculation: The server generates a prompt message.
[0211] Output: The generated prompt message is sent to the terminal.
[0212] Step 2:
[0213] The terminal displays the consent prompt received from the server to the user.
[0214] Input: The terminal receives the prompt message.
[0215] Data processing: The terminal displays a prompt message on the user interface.
[0216] Output: User's choice to accept or decline.
[0217] Step 3:
[0218] The user presses the consent button and the information is sent back to the server.
[0219] Input: User's consent or denial of information.
[0220] Data calculation: The terminal analyzes the user's selection and generates consent information.
[0221] Output: The consent information is sent to the server.
[0222] Step 4:
[0223] The server collects data such as the user's past search and purchase history to identify areas of interest.
[0224] Input: The server receives the user's activity data along with consent information.
[0225] Data Calculation: The server collects the user's past activity data from a database and uses analysis algorithms to identify areas of interest.
[0226] Output: A list of areas of interest is generated.
[0227] Step 5:
[0228] The device obtains its current location and sends that information to the server.
[0229] Input: The device receives current location information from the GPS module.
[0230] Data processing: The acquired location information is converted into a format suitable for the server.
[0231] Output: The converted current location information is sent to the server.
[0232] Step 6:
[0233] The server compares the received location information with the interest list, selects appropriate topics, and sends them to the device.
[0234] Input: The server receives the current location and a list of interests.
[0235] Data calculation: The server matches the location information with the list of areas of interest and generates related topics.
[0236] Output: A list of relevant topics is generated and sent to the terminal.
[0237] Step 7:
[0238] The terminal uses speech synthesis technology to speak to the user based on the topic data received from the server.
[0239] Input: The terminal receives topic data.
[0240] Data processing: Using voice synthesis technology, the topic data is converted into a voice message.
[0241] Output: The audio message is spoken to the user.
[0242] Step 8:
[0243] The user responds with voice, which is then analyzed by the device.
[0244] Input: The user's spoken response.
[0245] Data Computing: A speech recognition engine converts the user's response into text and analyzes it.
[0246] Output: The analysis results are sent to the server.
[0247] Step 9:
[0248] The server generates the next topic and response based on the analysis results and sends them back to the terminal.
[0249] Input: The server receives the user's voice analysis results.
[0250] Data calculation: The server generates an appropriate response based on the analysis results.
[0251] Output: The generated response is sent to the terminal.
[0252] Step 10:
[0253] When the autonomous driving system issues safety instructions, the device temporarily stops interacting with the user.
[0254] Input: The terminal receives safety instructions from the automated driving system.
[0255] Data processing: Pause the interaction.
[0256] Output: Voice interaction with the user stops during the safety instruction period.
[0257] Step 11:
[0258] Once the safety instructions from the automated driving system are complete, the device resumes interaction with the user.
[0259] Input: Notification of end of safety instruction.
[0260] Data processing: Restarting a stopped interaction.
[0261] Output: Voice interaction with the user is resumed.
[0262] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0263] System Overview
[0264] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[0265] Program processing overview and specific examples
[0266] Initial Setup
[0267] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[0268] 1. The server sends a prompt to the user's device asking for consent to data collection.
[0269] 2. The device displays a consent confirmation screen to the user.
[0270] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0271] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[0272] Example: If a user has made many music-related searches and purchases in the past, tag their interests as music-related and add music topics to the priority list.
[0273] Obtaining location information
[0274] The device obtains location information using GPS and sends that information to the server.
[0275] 1. The device periodically receives GPS signals and obtains its current location information.
[0276] 2. The device sends the acquired location information to the server.
[0277] Example: If your current location is near a tourist attraction, send that information.
[0278] Topic selection
[0279] The server selects topics based on interests and location.
[0280] 1. The server matches the location information with the user's interest list.
[0281] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[0282] Example: If the location is near a historical place and the user is interested in history, select topics about the history of that place.
[0283] Emotion recognition by emotion engine
[0284] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[0285] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[0286] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[0287] Example: If the user is tired or stressed, the emotion engine will recognize that emotional state.
[0288] Adjusting the topic based on emotions
[0289] The server adjusts the topic based on the user's emotional state.
[0290] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[0291] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[0292] For example, if the user is tired, select relaxing music or topics that will help them relax.
[0293] Start a conversation
[0294] The device will interact via voice based on the selected topic.
[0295] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0296] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[0297] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[0298] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0299] Linking with car navigation systems
[0300] The terminal monitors and controls the voice instructions of the car navigation device.
[0301] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0302] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0303] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0304] For example: "Let me tell you about the history of this area."
[0305] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[0306] The processing flow will be explained below.
[0307] Step 1:
[0308] The server confirms the user's consent
[0309] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[0310] 2. The device displays a consent confirmation screen to the user.
[0311] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0312] Step 2:
[0313] The server collects and analyzes the user's past activity data.
[0314] 1. Once the server receives the user's consent, it collects usage history data (e.g., search history, purchase history) of the relevant online service.
[0315] 2. The server analyzes the collected data and identifies the user's areas of interest.
[0316] Example: If a user does a lot of movie-related searches and purchases, tag their interests as movie-related.
[0317] 3. The server generates a priority list of interests based on the analysis results.
[0318] Step 3:
[0319] The device acquires location information using GPS
[0320] 1. The device periodically receives GPS signals and obtains its current location information.
[0321] 2. The device sends the acquired location information to the server.
[0322] Example: If your current location is near a tourist attraction, send that information.
[0323] Step 4:
[0324] The server selects topics based on interests and location information
[0325] 1. The server matches the location information with the user's interest list.
[0326] 2. The server selects interesting topics related to the location.
[0327] Example: If the location is near a historical landmark and the user is interested in history, select topics related to the history of that place.
[0328] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[0329] Step 5:
[0330] The device recognizes the user's emotions
[0331] 1. The device uses an emotion engine to analyze the user's voice and facial expressions.
[0332] 2. The device assesses the user's emotional state (e.g., happy, sad, tired, etc.).
[0333] 3. The device sends the evaluation results to the server.
[0334] Step 6:
[0335] The server adjusts the topic to match the emotional state
[0336] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[0337] 2. The server selects topics according to the user's emotional state and provides conversation content appropriate to the user's emotions.
[0338] For example, if the user is tired, select a topic or music that has a relaxing effect.
[0339] Step 7:
[0340] The device interacts with the user through voice
[0341] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0342] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[0343] 2. When the user responds by voice, the device analyzes the voice.
[0344] 3. The device sends the analysis results to the server and requests the next conversation content.
[0345] Step 8:
[0346] The terminal monitors and controls the voice instructions of the car navigation device
[0347] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0348] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0349] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0350] For example: "Let me tell you about the history of this area."
[0351] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[0352] Example 2
[0353] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0354] It is difficult for drivers to enjoy a comfortable long-distance drive without feeling drowsy. It is also not easy to provide appropriate interactions based on the user's interests, current location, and emotional state. In particular, current technology is not sufficient to synchronize the voice instructions of the car navigation system with the conversation between the user and the car.
[0355] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0356] In this invention, the server includes means for analyzing the user's interests and concerns based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for engaging in voice interaction with the user using the selected topics, and means for acquiring the user's emotional data using an emotion engine and adjusting the topics according to the emotional state. This makes it possible to provide appropriate interactions according to the user's interests, location information, and emotional state, prevent drowsiness during long-distance driving, and support a comfortable drive.
[0357] "User" means a person who operates or uses a particular system or device.
[0358] "Past activity" refers to behavioral data such as a user's past search history and purchase history.
[0359] "Means for analyzing interests and concerns" refers to technology or devices for analyzing a user's past activity data to identify what the user is interested in.
[0360] "Means for providing topics" refers to technologies and devices for presenting appropriate topics to users based on the analyzed data.
[0361] "Location Information" refers to a user's current physical location data obtained through technologies such as GPS.
[0362] "Means for obtaining location information" refers to technologies and devices that use GPS or other location measurement technologies to determine a user's current location.
[0363] "Means for selecting a topic" refers to a technology or device for selecting a topic appropriate for a user based on the acquired location information.
[0364] "Means for voice interaction" refers to technology and devices for using voice synthesis technology to have a voice dialogue with a user.
[0365] "Car navigation device" refers to an electronic device that displays the current location of a vehicle and provides directions to a destination.
[0366] The "means for controlling synchronization of voice utterances" refers to a technique or device for appropriately adjusting the voice instructions of the car navigation device and the voice dialogue with the user.
[0367] An "emotion engine" refers to technology or a device that analyzes a user's voice and facial expressions to assess their emotional state.
[0368] "Means for acquiring emotional data" refers to technology or devices for acquiring the user's emotional state in real time using an emotion engine.
[0369] "Means for adjusting the topic" refers to technology or devices for changing or adapting the topic provided to the user based on the acquired emotional data.
[0370] System Overview
[0371] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[0372] Program processing overview and specific examples
[0373] Initial Setup
[0374] The server verifies the user's consent and collects the necessary data.
[0375] 1. The server sends a prompt to the user's device asking for consent to data collection.
[0376] 2. The device displays a consent confirmation screen to the user.
[0377] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0378] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[0379] Examples:
[0380] If a user has made many music-related searches or purchases in the past, we tag their interests as music-related and add music topics to the priority list.
[0381] Obtaining location information
[0382] The device obtains location information using GPS and sends that information to the server.
[0383] 1. The device periodically receives GPS signals and obtains its current location information.
[0384] 2. The device sends the acquired location information to the server.
[0385] Examples:
[0386] If the location information is near a tourist attraction, the information is sent.
[0387] Topic selection
[0388] The server selects topics based on interests and location.
[0389] 1. The server matches the location information with the user's interest list.
[0390] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[0391] Examples:
[0392] If the location information is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[0393] Emotion recognition by emotion engine
[0394] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[0395] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[0396] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[0397] Examples:
[0398] If the user is tired or stressed, the emotion engine will recognize that emotional state.
[0399] Adjusting the topic based on emotions
[0400] The server adjusts the topic based on the user's emotional state.
[0401] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[0402] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[0403] Examples:
[0404] If the user is tired, select relaxing music or topics that will help them relax.
[0405] Start a conversation
[0406] The device will interact via voice based on the selected topic.
[0407] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0408] Examples:
[0409] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[0410] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[0411] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0412] Linking with car navigation systems
[0413] The terminal monitors and controls the voice instructions of the car navigation device.
[0414] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0415] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0416] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0417] Examples:
[0418] "Let me tell you about the history of this area," he resumes.
[0419] Prompt Sentence Examples
[0420] Examples:
[0421] A user is on a long drive and the car navigation system tells them to "turn left at the next traffic light."
[0422] Example prompt for a generative AI model:
[0423] "You're interested in history and are currently driving near a historical landmark. After the navigation system finishes giving you current directions, you can start a conversation about the history of that landmark."
[0424] This system can provide appropriate topics according to the user's emotional state and interests while driving, supporting a safe and comfortable drive.
[0425] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0426] Step 1:
[0427] The server sends a user consent prompt to the terminal.
[0428] Specific behavior:
[0429] The server creates a prompt to obtain the user's consent to data collection and sends the prompt to the terminal.
[0430] Input: Consent prompt template information, User ID.
[0431] Output: The user consent prompt sent to the device.
[0432] Step 2:
[0433] The terminal displays a consent confirmation screen to the user.
[0434] Specific behavior:
[0435] The terminal generates a consent confirmation screen based on the received consent prompt and displays it to the user.
[0436] Input: The consent prompt received from the server.
[0437] Output: The consent screen shown to the user.
[0438] Step 3:
[0439] The user presses the consent button and the information is sent back to the server via the terminal.
[0440] Specific behavior:
[0441] When the user presses the accept button, their input is captured on the device.
[0442] The terminal returns the user's consent information to the server.
[0443] Input: User consent.
[0444] Output: User consent information sent to the server.
[0445] Step 4:
[0446] The server collects the user's past search and purchase history, analyzes their areas of interest, and generates a topic list.
[0447] Specific behavior:
[0448] The server collects the user's past search history and purchase history from a database.
[0449] The server analyzes the collected data to identify the user's areas of interest.
[0450] The server generates a topic list appropriate for the user based on the identified areas of interest.
[0451] Input: User's past search history, purchase history.
[0452] Output: Analyzed areas of interest, generated topic list.
[0453] Examples:
[0454] If a user has made many music-related searches or purchases in the past, the server will use that information to add music-related topics to the list.
[0455] Step 5:
[0456] The device receives a GPS signal and sends the location information to the server.
[0457] Specific behavior:
[0458] The device periodically receives GPS signals to obtain its current location.
[0459] The terminal transmits the acquired location information to the server.
[0460] Input: GPS signal.
[0461] Output: Current location information sent to the server.
[0462] Examples:
[0463] If the location information is near a tourist attraction, the information is sent to the server.
[0464] Step 6:
[0465] The server compares the location information with the user's interest list and selects topics.
[0466] Specific behavior:
[0467] The server matches the received location information with the user's interest list.
[0468] The server selects interesting topics related to the location.
[0469] The topics are formatted into a conversational format and sent to the terminal.
[0470] Input: Location, user interest list.
[0471] Output: Conversational topic data sent to the device.
[0472] Examples:
[0473] If the current location is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[0474] Step 7:
[0475] The terminal recognizes the user's emotions and transmits the emotional state to the server.
[0476] Specific behavior:
[0477] The device acquires emotional data using an emotion engine that analyzes the user's voice and facial expressions.
[0478] The device evaluates the acquired emotion data and transmits the information to the server.
[0479] Input: User's voice and facial expression data.
[0480] Output: Emotion data sent to the server.
[0481] Examples:
[0482] Recognizes the user's emotional state when they are tired or stressed.
[0483] Step 8:
[0484] The server adjusts the topic based on the emotional state and sends it to the terminal.
[0485] Specific behavior:
[0486] The server evaluates the user's current emotional state based on the data received from the emotion engine.
[0487] The server selects a topic according to the emotional state and sends it to the terminal.
[0488] Input: Sentiment data, topic list.
[0489] Output: The adjusted topic data sent to the device.
[0490] Examples:
[0491] If the user is tired, select relaxing music or topics that will help them relax.
[0492] Step 9:
[0493] The device speaks the topic and engages in interaction.
[0494] Specific behavior:
[0495] The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0496] When the user responds by voice, the voice is analyzed and the analysis results are sent to the server.
[0497] Input: Topic data from the server, user voice input.
[0498] Output: Audio utterance, analysis results sent to the server.
[0499] Examples:
[0500] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[0501] Step 10:
[0502] The device monitors the car navigation system's voice instructions and controls the interaction.
[0503] Specific behavior:
[0504] The terminal constantly monitors the voice instructions of the car navigation device.
[0505] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[0506] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[0507] Input: Car navigation voice instructions.
[0508] Output: Pause and resume conversation.
[0509] Examples:
[0510] When the car navigation system says, "Turn right at the next traffic light," it pauses the conversation with the user, and then resumes after the instruction is completed, saying, "Let me tell you about the history of this area."
[0511] (Application example 2)
[0512] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0513] Currently, there are not enough systems proposed to prevent drowsiness during long-distance driving and ensure a comfortable and safe drive. Furthermore, no systems have been found that combine a variety of elements, such as providing topics based on the driver's emotional state and interests, selecting topics based on location information, and synchronizing voice output with a car navigation system. The present invention aims to solve these problems and provide a system that prevents drowsiness during long-distance driving and enables drivers to enjoy a fun and comfortable drive.
[0514] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's interests based on their past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for recognizing the user's emotions and adjusting the topics according to that state, and means for controlling synchronization of voice utterances with a car navigation device. This makes it possible to provide topics according to the interests and concerns of drivers who are driving long distances, select appropriate topics based on location information, adjust the content of the conversation according to their emotional state, and synchronize voice utterances with the navigation device to ensure safe driving.
[0515] "Means for analyzing a user's interests based on their past activities and providing appropriate topics" refers to a device or program that collects data such as a user's past search history and purchase history, analyzes it to identify areas and topics that the user is likely to be interested in, and generates and provides topics based on that data.
[0516] "Means for acquiring location information and selecting topics related to that information" refers to a device or program that identifies the user's current location using a location information acquisition means such as GPS, and selects appropriate topics based on historical background, tourist attractions, local characteristics, etc. related to that location information.
[0517] "Means for recognizing a user's emotions and adjusting the topic according to that state" refers to a device or program that uses voice and facial expression analysis technology to analyze the content of a user's speech and facial expressions in real time, recognizes the user's emotional state (e.g., joy, surprise, fatigue, stress, etc.), and selects the topic that best suits the recognized emotion.
[0518] The "means for controlling synchronization of voice utterances with a car navigation device" refers to a control device or program that monitors voice instructions issued by a car navigation system in real time, automatically stops voice interaction with the user when a navigation instruction is given, and resumes conversation after the instruction is completed.
[0519] "Control means for automatically stopping interaction with the user when the car navigation device gives a voice instruction" refers to a device or program that has the function of detecting when the car navigation system is about to give the next voice instruction and temporarily stopping the conversation with the user during that time.
[0520] The present invention is a system that aims to prevent drivers from getting drowsy during long-distance driving and to allow them to enjoy a fun and comfortable drive. The system is configured using the following hardware and software.
[0521] System Overview
[0522] 1. Server Role
[0523] Analyze and recommend topics based on user's past activities
[0524] The server collects the user's past search history and purchase history, and by analyzing this data, identifies the areas the user is interested in. Based on this data, it generates an appropriate topic list.
[0525] Example: If users are doing a lot of music-related searches and purchases, prioritize music topics in your list.
[0526] Obtain location information and select related topics
[0527] The server receives GPS location information sent from the device and selects topics related to a specific location based on that information.
[0528] For example: If the user is near a tourist attraction, choose a topic related to the history and attractions of that place.
[0529] User Emotion Recognition and Topic Adjustment
[0530] The server receives the user's voice or facial expression analysis data sent from the terminal, evaluates the user's emotional state using an emotion engine, and adjusts the topic based on this evaluation to provide conversation content appropriate to the user's emotional state.
[0531] Example: If the emotion engine detects that the user is tired, it will provide relaxing music or topics.
[0532] 2. Role of the terminal
[0533] Acquisition and transmission of GPS location information
[0534] The device periodically receives GPS signals, acquires its current location information, and sends it to the server.
[0535] Emotion recognition by emotion engine
[0536] The device is equipped with a facial recognition camera and microphone to analyze the user's voice and facial expressions in real time, and the analysis results are sent to a server.
[0537] Conducting a voice interaction
[0538] The device uses speech synthesis technology to converse with the user based on the topic data received from the server. When the user responds by voice, the device analyzes the voice and sends the analysis results back to the server.
[0539] For example, you could say, "Hello. There is a historic building nearby. Do you know about this place?"
[0540] 3. Linkage with car navigation devices
[0541] The terminal works in conjunction with the car navigation system, automatically pausing the conversation with the user when a voice command is given, and resuming the conversation after the command is completed.
[0542] Prompt Sentence Examples
[0543] "There's a historic building nearby. Would you like to know about this place? And if you're interested in music, would you like to tell me about a new hit song?"
[0544] This system can select appropriate topics based on the user's past activity information, current location information, and emotional information, supporting a fun and comfortable drive. By providing interesting topics based on location and emotional information, the user can stay relaxed and alert while driving. This invention achieves more personalized interactions by using a generative AI model and an emotion recognition engine.
[0545] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0546] Step 1:
[0547] Consent verification and user data collection
[0548] Input: User ID
[0549] How it works: When a user launches the app, the server sends a consent prompt to the device.
[0550] When the user presses the consent button, the information is sent back to the server via the terminal.
[0551] Output: User consent information
[0552] Data processing: The server collects and analyzes users' past search history and purchase history.
[0553] Output: User's interests and a list of topics based on them
[0554] Step 2:
[0555] Acquiring and sending location information
[0556] Input: None
[0557] How it works: The device periodically receives GPS signals and obtains its current location.
[0558] Output: Location information
[0559] Data processing: The device sends the acquired location information to the server.
[0560] Output: Location data
[0561] Step 3:
[0562] Topic selection
[0563] Input: User's interest list, location
[0564] How it works: The server matches the location information with the user's interest list.
[0565] The server selects interesting topics related to the location and formats them into conversational format.
[0566] Output: topic list
[0567] Data processing: Selecting topics based on user interests and location information
[0568] Output: A formatted list of conversation topics
[0569] Step 4:
[0570] emotion recognition
[0571] Input: Voice data, facial expression data
[0572] Operation:
[0573] The device collects the user's voice and facial expressions using a microphone and a facial recognition camera.
[0574] The device uses an emotion engine to analyze emotions in real time.
[0575] Output: Emotion data
[0576] Data processing: Emotion evaluation as a result of analyzing voice data and facial expression data
[0577] Output: User's emotional state
[0578] Step 5:
[0579] Adjusting the topic based on emotions
[0580] Input: Topic list, sentiment data
[0581] Operation:
[0582] The server evaluates the user's current emotional state based on the emotion engine's evaluation.
[0583] The server adjusts the topic list according to the emotional state.
[0584] Output: Adjusted topic list
[0585] Data processing: Selecting topics optimized for emotions
[0586] Output: Adjusted topic list
[0587] Step 6:
[0588] Voice Interactions
[0589] Input: adjusted topic list
[0590] Operation:
[0591] The device uses speech synthesis technology to speak aloud based on topic data.
[0592] When the user responds by voice, the terminal analyzes the voice and sends the analysis results to the server.
[0593] Output: Audio data, analysis results
[0594] Data processing: Selecting the next topic based on the results of voice data analysis
[0595] Output: Next topic or response
[0596] Step 7:
[0597] Linking with car navigation systems
[0598] Input: Car navigation voice instructions
[0599] Operation:
[0600] The terminal constantly monitors the voice instructions of the car navigation device.
[0601] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[0602] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[0603] Output: Trigger for the end of voice command
[0604] Data processing: Conversation control based on voice instructions for car navigation
[0605] Output: The resumed conversation
[0606] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0607] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0608] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0609] [Second embodiment]
[0610] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0611] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0612] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0613] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0614] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0615] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0616] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0617] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0618] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0619] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0620] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0621] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0622] System Overview
[0623] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to allow them to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities and provides appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, thereby ensuring driving safety.
[0624] Program processing overview and specific examples
[0625] Initial Setup
[0626] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[0627] 1. The server sends a prompt to the device asking the user for consent to data collection.
[0628] 2. When the user presses the consent button, the information is sent back to the server.
[0629] 3. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[0630] Example: If a user has made a lot of travel-related searches and purchases in the past, use that information to prioritize travel topics.
[0631] Obtaining location information
[0632] The device obtains its current location and sends that information to the server.
[0633] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[0634] 2. The device sends the acquired location information to the server.
[0635] Example: If the user is near a tourist attraction in Tokyo, the current location information is sent to the server that the user is near Tokyo Tower.
[0636] Topic selection
[0637] The server selects topics based on interests and location.
[0638] 1. The server matches the received location information with the user's interest list.
[0639] 2. The server picks up topics related to the user's interests and location information and sends them to the device.
[0640] Example: If the current location is near Tokyo Tower and the user is interested in architecture, select the topic "The history of the architecture of Tokyo Tower."
[0641] Start a conversation
[0642] The device will interact via voice based on the selected topic.
[0643] 1. The device uses speech synthesis technology to speak to the user based on the topic data received from the server.
[0644] Example: The device says, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[0645] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[0646] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0647] Linking with car navigation systems
[0648] The device synchronizes with the car navigation system's speech.
[0649] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0650] 2. When the car navigation system issues voice instructions, the device automatically interrupts interaction with the user.
[0651] Example: The device pauses voice conversation while the car navigation system says, "Turn right at the next traffic light."
[0652] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0653] The system takes into account the user's individual interests and current geographical environment to provide a safe and comfortable drive while preventing drowsiness while driving.
[0654] The processing flow will be explained below.
[0655] Step 1:
[0656] The server confirms the user's consent
[0657] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[0658] 2. The device displays a consent confirmation screen to the user.
[0659] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0660] Step 2:
[0661] The server collects and analyzes the user's past activity data.
[0662] 1. Once the server receives the user's consent, it will collect usage history data for LINE Yahoo! services (e.g., search history, purchase history).
[0663] 2. The server analyzes the collected data and identifies the user's areas of interest.
[0664] Example: If a user searches for and purchases a lot of automotive-related items, tag their interests as automotive-related.
[0665] 3. The server generates a priority list of interests based on the analysis results.
[0666] Step 3:
[0667] The device acquires location information using GPS
[0668] 1. The device periodically receives GPS signals and obtains its current location information.
[0669] 2. The device sends the acquired location information to the server.
[0670] Example: If your current location is a tourist spot in a city, send that information.
[0671] Step 4:
[0672] The server selects topics based on interests and location information
[0673] 1. The server matches the location information with the user's interest list.
[0674] 2. The server selects interesting topics related to the location.
[0675] Example: If the location is near Tokyo Tower and the user is interested in architecture, select topics related to the history of architecture.
[0676] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[0677] Step 5:
[0678] The device interacts with the user through voice
[0679] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0680] Example: Say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[0681] 2. When the user responds by voice, the device analyzes the voice.
[0682] 3. The device sends the analysis results to the server and requests the next conversation content.
[0683] Step 6:
[0684] The terminal monitors and controls the voice instructions of the car navigation device
[0685] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0686] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0687] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0688] For example: "Let's talk about the construction of Tokyo Tower."
[0689] This allows users to avoid drowsiness while driving and enjoy conversations about interesting topics. In addition, safe driving is ensured by linking with car navigation systems.
[0690] Example 1
[0691] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0692] Conventional long-distance driving assistance systems have had problems with drivers feeling drowsy, reducing driving safety. They also lacked entertainment and information provision while driving, making it impossible to provide topics that matched the driver's interests. Furthermore, if voice utterances were not properly synchronized with the car navigation system, the driver's attention could be distracted, potentially leading to an accident.
[0693] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0694] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, means for collecting data with the user's consent, and means for analyzing the user's responses and generating the next topic or response based on the analysis results. This allows for a safe and comfortable drive while preventing drowsiness while taking into account the user's individual interests and current geographical environment.
[0695] "Means of analyzing a user's interests and concerns based on their past activities and providing appropriate topics" refers to a system that collects past behavioral data such as a user's search history and purchase history, analyzes this data to identify the user's interests and concerns, and automatically selects and provides related topics based on that.
[0696] "Means for obtaining location information and selecting topics related to that information" refers to a mechanism in which the user's device obtains current geographical location information using GPS or other means, and selects interesting related topics based on this location information.
[0697] "Means for interacting with the user via voice using a selected topic" refers to a system that uses voice synthesis technology to provide information to the user via voice based on a selected topic, and allows the user to respond via voice, thereby achieving two-way communication.
[0698] The "means for controlling synchronization of voice utterances with the car navigation device" is a mechanism for starting or interrupting voice interaction with the user at appropriate timing in synchronization with voice instructions from the car navigation device.
[0699] "Means for collecting data with user consent" refers to a mechanism by which the system sends a prompt to the user requesting consent for data collection, obtains consent from the user, and collects the necessary data based on that consent.
[0700] "Means for analyzing the user's response and generating the next topic or response based on the analysis results" refers to a mechanism that analyzes the content of the user's voice response and generates the next topic or an appropriate response based on the analysis results.
[0701] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user through voice and further controls the synchronization of voice utterances with the car navigation system to ensure driving safety.
[0702] Initial Setup
[0703] The server confirms the user's consent and collects the necessary data. First, the server sends a prompt to the device to obtain the user's consent. This prompt includes a message such as "Please consent to data collection." When the user presses the consent button, the information is returned to the server, which stores it in a database. The server then collects the user's past search history, purchase history, etc., and uses this data to analyze areas of interest and generate a topic list. A natural language processing algorithm is preferably used in this process. Specific techniques that can be used include data analysis tools such as Python and R.
[0704] Obtaining location information
[0705] The device obtains its current location and sends that information to the server. The device (car navigation system or smartphone) periodically obtains its current location information using a GPS module. For example, it can be set to update its location information every 5 minutes. The obtained location information is sent to the server via an HTTP POST request, and the server stores this information in a database.
[0706] Topic selection
[0707] The server compares the received location information with the user's interest list and selects a topic based on that interest. The server uses database searches to compare the location information with the user's areas of interest. The selected topic is sent to the device in JSON format or similar. For example, if the current location is near Tokyo Tower and the user is interested in architecture, the topic "The architectural history of Tokyo Tower" will be selected.
[0708] Start a conversation
[0709] The device uses speech synthesis technology to speak to the user based on the selected topic. Specifically, it uses the Google Text-to-Speech API to say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" When the user responds verbally, the device analyzes this speech and converts it into text data using the Google Speech-to-Text API or similar. The analyzed data is sent to the server, which generates the next topic and response based on the analysis results. At this stage, a generative AI model (such as GPT-4) is used.
[0710] Linking with car navigation systems
[0711] The device constantly monitors voice instructions from the car navigation system. This is achieved through API integration between the car navigation system and the device. When the car navigation system issues a voice instruction, the device automatically suspends interaction with the user. For example, when the car navigation system says, "Turn right at the next traffic light," the voice conversation is paused. When the car navigation system finishes issuing the instruction, the device resumes the conversation with the user.
[0712] Specific examples and examples of prompts for generative AI models
[0713] As a concrete example, if a user asks, "How tall is Tokyo Tower?", the server will generate a response based on the analysis results and reply, "Tokyo Tower is approximately 333 meters tall."
[0714] An example of a prompt for a generative AI model is:
[0715] "Based on your past search history, what are some travel topics that have recently interested you?"
[0716] This system takes into account the user's individual interests and geographical environment, making it possible to provide a safe and comfortable drive while preventing drowsiness while driving.
[0717] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0718] Step 1: Confirm user agreement
[0719] The server sends a prompt to the device to obtain the user's consent. This uses an HTTP request and includes the message "Please consent to data collection." The input is a consent confirmation message, and the output is a prompt display. Specifically, the server sends an HTTP request to the user's device, and the consent confirmation message is displayed on the device.
[0720] Step 2: Submit consent data
[0721] When the device receives the user's consent, it sends it back to the server. Specifically, it uses an HTTP POST request to send the user's consent data. The input is the user's consent, and the output is the transmission of the consent data to the server. Specifically, the device recognizes the user's "Agree" operation and sends the consent data to the server.
[0722] Step 3: Collect and analyze user data
[0723] The server collects the user's past search history and purchase history and analyzes their areas of interest based on this data. The collected data includes search history, purchase history, etc., and analyzes the text data using a natural language processing algorithm. The input is the user's past search history and purchase history, and the output is the user's areas of interest and a topic list. Specifically, the server retrieves the user's history from the database and inputs it into the interest analysis algorithm.
[0724] Step 4: Obtaining location information
[0725] Devices (car navigation systems and smartphones) periodically obtain their current location information using a GPS module. The input is GPS location data, and the output is sending the location information to a server. Specifically, the device updates its location information every five minutes and sends that data to the server via an HTTP POST request.
[0726] Step 5: Send location information
[0727] The device sends the acquired location information to the server. This information is sent to the server in JSON format or similar. The acquired location information is used as input, and the location information is sent to the server as output. Specifically, the device sends coordinate data from the GPS to the server.
[0728] Step 6: Matching and selecting topics
[0729] The server compares the received location information with the user's interest list and selects relevant topics. This process uses a database search. The location information and interest list are input, and the selected topics are output and sent to the device. Specifically, the server searches the database using an SQL query to extract relevant topics.
[0730] Step 7: Submit topic data
[0731] The server sends the selected topic to the terminal using a data format such as JSON. The selected topic is used as input, and the topic data is sent to the terminal as output.
[0732] Step 8: Uttering topic data
[0733] The device uses speech synthesis technology to speak to the user based on the topic data received from the server. The received topic data is used as input, and the voice speech is used as output. Specifically, using the Google Text-to-Speech API, the device speaks, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[0734] Step 9: Analyze user responses
[0735] The user responds by voice, and the device analyzes the content. The user's voice response is the input, and the analysis results are sent to the server as the output. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text and send it to the server.
[0736] Step 10: Generate response data
[0737] The server generates the next topic and response based on the analysis results and sends them to the device. The input is the analysis result of the user's response, and the output is the generated response. Specifically, a generative AI model (e.g., GPT-4) is used to generate an appropriate response.
[0738] Step 11: Monitor car navigation voice instructions
[0739] The device constantly monitors the voice instructions from the car navigation system, and when these are heard, the device pauses the interaction. The input is the voice instructions from the car navigation system, and the output is pausing the interaction. Specifically, the device monitors the voice instructions through the car navigation API, and pauses the interaction when it says "Turn right at the next traffic light."
[0740] Step 12: Restart the conversation after receiving instructions from the car navigation system
[0741] When the car navigation system's voice instructions end, the device resumes the conversation with the user. The input is the end of the car navigation system's voice instructions, and the output is the resumption of the interaction. Specifically, the device confirms that the car navigation system's instructions have ended, and resumes the voice interaction based on the previous topic and the next topic.
[0742] (Application example 1)
[0743] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0744] During long-distance driving, drivers can become bored or drowsy, increasing the risk of accidents. Even in self-driving vehicles, passengers can lose interest and become bored. It is essential to prevent this and ensure a safe and comfortable drive.
[0745] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0746] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, and means for synchronizing the interaction with safety instructions from the automated driving system and pausing the interaction under certain conditions, thereby enabling drivers and passengers to maintain their interest even during long-distance driving, ensuring a comfortable and safe drive.
[0747] "User's past activities" is a general term for behavioral data such as a user's past search history, purchase history, and browsing history.
[0748] "Interests and concerns" refers to the tendency or degree of interest a user has in a particular theme or topic.
[0749] "Location information" refers to information indicating a user's current location and past travel routes obtained using GPS or other location measurement technologies.
[0750] "Topics" refer to information and content of conversations provided to users, and are selected based on the user's interests and location information.
[0751] "Voice interaction" refers to communication between a user and a system through voice, including the user asking questions and responding by voice.
[0752] A "car navigation device" is a device that displays map information for a vehicle and provides route guidance to a destination.
[0753] "Synchronization of voice utterances" refers to the timely coordination of voice guidance from the car navigation device and voice interaction with the user.
[0754] An "automated driving system" is a system that enables a vehicle to drive automatically, controlling the vehicle without the need for driver intervention.
[0755] "Safety instructions" refers to guidance and warning information provided by the automated driving system to ensure operational safety.
[0756] "Means for temporarily suspending interaction under certain conditions" refers to a control method for temporarily halting voice dialogue with the user when a safety instruction is issued by a car navigation device or an automated driving system.
[0757] This invention provides a system for preventing drowsiness during long-distance driving and allowing drivers to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities, provides appropriate topics, acquires location information and selects topics related to that information, and interacts with the user through voice using the selected topics. It also includes a function for controlling the synchronization of voice utterances with a car navigation system and synchronizing interactions with safety instructions from an automated driving system.
[0758] System configuration
[0759] The system is realized with the following components:
[0760] 1. Server: Collects and analyzes data such as your past search and purchase history to identify your interests.
[0761] 2. Terminal: Includes the dashboard, GPS module, and speech synthesis and recognition system installed in the vehicle.
[0762] 3. Car navigation device: Displays map information for the vehicle and provides guidance on the route to the destination.
[0763] 4. Autonomous driving system: A system in which a vehicle drives itself.
[0764] 5. Speech synthesis and recognition engine: An engine for realizing voice interaction with the user. Specifically, Google Cloud Text-to-Speech and Speech-to-Text APIs are used.
[0765] Program processing overview
[0766] Below is a description of how each component works together to process and calculate data:
[0767] Initial Setup
[0768] The server verifies the user's consent and collects the necessary data. The specific steps are as follows:
[0769] 1. The device prompts the user to consent to data collection.
[0770] 2. The user consents and the information is returned to the server.
[0771] 3. The server collects and analyzes the user's past activity data, identifies areas of interest, and generates a list of topics.
[0772] Obtaining location information
[0773] The device obtains its current location and sends that information to the server. The specific steps are as follows:
[0774] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[0775] 2. The acquired location information is sent to the server.
[0776] Topic selection
[0777] The server selects topics based on interests and location, as follows:
[0778] 1. The server matches the received location information with the user's interest list.
[0779] 2. Pick up topics related to the user's interests and location information and send them to the device.
[0780] Start a conversation
[0781] The device will then interact with the user through voice based on the selected topic. The specific steps are as follows:
[0782] 1. The device uses speech synthesis technology to provide topics to the user.
[0783] 2. The user responds verbally, and the device analyzes the speech.
[0784] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0785] Cooperation with autonomous driving systems
[0786] The terminal synchronizes the safety instructions and interactions of the automated driving system. The specific steps are as follows:
[0787] 1. The device constantly monitors the voice instructions of the autonomous driving system.
[0788] 2. When an autonomous driving safety instruction is issued, the system pauses interaction with the user.
[0789] 3. When the instruction is completed, the interaction resumes.
[0790] Specific examples
[0791] For example, if a user has been interested in travel in the past and is currently near Tokyo Tower, the server will select a topic about the "history of Tokyo Tower's architecture." The device will then say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" If the autonomous driving system issues the instruction "Turn right at the next traffic light," the infotainment system will pause and resume the conversation after the instruction is completed.
[0792] Prompt Sentence Examples
[0793] Here are some example prompts to input to a generative AI model:
[0794] "Design an infotainment system for autonomous vehicles that can provide topical information related to the driver's current location and enable voice interaction based on the driver's interests."
[0795] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0796] Step 1:
[0797] The server sends a prompt to the terminal asking the user for consent to data collection.
[0798] Input: The server receives the data needed to generate the consent prompt (e.g., the consent message).
[0799] Data calculation: The server generates a prompt message.
[0800] Output: The generated prompt message is sent to the terminal.
[0801] Step 2:
[0802] The terminal displays the consent prompt received from the server to the user.
[0803] Input: The terminal receives the prompt message.
[0804] Data processing: The terminal displays a prompt message on the user interface.
[0805] Output: User's choice to accept or decline.
[0806] Step 3:
[0807] The user presses the consent button and the information is sent back to the server.
[0808] Input: User's consent or denial of information.
[0809] Data calculation: The terminal analyzes the user's selection and generates consent information.
[0810] Output: The consent information is sent to the server.
[0811] Step 4:
[0812] The server collects data such as the user's past search and purchase history to identify areas of interest.
[0813] Input: The server receives the user's activity data along with consent information.
[0814] Data Calculation: The server collects the user's past activity data from a database and uses analysis algorithms to identify areas of interest.
[0815] Output: A list of areas of interest is generated.
[0816] Step 5:
[0817] The device obtains its current location and sends that information to the server.
[0818] Input: The device receives current location information from the GPS module.
[0819] Data processing: The acquired location information is converted into a format suitable for the server.
[0820] Output: The converted current location information is sent to the server.
[0821] Step 6:
[0822] The server compares the received location information with the interest list, selects appropriate topics, and sends them to the device.
[0823] Input: The server receives the current location and a list of interests.
[0824] Data calculation: The server matches the location information with the list of areas of interest and generates related topics.
[0825] Output: A list of relevant topics is generated and sent to the terminal.
[0826] Step 7:
[0827] The terminal uses speech synthesis technology to speak to the user based on the topic data received from the server.
[0828] Input: The terminal receives topic data.
[0829] Data processing: Using voice synthesis technology, the topic data is converted into a voice message.
[0830] Output: The audio message is spoken to the user.
[0831] Step 8:
[0832] The user responds with voice, which is then analyzed by the device.
[0833] Input: The user's spoken response.
[0834] Data Computing: A speech recognition engine converts the user's response into text and analyzes it.
[0835] Output: The analysis results are sent to the server.
[0836] Step 9:
[0837] The server generates the next topic and response based on the analysis results and sends them back to the terminal.
[0838] Input: The server receives the user's voice analysis results.
[0839] Data calculation: The server generates an appropriate response based on the analysis results.
[0840] Output: The generated response is sent to the terminal.
[0841] Step 10:
[0842] When the autonomous driving system issues safety instructions, the device temporarily stops interacting with the user.
[0843] Input: The terminal receives safety instructions from the automated driving system.
[0844] Data processing: Pause the interaction.
[0845] Output: Voice interaction with the user stops during the safety instruction period.
[0846] Step 11:
[0847] Once the safety instructions from the automated driving system are complete, the device resumes interaction with the user.
[0848] Input: Notification of end of safety instruction.
[0849] Data processing: Restarting a stopped interaction.
[0850] Output: Voice interaction with the user is resumed.
[0851] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0852] System Overview
[0853] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[0854] Program processing overview and specific examples
[0855] Initial Setup
[0856] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[0857] 1. The server sends a prompt to the user's device asking for consent to data collection.
[0858] 2. The device displays a consent confirmation screen to the user.
[0859] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0860] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[0861] Example: If a user has made many music-related searches and purchases in the past, tag their interests as music-related and add music topics to the priority list.
[0862] Obtaining location information
[0863] The device obtains location information using GPS and sends that information to the server.
[0864] 1. The device periodically receives GPS signals and obtains its current location information.
[0865] 2. The device sends the acquired location information to the server.
[0866] Example: If your current location is near a tourist attraction, send that information.
[0867] Topic selection
[0868] The server selects topics based on interests and location.
[0869] 1. The server matches the location information with the user's interest list.
[0870] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[0871] Example: If the location is near a historical place and the user is interested in history, select topics about the history of that place.
[0872] Emotion recognition by emotion engine
[0873] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[0874] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[0875] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[0876] Example: If the user is tired or stressed, the emotion engine will recognize that emotional state.
[0877] Adjusting the topic based on emotions
[0878] The server adjusts the topic based on the user's emotional state.
[0879] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[0880] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[0881] For example, if the user is tired, select relaxing music or topics that will help them relax.
[0882] Start a conversation
[0883] The device will interact via voice based on the selected topic.
[0884] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0885] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[0886] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[0887] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[0888] Linking with car navigation systems
[0889] The terminal monitors and controls the voice instructions of the car navigation device.
[0890] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0891] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0892] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0893] For example: "Let me tell you about the history of this area."
[0894] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[0895] The processing flow will be explained below.
[0896] Step 1:
[0897] The server confirms the user's consent
[0898] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[0899] 2. The device displays a consent confirmation screen to the user.
[0900] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0901] Step 2:
[0902] The server collects and analyzes the user's past activity data.
[0903] 1. Once the server receives the user's consent, it collects usage history data (e.g., search history, purchase history) of the relevant online service.
[0904] 2. The server analyzes the collected data and identifies the user's areas of interest.
[0905] Example: If a user does a lot of movie-related searches and purchases, tag their interests as movie-related.
[0906] 3. The server generates a priority list of interests based on the analysis results.
[0907] Step 3:
[0908] The device acquires location information using GPS
[0909] 1. The device periodically receives GPS signals and obtains its current location information.
[0910] 2. The device sends the acquired location information to the server.
[0911] Example: If your current location is near a tourist attraction, send that information.
[0912] Step 4:
[0913] The server selects topics based on interests and location information
[0914] 1. The server matches the location information with the user's interest list.
[0915] 2. The server selects interesting topics related to the location.
[0916] Example: If the location is near a historical landmark and the user is interested in history, select topics related to the history of that place.
[0917] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[0918] Step 5:
[0919] The device recognizes the user's emotions
[0920] 1. The device uses an emotion engine to analyze the user's voice and facial expressions.
[0921] 2. The device assesses the user's emotional state (e.g., happy, sad, tired, etc.).
[0922] 3. The device sends the evaluation results to the server.
[0923] Step 6:
[0924] The server adjusts the topic to match the emotional state
[0925] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[0926] 2. The server selects topics according to the user's emotional state and provides conversation content appropriate to the user's emotions.
[0927] For example, if the user is tired, select a topic or music that has a relaxing effect.
[0928] Step 7:
[0929] The device interacts with the user through voice
[0930] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0931] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[0932] 2. When the user responds by voice, the device analyzes the voice.
[0933] 3. The device sends the analysis results to the server and requests the next conversation content.
[0934] Step 8:
[0935] The terminal monitors and controls the voice instructions of the car navigation device
[0936] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[0937] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[0938] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[0939] For example: "Let me tell you about the history of this area."
[0940] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[0941] Example 2
[0942] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0943] It is difficult for drivers to enjoy a comfortable long-distance drive without feeling drowsy. It is also not easy to provide appropriate interactions based on the user's interests, current location, and emotional state. In particular, current technology is not sufficient to synchronize the voice instructions of the car navigation system with the conversation between the user and the car.
[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0945] In this invention, the server includes means for analyzing the user's interests and concerns based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for engaging in voice interaction with the user using the selected topics, and means for acquiring the user's emotional data using an emotion engine and adjusting the topics according to the emotional state. This makes it possible to provide appropriate interactions according to the user's interests, location information, and emotional state, prevent drowsiness during long-distance driving, and support a comfortable drive.
[0946] "User" means a person who operates or uses a particular system or device.
[0947] "Past activity" refers to behavioral data such as a user's past search history and purchase history.
[0948] "Means for analyzing interests and concerns" refers to technology or devices for analyzing a user's past activity data to identify what the user is interested in.
[0949] "Means for providing topics" refers to technologies and devices for presenting appropriate topics to users based on the analyzed data.
[0950] "Location Information" refers to a user's current physical location data obtained through technologies such as GPS.
[0951] "Means for obtaining location information" refers to technologies and devices that use GPS or other location measurement technologies to determine a user's current location.
[0952] "Means for selecting a topic" refers to a technology or device for selecting a topic appropriate for a user based on the acquired location information.
[0953] "Means for voice interaction" refers to technology and devices for using voice synthesis technology to have a voice dialogue with a user.
[0954] "Car navigation device" refers to an electronic device that displays the current location of a vehicle and provides directions to a destination.
[0955] The "means for controlling synchronization of voice utterances" refers to a technique or device for appropriately adjusting the voice instructions of the car navigation device and the voice dialogue with the user.
[0956] An "emotion engine" refers to technology or a device that analyzes a user's voice and facial expressions to assess their emotional state.
[0957] "Means for acquiring emotional data" refers to technology or devices for acquiring the user's emotional state in real time using an emotion engine.
[0958] "Means for adjusting the topic" refers to technology or devices for changing or adapting the topic provided to the user based on the acquired emotional data.
[0959] System Overview
[0960] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[0961] Program processing overview and specific examples
[0962] Initial Setup
[0963] The server verifies the user's consent and collects the necessary data.
[0964] 1. The server sends a prompt to the user's device asking for consent to data collection.
[0965] 2. The device displays a consent confirmation screen to the user.
[0966] 3. When the user presses the consent button, the information is sent back to the server via the device.
[0967] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[0968] Examples:
[0969] If a user has made many music-related searches or purchases in the past, we tag their interests as music-related and add music topics to the priority list.
[0970] Obtaining location information
[0971] The device obtains location information using GPS and sends that information to the server.
[0972] 1. The device periodically receives GPS signals and obtains its current location information.
[0973] 2. The device sends the acquired location information to the server.
[0974] Examples:
[0975] If the location information is near a tourist attraction, the information is sent.
[0976] Topic selection
[0977] The server selects topics based on interests and location.
[0978] 1. The server matches the location information with the user's interest list.
[0979] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[0980] Examples:
[0981] If the location information is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[0982] Emotion recognition by emotion engine
[0983] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[0984] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[0985] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[0986] Examples:
[0987] If the user is tired or stressed, the emotion engine will recognize that emotional state.
[0988] Adjusting the topic based on emotions
[0989] The server adjusts the topic based on the user's emotional state.
[0990] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[0991] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[0992] Examples:
[0993] If the user is tired, select relaxing music or topics that will help them relax.
[0994] Start a conversation
[0995] The device will interact via voice based on the selected topic.
[0996] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[0997] Examples:
[0998] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[0999] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[1000] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1001] Linking with car navigation systems
[1002] The terminal monitors and controls the voice instructions of the car navigation device.
[1003] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1004] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[1005] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1006] Examples:
[1007] "Let me tell you about the history of this area," he resumes.
[1008] Prompt Sentence Examples
[1009] Examples:
[1010] A user is on a long drive and the car navigation system tells them to "turn left at the next traffic light."
[1011] Example prompt for a generative AI model:
[1012] "You're interested in history and are currently driving near a historical landmark. After the navigation system finishes giving you current directions, you can start a conversation about the history of that landmark."
[1013] This system can provide appropriate topics according to the user's emotional state and interests while driving, supporting a safe and comfortable drive.
[1014] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1015] Step 1:
[1016] The server sends a user consent prompt to the terminal.
[1017] Specific behavior:
[1018] The server creates a prompt to obtain the user's consent to data collection and sends the prompt to the terminal.
[1019] Input: Consent prompt template information, User ID.
[1020] Output: The user consent prompt sent to the device.
[1021] Step 2:
[1022] The terminal displays a consent confirmation screen to the user.
[1023] Specific behavior:
[1024] The terminal generates a consent confirmation screen based on the received consent prompt and displays it to the user.
[1025] Input: The consent prompt received from the server.
[1026] Output: The consent screen shown to the user.
[1027] Step 3:
[1028] The user presses the consent button and the information is sent back to the server via the terminal.
[1029] Specific behavior:
[1030] When the user presses the accept button, their input is captured on the device.
[1031] The terminal returns the user's consent information to the server.
[1032] Input: User consent.
[1033] Output: User consent information sent to the server.
[1034] Step 4:
[1035] The server collects the user's past search and purchase history, analyzes their areas of interest, and generates a topic list.
[1036] Specific behavior:
[1037] The server collects the user's past search history and purchase history from a database.
[1038] The server analyzes the collected data to identify the user's areas of interest.
[1039] The server generates a topic list appropriate for the user based on the identified areas of interest.
[1040] Input: User's past search history, purchase history.
[1041] Output: Analyzed areas of interest, generated topic list.
[1042] Examples:
[1043] If a user has made many music-related searches or purchases in the past, the server will use that information to add music-related topics to the list.
[1044] Step 5:
[1045] The device receives a GPS signal and sends the location information to the server.
[1046] Specific behavior:
[1047] The device periodically receives GPS signals to obtain its current location.
[1048] The terminal transmits the acquired location information to the server.
[1049] Input: GPS signal.
[1050] Output: Current location information sent to the server.
[1051] Examples:
[1052] If the location information is near a tourist attraction, the information is sent to the server.
[1053] Step 6:
[1054] The server compares the location information with the user's interest list and selects topics.
[1055] Specific behavior:
[1056] The server matches the received location information with the user's interest list.
[1057] The server selects interesting topics related to the location.
[1058] The topics are formatted into a conversational format and sent to the terminal.
[1059] Input: Location, user interest list.
[1060] Output: Conversational topic data sent to the device.
[1061] Examples:
[1062] If the current location is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[1063] Step 7:
[1064] The terminal recognizes the user's emotions and transmits the emotional state to the server.
[1065] Specific behavior:
[1066] The device acquires emotional data using an emotion engine that analyzes the user's voice and facial expressions.
[1067] The device evaluates the acquired emotion data and transmits the information to the server.
[1068] Input: User's voice and facial expression data.
[1069] Output: Emotion data sent to the server.
[1070] Examples:
[1071] Recognizes the user's emotional state when they are tired or stressed.
[1072] Step 8:
[1073] The server adjusts the topic based on the emotional state and sends it to the terminal.
[1074] Specific behavior:
[1075] The server evaluates the user's current emotional state based on the data received from the emotion engine.
[1076] The server selects a topic according to the emotional state and sends it to the terminal.
[1077] Input: Sentiment data, topic list.
[1078] Output: The adjusted topic data sent to the device.
[1079] Examples:
[1080] If the user is tired, select relaxing music or topics that will help them relax.
[1081] Step 9:
[1082] The device speaks the topic and engages in interaction.
[1083] Specific behavior:
[1084] The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1085] When the user responds by voice, the voice is analyzed and the analysis results are sent to the server.
[1086] Input: Topic data from the server, user voice input.
[1087] Output: Audio utterance, analysis results sent to the server.
[1088] Examples:
[1089] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[1090] Step 10:
[1091] The device monitors the car navigation system's voice instructions and controls the interaction.
[1092] Specific behavior:
[1093] The terminal constantly monitors the voice instructions of the car navigation device.
[1094] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[1095] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[1096] Input: Car navigation voice instructions.
[1097] Output: Pause and resume conversation.
[1098] Examples:
[1099] When the car navigation system says, "Turn right at the next traffic light," it pauses the conversation with the user, and then resumes after the instruction is completed, saying, "Let me tell you about the history of this area."
[1100] (Application example 2)
[1101] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1102] Currently, there are not enough systems proposed to prevent drowsiness during long-distance driving and ensure a comfortable and safe drive. Furthermore, no systems have been found that combine a variety of elements, such as providing topics based on the driver's emotional state and interests, selecting topics based on location information, and synchronizing voice output with a car navigation system. The present invention aims to solve these problems and provide a system that prevents drowsiness during long-distance driving and enables drivers to enjoy a fun and comfortable drive.
[1103] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's interests based on their past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for recognizing the user's emotions and adjusting the topics according to that state, and means for controlling synchronization of voice utterances with a car navigation device. This makes it possible to provide topics according to the interests and concerns of drivers who are driving long distances, select appropriate topics based on location information, adjust the content of the conversation according to their emotional state, and synchronize voice utterances with the navigation device to ensure safe driving.
[1104] "Means for analyzing a user's interests based on their past activities and providing appropriate topics" refers to a device or program that collects data such as a user's past search history and purchase history, analyzes it to identify areas and topics that the user is likely to be interested in, and generates and provides topics based on that data.
[1105] "Means for acquiring location information and selecting topics related to that information" refers to a device or program that identifies the user's current location using a location information acquisition means such as GPS, and selects appropriate topics based on historical background, tourist attractions, local characteristics, etc. related to that location information.
[1106] "Means for recognizing a user's emotions and adjusting the topic according to that state" refers to a device or program that uses voice and facial expression analysis technology to analyze the content of a user's speech and facial expressions in real time, recognizes the user's emotional state (e.g., joy, surprise, fatigue, stress, etc.), and selects the topic that best suits the recognized emotion.
[1107] The "means for controlling synchronization of voice utterances with a car navigation device" refers to a control device or program that monitors voice instructions issued by a car navigation system in real time, automatically stops voice interaction with the user when a navigation instruction is given, and resumes conversation after the instruction is completed.
[1108] "Control means for automatically stopping interaction with the user when the car navigation device gives a voice instruction" refers to a device or program that has the function of detecting when the car navigation system is about to give the next voice instruction and temporarily stopping the conversation with the user during that time.
[1109] The present invention is a system that aims to prevent drivers from getting drowsy during long-distance driving and to allow them to enjoy a fun and comfortable drive. The system is configured using the following hardware and software.
[1110] System Overview
[1111] 1. Server Role
[1112] Analyze and recommend topics based on user's past activities
[1113] The server collects the user's past search history and purchase history, and by analyzing this data, identifies the areas the user is interested in. Based on this data, it generates an appropriate topic list.
[1114] Example: If users are doing a lot of music-related searches and purchases, prioritize music topics in your list.
[1115] Obtain location information and select related topics
[1116] The server receives GPS location information sent from the device and selects topics related to a specific location based on that information.
[1117] For example: If the user is near a tourist attraction, choose a topic related to the history and attractions of that place.
[1118] User Emotion Recognition and Topic Adjustment
[1119] The server receives the user's voice or facial expression analysis data sent from the terminal, evaluates the user's emotional state using an emotion engine, and adjusts the topic based on this evaluation to provide conversation content appropriate to the user's emotional state.
[1120] Example: If the emotion engine detects that the user is tired, it will provide relaxing music or topics.
[1121] 2. Role of the terminal
[1122] Acquisition and transmission of GPS location information
[1123] The device periodically receives GPS signals, acquires its current location information, and sends it to the server.
[1124] Emotion recognition by emotion engine
[1125] The device is equipped with a facial recognition camera and microphone to analyze the user's voice and facial expressions in real time, and the analysis results are sent to a server.
[1126] Conducting a voice interaction
[1127] The device uses speech synthesis technology to converse with the user based on the topic data received from the server. When the user responds by voice, the device analyzes the voice and sends the analysis results back to the server.
[1128] For example, you could say, "Hello. There is a historic building nearby. Do you know about this place?"
[1129] 3. Linkage with car navigation devices
[1130] The terminal works in conjunction with the car navigation system, automatically pausing the conversation with the user when a voice command is given, and resuming the conversation after the command is completed.
[1131] Prompt Sentence Examples
[1132] "There's a historic building nearby. Would you like to know about this place? And if you're interested in music, would you like to tell me about a new hit song?"
[1133] This system can select appropriate topics based on the user's past activity information, current location information, and emotional information, supporting a fun and comfortable drive. By providing interesting topics based on location and emotional information, the user can stay relaxed and alert while driving. This invention achieves more personalized interactions by using a generative AI model and an emotion recognition engine.
[1134] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1135] Step 1:
[1136] Consent verification and user data collection
[1137] Input: User ID
[1138] How it works: When a user launches the app, the server sends a consent prompt to the device.
[1139] When the user presses the consent button, the information is sent back to the server via the terminal.
[1140] Output: User consent information
[1141] Data processing: The server collects and analyzes users' past search history and purchase history.
[1142] Output: User's interests and a list of topics based on them
[1143] Step 2:
[1144] Acquiring and sending location information
[1145] Input: None
[1146] How it works: The device periodically receives GPS signals and obtains its current location.
[1147] Output: Location information
[1148] Data processing: The device sends the acquired location information to the server.
[1149] Output: Location data
[1150] Step 3:
[1151] Topic selection
[1152] Input: User's interest list, location
[1153] How it works: The server matches the location information with the user's interest list.
[1154] The server selects interesting topics related to the location and formats them into conversational format.
[1155] Output: topic list
[1156] Data processing: Selecting topics based on user interests and location information
[1157] Output: A formatted list of conversation topics
[1158] Step 4:
[1159] emotion recognition
[1160] Input: Voice data, facial expression data
[1161] Operation:
[1162] The device collects the user's voice and facial expressions using a microphone and a facial recognition camera.
[1163] The device uses an emotion engine to analyze emotions in real time.
[1164] Output: Emotion data
[1165] Data processing: Emotion evaluation as a result of analyzing voice data and facial expression data
[1166] Output: User's emotional state
[1167] Step 5:
[1168] Adjusting the topic based on emotions
[1169] Input: Topic list, sentiment data
[1170] Operation:
[1171] The server evaluates the user's current emotional state based on the emotion engine's evaluation.
[1172] The server adjusts the topic list according to the emotional state.
[1173] Output: Adjusted topic list
[1174] Data processing: Selecting topics optimized for emotions
[1175] Output: Adjusted topic list
[1176] Step 6:
[1177] Voice Interactions
[1178] Input: adjusted topic list
[1179] Operation:
[1180] The device uses speech synthesis technology to speak aloud based on topic data.
[1181] When the user responds by voice, the terminal analyzes the voice and sends the analysis results to the server.
[1182] Output: Audio data, analysis results
[1183] Data processing: Selecting the next topic based on the results of voice data analysis
[1184] Output: Next topic or response
[1185] Step 7:
[1186] Linking with car navigation systems
[1187] Input: Car navigation voice instructions
[1188] Operation:
[1189] The terminal constantly monitors the voice instructions of the car navigation device.
[1190] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[1191] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[1192] Output: Trigger for the end of voice command
[1193] Data processing: Conversation control based on voice instructions for car navigation
[1194] Output: The resumed conversation
[1195] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1196] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1197] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1198] [Third embodiment]
[1199] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1200] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1201] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1202] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1203] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1205] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1206] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1207] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1208] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1209] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1210] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1211] System Overview
[1212] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to allow them to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities and provides appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, thereby ensuring driving safety.
[1213] Program processing overview and specific examples
[1214] Initial Setup
[1215] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[1216] 1. The server sends a prompt to the device asking the user for consent to data collection.
[1217] 2. When the user presses the consent button, the information is sent back to the server.
[1218] 3. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[1219] Example: If a user has made a lot of travel-related searches and purchases in the past, use that information to prioritize travel topics.
[1220] Obtaining location information
[1221] The device obtains its current location and sends that information to the server.
[1222] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[1223] 2. The device sends the acquired location information to the server.
[1224] Example: If the user is near a tourist attraction in Tokyo, the current location information is sent to the server that the user is near Tokyo Tower.
[1225] Topic selection
[1226] The server selects topics based on interests and location.
[1227] 1. The server matches the received location information with the user's interest list.
[1228] 2. The server picks up topics related to the user's interests and location information and sends them to the device.
[1229] Example: If the current location is near Tokyo Tower and the user is interested in architecture, select the topic "The history of the architecture of Tokyo Tower."
[1230] Start a conversation
[1231] The device will interact via voice based on the selected topic.
[1232] 1. The device uses speech synthesis technology to speak to the user based on the topic data received from the server.
[1233] Example: The device says, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[1234] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[1235] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1236] Linking with car navigation systems
[1237] The device synchronizes with the car navigation system's speech.
[1238] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1239] 2. When the car navigation system issues voice instructions, the device automatically interrupts interaction with the user.
[1240] Example: The device pauses voice conversation while the car navigation system says, "Turn right at the next traffic light."
[1241] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1242] The system takes into account the user's individual interests and current geographical environment to provide a safe and comfortable drive while preventing drowsiness while driving.
[1243] The processing flow will be explained below.
[1244] Step 1:
[1245] The server confirms the user's consent
[1246] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[1247] 2. The device displays a consent confirmation screen to the user.
[1248] 3. When the user presses the consent button, the information is sent back to the server via the device.
[1249] Step 2:
[1250] The server collects and analyzes the user's past activity data.
[1251] 1. Once the server receives the user's consent, it will collect usage history data for LINE Yahoo! services (e.g., search history, purchase history).
[1252] 2. The server analyzes the collected data and identifies the user's areas of interest.
[1253] Example: If a user searches for and purchases a lot of automotive-related items, tag their interests as automotive-related.
[1254] 3. The server generates a priority list of interests based on the analysis results.
[1255] Step 3:
[1256] The device acquires location information using GPS
[1257] 1. The device periodically receives GPS signals and obtains its current location information.
[1258] 2. The device sends the acquired location information to the server.
[1259] Example: If your current location is a tourist spot in a city, send that information.
[1260] Step 4:
[1261] The server selects topics based on interests and location information
[1262] 1. The server matches the location information with the user's interest list.
[1263] 2. The server selects interesting topics related to the location.
[1264] Example: If the location is near Tokyo Tower and the user is interested in architecture, select topics related to the history of architecture.
[1265] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[1266] Step 5:
[1267] The device interacts with the user through voice
[1268] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1269] Example: Say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[1270] 2. When the user responds by voice, the device analyzes the voice.
[1271] 3. The device sends the analysis results to the server and requests the next conversation content.
[1272] Step 6:
[1273] The terminal monitors and controls the voice instructions of the car navigation device
[1274] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1275] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[1276] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1277] For example: "Let's talk about the construction of Tokyo Tower."
[1278] This allows users to avoid drowsiness while driving and enjoy conversations about interesting topics. In addition, safe driving is ensured by linking with car navigation systems.
[1279] Example 1
[1280] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1281] Conventional long-distance driving assistance systems have had problems with drivers feeling drowsy, reducing driving safety. They also lacked entertainment and information provision while driving, making it impossible to provide topics that matched the driver's interests. Furthermore, if voice utterances were not properly synchronized with the car navigation system, the driver's attention could be distracted, potentially leading to an accident.
[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1283] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, means for collecting data with the user's consent, and means for analyzing the user's responses and generating the next topic or response based on the analysis results. This allows for a safe and comfortable drive while preventing drowsiness while taking into account the user's individual interests and current geographical environment.
[1284] "Means of analyzing a user's interests and concerns based on their past activities and providing appropriate topics" refers to a system that collects past behavioral data such as a user's search history and purchase history, analyzes this data to identify the user's interests and concerns, and automatically selects and provides related topics based on that.
[1285] "Means for obtaining location information and selecting topics related to that information" refers to a mechanism in which the user's device obtains current geographical location information using GPS or other means, and selects interesting related topics based on this location information.
[1286] "Means for interacting with the user via voice using a selected topic" refers to a system that uses voice synthesis technology to provide information to the user via voice based on a selected topic, and allows the user to respond via voice, thereby achieving two-way communication.
[1287] The "means for controlling synchronization of voice utterances with the car navigation device" is a mechanism for starting or interrupting voice interaction with the user at appropriate timing in synchronization with voice instructions from the car navigation device.
[1288] "Means for collecting data with user consent" refers to a mechanism by which the system sends a prompt to the user requesting consent for data collection, obtains consent from the user, and collects the necessary data based on that consent.
[1289] "Means for analyzing the user's response and generating the next topic or response based on the analysis results" refers to a mechanism that analyzes the content of the user's voice response and generates the next topic or an appropriate response based on the analysis results.
[1290] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user through voice and further controls the synchronization of voice utterances with the car navigation system to ensure driving safety.
[1291] Initial Setup
[1292] The server confirms the user's consent and collects the necessary data. First, the server sends a prompt to the device to obtain the user's consent. This prompt includes a message such as "Please consent to data collection." When the user presses the consent button, the information is returned to the server, which stores it in a database. The server then collects the user's past search history, purchase history, etc., and uses this data to analyze areas of interest and generate a topic list. A natural language processing algorithm is preferably used in this process. Specific techniques that can be used include data analysis tools such as Python and R.
[1293] Obtaining location information
[1294] The device obtains its current location and sends that information to the server. The device (car navigation system or smartphone) periodically obtains its current location information using a GPS module. For example, it can be set to update its location information every 5 minutes. The obtained location information is sent to the server via an HTTP POST request, and the server stores this information in a database.
[1295] Topic selection
[1296] The server compares the received location information with the user's interest list and selects a topic based on that interest. The server uses database searches to compare the location information with the user's areas of interest. The selected topic is sent to the device in JSON format or similar. For example, if the current location is near Tokyo Tower and the user is interested in architecture, the topic "The architectural history of Tokyo Tower" will be selected.
[1297] Start a conversation
[1298] The device uses speech synthesis technology to speak to the user based on the selected topic. Specifically, it uses the Google Text-to-Speech API to say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" When the user responds verbally, the device analyzes this speech and converts it into text data using the Google Speech-to-Text API or similar. The analyzed data is sent to the server, which generates the next topic and response based on the analysis results. At this stage, a generative AI model (such as GPT-4) is used.
[1299] Linking with car navigation systems
[1300] The device constantly monitors voice instructions from the car navigation system. This is achieved through API integration between the car navigation system and the device. When the car navigation system issues a voice instruction, the device automatically suspends interaction with the user. For example, when the car navigation system says, "Turn right at the next traffic light," the voice conversation is paused. When the car navigation system finishes issuing the instruction, the device resumes the conversation with the user.
[1301] Specific examples and examples of prompts for generative AI models
[1302] As a concrete example, if a user asks, "How tall is Tokyo Tower?", the server will generate a response based on the analysis results and reply, "Tokyo Tower is approximately 333 meters tall."
[1303] An example of a prompt for a generative AI model is:
[1304] "Based on your past search history, what are some travel topics that have recently interested you?"
[1305] This system takes into account the user's individual interests and geographical environment, making it possible to provide a safe and comfortable drive while preventing drowsiness while driving.
[1306] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1307] Step 1: Confirm user agreement
[1308] The server sends a prompt to the device to obtain the user's consent. This uses an HTTP request and includes the message "Please consent to data collection." The input is a consent confirmation message, and the output is a prompt display. Specifically, the server sends an HTTP request to the user's device, and the consent confirmation message is displayed on the device.
[1309] Step 2: Submit consent data
[1310] When the device receives the user's consent, it sends it back to the server. Specifically, it uses an HTTP POST request to send the user's consent data. The input is the user's consent, and the output is the transmission of the consent data to the server. Specifically, the device recognizes the user's "Agree" operation and sends the consent data to the server.
[1311] Step 3: Collect and analyze user data
[1312] The server collects the user's past search history and purchase history and analyzes their areas of interest based on this data. The collected data includes search history, purchase history, etc., and analyzes the text data using a natural language processing algorithm. The input is the user's past search history and purchase history, and the output is the user's areas of interest and a topic list. Specifically, the server retrieves the user's history from the database and inputs it into the interest analysis algorithm.
[1313] Step 4: Obtaining location information
[1314] Devices (car navigation systems and smartphones) periodically obtain their current location information using a GPS module. The input is GPS location data, and the output is sending the location information to a server. Specifically, the device updates its location information every five minutes and sends that data to the server via an HTTP POST request.
[1315] Step 5: Send location information
[1316] The device sends the acquired location information to the server. This information is sent to the server in JSON format or similar. The acquired location information is used as input, and the location information is sent to the server as output. Specifically, the device sends coordinate data from the GPS to the server.
[1317] Step 6: Matching and selecting topics
[1318] The server compares the received location information with the user's interest list and selects relevant topics. This process uses a database search. The location information and interest list are input, and the selected topics are output and sent to the device. Specifically, the server searches the database using an SQL query to extract relevant topics.
[1319] Step 7: Submit topic data
[1320] The server sends the selected topic to the terminal using a data format such as JSON. The selected topic is used as input, and the topic data is sent to the terminal as output.
[1321] Step 8: Uttering topic data
[1322] The device uses speech synthesis technology to speak to the user based on the topic data received from the server. The received topic data is used as input, and the voice speech is used as output. Specifically, using the Google Text-to-Speech API, the device speaks, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[1323] Step 9: Analyze user responses
[1324] The user responds by voice, and the device analyzes the content. The user's voice response is the input, and the analysis results are sent to the server as the output. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text and send it to the server.
[1325] Step 10: Generate response data
[1326] The server generates the next topic and response based on the analysis results and sends them to the device. The input is the analysis result of the user's response, and the output is the generated response. Specifically, a generative AI model (e.g., GPT-4) is used to generate an appropriate response.
[1327] Step 11: Monitor car navigation voice instructions
[1328] The device constantly monitors the voice instructions from the car navigation system, and when these are heard, the device pauses the interaction. The input is the voice instructions from the car navigation system, and the output is pausing the interaction. Specifically, the device monitors the voice instructions through the car navigation API, and pauses the interaction when it says "Turn right at the next traffic light."
[1329] Step 12: Restart the conversation after receiving instructions from the car navigation system
[1330] When the car navigation system's voice instructions end, the device resumes the conversation with the user. The input is the end of the car navigation system's voice instructions, and the output is the resumption of the interaction. Specifically, the device confirms that the car navigation system's instructions have ended, and resumes the voice interaction based on the previous topic and the next topic.
[1331] (Application example 1)
[1332] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1333] During long-distance driving, drivers can become bored or drowsy, increasing the risk of accidents. Even in self-driving vehicles, passengers can lose interest and become bored. It is essential to prevent this and ensure a safe and comfortable drive.
[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1335] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, and means for synchronizing the interaction with safety instructions from the automated driving system and pausing the interaction under certain conditions, thereby enabling drivers and passengers to maintain their interest even during long-distance driving, ensuring a comfortable and safe drive.
[1336] "User's past activities" is a general term for behavioral data such as a user's past search history, purchase history, and browsing history.
[1337] "Interests and concerns" refers to the tendency or degree of interest a user has in a particular theme or topic.
[1338] "Location information" refers to information indicating a user's current location and past travel routes obtained using GPS or other location measurement technologies.
[1339] "Topics" refer to information and content of conversations provided to users, and are selected based on the user's interests and location information.
[1340] "Voice interaction" refers to communication between a user and a system through voice, including the user asking questions and responding by voice.
[1341] A "car navigation device" is a device that displays map information for a vehicle and provides route guidance to a destination.
[1342] "Synchronization of voice utterances" refers to the timely coordination of voice guidance from the car navigation device and voice interaction with the user.
[1343] An "automated driving system" is a system that enables a vehicle to drive automatically, controlling the vehicle without the need for driver intervention.
[1344] "Safety instructions" refers to guidance and warning information provided by the automated driving system to ensure operational safety.
[1345] "Means for temporarily suspending interaction under certain conditions" refers to a control method for temporarily halting voice dialogue with the user when a safety instruction is issued by a car navigation device or an automated driving system.
[1346] This invention provides a system for preventing drowsiness during long-distance driving and allowing drivers to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities, provides appropriate topics, acquires location information and selects topics related to that information, and interacts with the user through voice using the selected topics. It also includes a function for controlling the synchronization of voice utterances with a car navigation system and synchronizing interactions with safety instructions from an automated driving system.
[1347] System configuration
[1348] The system is realized with the following components:
[1349] 1. Server: Collects and analyzes data such as your past search and purchase history to identify your interests.
[1350] 2. Terminal: Includes the dashboard, GPS module, and speech synthesis and recognition system installed in the vehicle.
[1351] 3. Car navigation device: Displays map information for the vehicle and provides guidance on the route to the destination.
[1352] 4. Autonomous driving system: A system in which a vehicle drives itself.
[1353] 5. Speech synthesis and recognition engine: An engine for realizing voice interaction with the user. Specifically, Google Cloud Text-to-Speech and Speech-to-Text APIs are used.
[1354] Program processing overview
[1355] Below is a description of how each component works together to process and calculate data:
[1356] Initial Setup
[1357] The server verifies the user's consent and collects the necessary data. The specific steps are as follows:
[1358] 1. The device prompts the user to consent to data collection.
[1359] 2. The user consents and the information is returned to the server.
[1360] 3. The server collects and analyzes the user's past activity data, identifies areas of interest, and generates a list of topics.
[1361] Obtaining location information
[1362] The device obtains its current location and sends that information to the server. The specific steps are as follows:
[1363] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[1364] 2. The acquired location information is sent to the server.
[1365] Topic selection
[1366] The server selects topics based on interests and location, as follows:
[1367] 1. The server matches the received location information with the user's interest list.
[1368] 2. Pick up topics related to the user's interests and location information and send them to the device.
[1369] Start a conversation
[1370] The device will then interact with the user through voice based on the selected topic. The specific steps are as follows:
[1371] 1. The device uses speech synthesis technology to provide topics to the user.
[1372] 2. The user responds verbally, and the device analyzes the speech.
[1373] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1374] Cooperation with autonomous driving systems
[1375] The terminal synchronizes the safety instructions and interactions of the automated driving system. The specific steps are as follows:
[1376] 1. The device constantly monitors the voice instructions of the autonomous driving system.
[1377] 2. When an autonomous driving safety instruction is issued, the system pauses interaction with the user.
[1378] 3. When the instruction is completed, the interaction resumes.
[1379] Specific examples
[1380] For example, if a user has been interested in travel in the past and is currently near Tokyo Tower, the server will select a topic about the "history of Tokyo Tower's architecture." The device will then say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" If the autonomous driving system issues the instruction "Turn right at the next traffic light," the infotainment system will pause and resume the conversation after the instruction is completed.
[1381] Prompt Sentence Examples
[1382] Here are some example prompts to input to a generative AI model:
[1383] "Design an infotainment system for autonomous vehicles that can provide topical information related to the driver's current location and enable voice interaction based on the driver's interests."
[1384] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1385] Step 1:
[1386] The server sends a prompt to the terminal asking the user for consent to data collection.
[1387] Input: The server receives the data needed to generate the consent prompt (e.g., the consent message).
[1388] Data calculation: The server generates a prompt message.
[1389] Output: The generated prompt message is sent to the terminal.
[1390] Step 2:
[1391] The terminal displays the consent prompt received from the server to the user.
[1392] Input: The terminal receives the prompt message.
[1393] Data processing: The terminal displays a prompt message on the user interface.
[1394] Output: User's choice to accept or decline.
[1395] Step 3:
[1396] The user presses the consent button and the information is sent back to the server.
[1397] Input: User's consent or denial of information.
[1398] Data calculation: The terminal analyzes the user's selection and generates consent information.
[1399] Output: The consent information is sent to the server.
[1400] Step 4:
[1401] The server collects data such as the user's past search and purchase history to identify areas of interest.
[1402] Input: The server receives the user's activity data along with consent information.
[1403] Data Calculation: The server collects the user's past activity data from a database and uses analysis algorithms to identify areas of interest.
[1404] Output: A list of areas of interest is generated.
[1405] Step 5:
[1406] The device obtains its current location and sends that information to the server.
[1407] Input: The device receives current location information from the GPS module.
[1408] Data processing: The acquired location information is converted into a format suitable for the server.
[1409] Output: The converted current location information is sent to the server.
[1410] Step 6:
[1411] The server compares the received location information with the interest list, selects appropriate topics, and sends them to the device.
[1412] Input: The server receives the current location and a list of interests.
[1413] Data calculation: The server matches the location information with the list of areas of interest and generates related topics.
[1414] Output: A list of relevant topics is generated and sent to the terminal.
[1415] Step 7:
[1416] The terminal uses speech synthesis technology to speak to the user based on the topic data received from the server.
[1417] Input: The terminal receives topic data.
[1418] Data processing: Using voice synthesis technology, the topic data is converted into a voice message.
[1419] Output: The audio message is spoken to the user.
[1420] Step 8:
[1421] The user responds with voice, which is then analyzed by the device.
[1422] Input: The user's spoken response.
[1423] Data Computing: A speech recognition engine converts the user's response into text and analyzes it.
[1424] Output: The analysis results are sent to the server.
[1425] Step 9:
[1426] The server generates the next topic and response based on the analysis results and sends them back to the terminal.
[1427] Input: The server receives the user's voice analysis results.
[1428] Data calculation: The server generates an appropriate response based on the analysis results.
[1429] Output: The generated response is sent to the terminal.
[1430] Step 10:
[1431] When the autonomous driving system issues safety instructions, the device temporarily stops interacting with the user.
[1432] Input: The terminal receives safety instructions from the automated driving system.
[1433] Data processing: Pause the interaction.
[1434] Output: Voice interaction with the user stops during the safety instruction period.
[1435] Step 11:
[1436] Once the safety instructions from the automated driving system are complete, the device resumes interaction with the user.
[1437] Input: Notification of end of safety instruction.
[1438] Data processing: Restarting a stopped interaction.
[1439] Output: Voice interaction with the user is resumed.
[1440] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1441] System Overview
[1442] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[1443] Program processing overview and specific examples
[1444] Initial Setup
[1445] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[1446] 1. The server sends a prompt to the user's device asking for consent to data collection.
[1447] 2. The device displays a consent confirmation screen to the user.
[1448] 3. When the user presses the consent button, the information is sent back to the server via the device.
[1449] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[1450] Example: If a user has made many music-related searches and purchases in the past, tag their interests as music-related and add music topics to the priority list.
[1451] Obtaining location information
[1452] The device obtains location information using GPS and sends that information to the server.
[1453] 1. The device periodically receives GPS signals and obtains its current location information.
[1454] 2. The device sends the acquired location information to the server.
[1455] Example: If your current location is near a tourist attraction, send that information.
[1456] Topic selection
[1457] The server selects topics based on interests and location.
[1458] 1. The server matches the location information with the user's interest list.
[1459] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[1460] Example: If the location is near a historical place and the user is interested in history, select topics about the history of that place.
[1461] Emotion recognition by emotion engine
[1462] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[1463] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[1464] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[1465] Example: If the user is tired or stressed, the emotion engine will recognize that emotional state.
[1466] Adjusting the topic based on emotions
[1467] The server adjusts the topic based on the user's emotional state.
[1468] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[1469] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[1470] For example, if the user is tired, select relaxing music or topics that will help them relax.
[1471] Start a conversation
[1472] The device will interact via voice based on the selected topic.
[1473] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1474] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[1475] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[1476] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1477] Linking with car navigation systems
[1478] The terminal monitors and controls the voice instructions of the car navigation device.
[1479] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1480] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[1481] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1482] For example: "Let me tell you about the history of this area."
[1483] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[1484] The processing flow will be explained below.
[1485] Step 1:
[1486] The server confirms the user's consent
[1487] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[1488] 2. The device displays a consent confirmation screen to the user.
[1489] 3. When the user presses the consent button, the information is sent back to the server via the device.
[1490] Step 2:
[1491] The server collects and analyzes the user's past activity data.
[1492] 1. Once the server receives the user's consent, it collects usage history data (e.g., search history, purchase history) of the relevant online service.
[1493] 2. The server analyzes the collected data and identifies the user's areas of interest.
[1494] Example: If a user does a lot of movie-related searches and purchases, tag their interests as movie-related.
[1495] 3. The server generates a priority list of interests based on the analysis results.
[1496] Step 3:
[1497] The device acquires location information using GPS
[1498] 1. The device periodically receives GPS signals and obtains its current location information.
[1499] 2. The device sends the acquired location information to the server.
[1500] Example: If your current location is near a tourist attraction, send that information.
[1501] Step 4:
[1502] The server selects topics based on interests and location information
[1503] 1. The server matches the location information with the user's interest list.
[1504] 2. The server selects interesting topics related to the location.
[1505] Example: If the location is near a historical landmark and the user is interested in history, select topics related to the history of that place.
[1506] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[1507] Step 5:
[1508] The device recognizes the user's emotions
[1509] 1. The device uses an emotion engine to analyze the user's voice and facial expressions.
[1510] 2. The device assesses the user's emotional state (e.g., happy, sad, tired, etc.).
[1511] 3. The device sends the evaluation results to the server.
[1512] Step 6:
[1513] The server adjusts the topic to match the emotional state
[1514] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[1515] 2. The server selects topics according to the user's emotional state and provides conversation content appropriate to the user's emotions.
[1516] For example, if the user is tired, select a topic or music that has a relaxing effect.
[1517] Step 7:
[1518] The device interacts with the user through voice
[1519] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1520] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[1521] 2. When the user responds by voice, the device analyzes the voice.
[1522] 3. The device sends the analysis results to the server and requests the next conversation content.
[1523] Step 8:
[1524] The terminal monitors and controls the voice instructions of the car navigation device
[1525] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1526] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[1527] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1528] For example: "Let me tell you about the history of this area."
[1529] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[1530] Example 2
[1531] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1532] It is difficult for drivers to enjoy a comfortable long-distance drive without feeling drowsy. It is also not easy to provide appropriate interactions based on the user's interests, current location, and emotional state. In particular, current technology is not sufficient to synchronize the voice instructions of the car navigation system with the conversation between the user and the car.
[1533] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1534] In this invention, the server includes means for analyzing the user's interests and concerns based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for engaging in voice interaction with the user using the selected topics, and means for acquiring the user's emotional data using an emotion engine and adjusting the topics according to the emotional state. This makes it possible to provide appropriate interactions according to the user's interests, location information, and emotional state, prevent drowsiness during long-distance driving, and support a comfortable drive.
[1535] "User" means a person who operates or uses a particular system or device.
[1536] "Past activity" refers to behavioral data such as a user's past search history and purchase history.
[1537] "Means for analyzing interests and concerns" refers to technology or devices for analyzing a user's past activity data to identify what the user is interested in.
[1538] "Means for providing topics" refers to technologies and devices for presenting appropriate topics to users based on the analyzed data.
[1539] "Location Information" refers to a user's current physical location data obtained through technologies such as GPS.
[1540] "Means for obtaining location information" refers to technologies and devices that use GPS or other location measurement technologies to determine a user's current location.
[1541] "Means for selecting a topic" refers to a technology or device for selecting a topic appropriate for a user based on the acquired location information.
[1542] "Means for voice interaction" refers to technology and devices for using voice synthesis technology to have a voice dialogue with a user.
[1543] "Car navigation device" refers to an electronic device that displays the current location of a vehicle and provides directions to a destination.
[1544] The "means for controlling synchronization of voice utterances" refers to a technique or device for appropriately adjusting the voice instructions of the car navigation device and the voice dialogue with the user.
[1545] An "emotion engine" refers to technology or a device that analyzes a user's voice and facial expressions to assess their emotional state.
[1546] "Means for acquiring emotional data" refers to technology or devices for acquiring the user's emotional state in real time using an emotion engine.
[1547] "Means for adjusting the topic" refers to technology or devices for changing or adapting the topic provided to the user based on the acquired emotional data.
[1548] System Overview
[1549] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[1550] Program processing overview and specific examples
[1551] Initial Setup
[1552] The server verifies the user's consent and collects the necessary data.
[1553] 1. The server sends a prompt to the user's device asking for consent to data collection.
[1554] 2. The device displays a consent confirmation screen to the user.
[1555] 3. When the user presses the consent button, the information is sent back to the server via the device.
[1556] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[1557] Examples:
[1558] If a user has made many music-related searches or purchases in the past, we tag their interests as music-related and add music topics to the priority list.
[1559] Obtaining location information
[1560] The device obtains location information using GPS and sends that information to the server.
[1561] 1. The device periodically receives GPS signals and obtains its current location information.
[1562] 2. The device sends the acquired location information to the server.
[1563] Examples:
[1564] If the location information is near a tourist attraction, the information is sent.
[1565] Topic selection
[1566] The server selects topics based on interests and location.
[1567] 1. The server matches the location information with the user's interest list.
[1568] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[1569] Examples:
[1570] If the location information is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[1571] Emotion recognition by emotion engine
[1572] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[1573] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[1574] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[1575] Examples:
[1576] If the user is tired or stressed, the emotion engine will recognize that emotional state.
[1577] Adjusting the topic based on emotions
[1578] The server adjusts the topic based on the user's emotional state.
[1579] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[1580] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[1581] Examples:
[1582] If the user is tired, select relaxing music or topics that will help them relax.
[1583] Start a conversation
[1584] The device will interact via voice based on the selected topic.
[1585] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1586] Examples:
[1587] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[1588] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[1589] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1590] Linking with car navigation systems
[1591] The terminal monitors and controls the voice instructions of the car navigation device.
[1592] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1593] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[1594] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1595] Examples:
[1596] "Let me tell you about the history of this area," he resumes.
[1597] Prompt Sentence Examples
[1598] Examples:
[1599] A user is on a long drive and the car navigation system tells them to "turn left at the next traffic light."
[1600] Example prompt for a generative AI model:
[1601] "You're interested in history and are currently driving near a historical landmark. After the navigation system finishes giving you current directions, you can start a conversation about the history of that landmark."
[1602] This system can provide appropriate topics according to the user's emotional state and interests while driving, supporting a safe and comfortable drive.
[1603] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1604] Step 1:
[1605] The server sends a user consent prompt to the terminal.
[1606] Specific behavior:
[1607] The server creates a prompt to obtain the user's consent to data collection and sends the prompt to the terminal.
[1608] Input: Consent prompt template information, User ID.
[1609] Output: The user consent prompt sent to the device.
[1610] Step 2:
[1611] The terminal displays a consent confirmation screen to the user.
[1612] Specific behavior:
[1613] The terminal generates a consent confirmation screen based on the received consent prompt and displays it to the user.
[1614] Input: The consent prompt received from the server.
[1615] Output: The consent screen shown to the user.
[1616] Step 3:
[1617] The user presses the consent button and the information is sent back to the server via the terminal.
[1618] Specific behavior:
[1619] When the user presses the accept button, their input is captured on the device.
[1620] The terminal returns the user's consent information to the server.
[1621] Input: User consent.
[1622] Output: User consent information sent to the server.
[1623] Step 4:
[1624] The server collects the user's past search and purchase history, analyzes their areas of interest, and generates a topic list.
[1625] Specific behavior:
[1626] The server collects the user's past search history and purchase history from a database.
[1627] The server analyzes the collected data to identify the user's areas of interest.
[1628] The server generates a topic list appropriate for the user based on the identified areas of interest.
[1629] Input: User's past search history, purchase history.
[1630] Output: Analyzed areas of interest, generated topic list.
[1631] Examples:
[1632] If a user has made many music-related searches or purchases in the past, the server will use that information to add music-related topics to the list.
[1633] Step 5:
[1634] The device receives a GPS signal and sends the location information to the server.
[1635] Specific behavior:
[1636] The device periodically receives GPS signals to obtain its current location.
[1637] The terminal transmits the acquired location information to the server.
[1638] Input: GPS signal.
[1639] Output: Current location information sent to the server.
[1640] Examples:
[1641] If the location information is near a tourist attraction, the information is sent to the server.
[1642] Step 6:
[1643] The server compares the location information with the user's interest list and selects topics.
[1644] Specific behavior:
[1645] The server matches the received location information with the user's interest list.
[1646] The server selects interesting topics related to the location.
[1647] The topics are formatted into a conversational format and sent to the terminal.
[1648] Input: Location, user interest list.
[1649] Output: Conversational topic data sent to the device.
[1650] Examples:
[1651] If the current location is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[1652] Step 7:
[1653] The terminal recognizes the user's emotions and transmits the emotional state to the server.
[1654] Specific behavior:
[1655] The device acquires emotional data using an emotion engine that analyzes the user's voice and facial expressions.
[1656] The device evaluates the acquired emotion data and transmits the information to the server.
[1657] Input: User's voice and facial expression data.
[1658] Output: Emotion data sent to the server.
[1659] Examples:
[1660] Recognizes the user's emotional state when they are tired or stressed.
[1661] Step 8:
[1662] The server adjusts the topic based on the emotional state and sends it to the terminal.
[1663] Specific behavior:
[1664] The server evaluates the user's current emotional state based on the data received from the emotion engine.
[1665] The server selects a topic according to the emotional state and sends it to the terminal.
[1666] Input: Sentiment data, topic list.
[1667] Output: The adjusted topic data sent to the device.
[1668] Examples:
[1669] If the user is tired, select relaxing music or topics that will help them relax.
[1670] Step 9:
[1671] The device speaks the topic and engages in interaction.
[1672] Specific behavior:
[1673] The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1674] When the user responds by voice, the voice is analyzed and the analysis results are sent to the server.
[1675] Input: Topic data from the server, user voice input.
[1676] Output: Audio utterance, analysis results sent to the server.
[1677] Examples:
[1678] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[1679] Step 10:
[1680] The device monitors the car navigation system's voice instructions and controls the interaction.
[1681] Specific behavior:
[1682] The terminal constantly monitors the voice instructions of the car navigation device.
[1683] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[1684] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[1685] Input: Car navigation voice instructions.
[1686] Output: Pause and resume conversation.
[1687] Examples:
[1688] When the car navigation system says, "Turn right at the next traffic light," it pauses the conversation with the user, and then resumes after the instruction is completed, saying, "Let me tell you about the history of this area."
[1689] (Application example 2)
[1690] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1691] Currently, there are not enough systems proposed to prevent drowsiness during long-distance driving and ensure a comfortable and safe drive. Furthermore, no systems have been found that combine a variety of elements, such as providing topics based on the driver's emotional state and interests, selecting topics based on location information, and synchronizing voice output with a car navigation system. The present invention aims to solve these problems and provide a system that prevents drowsiness during long-distance driving and enables drivers to enjoy a fun and comfortable drive.
[1692] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's interests based on their past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for recognizing the user's emotions and adjusting the topics according to that state, and means for controlling synchronization of voice utterances with a car navigation device. This makes it possible to provide topics according to the interests and concerns of drivers who are driving long distances, select appropriate topics based on location information, adjust the content of the conversation according to their emotional state, and synchronize voice utterances with the navigation device to ensure safe driving.
[1693] "Means for analyzing a user's interests based on their past activities and providing appropriate topics" refers to a device or program that collects data such as a user's past search history and purchase history, analyzes it to identify areas and topics that the user is likely to be interested in, and generates and provides topics based on that data.
[1694] "Means for acquiring location information and selecting topics related to that information" refers to a device or program that identifies the user's current location using a location information acquisition means such as GPS, and selects appropriate topics based on historical background, tourist attractions, local characteristics, etc. related to that location information.
[1695] "Means for recognizing a user's emotions and adjusting the topic according to that state" refers to a device or program that uses voice and facial expression analysis technology to analyze the content of a user's speech and facial expressions in real time, recognizes the user's emotional state (e.g., joy, surprise, fatigue, stress, etc.), and selects the topic that best suits the recognized emotion.
[1696] The "means for controlling synchronization of voice utterances with a car navigation device" refers to a control device or program that monitors voice instructions issued by a car navigation system in real time, automatically stops voice interaction with the user when a navigation instruction is given, and resumes conversation after the instruction is completed.
[1697] "Control means for automatically stopping interaction with the user when the car navigation device gives a voice instruction" refers to a device or program that has the function of detecting when the car navigation system is about to give the next voice instruction and temporarily stopping the conversation with the user during that time.
[1698] The present invention is a system that aims to prevent drivers from getting drowsy during long-distance driving and to allow them to enjoy a fun and comfortable drive. The system is configured using the following hardware and software.
[1699] System Overview
[1700] 1. Server Role
[1701] Analyze and recommend topics based on user's past activities
[1702] The server collects the user's past search history and purchase history, and by analyzing this data, identifies the areas the user is interested in. Based on this data, it generates an appropriate topic list.
[1703] Example: If users are doing a lot of music-related searches and purchases, prioritize music topics in your list.
[1704] Obtain location information and select related topics
[1705] The server receives GPS location information sent from the device and selects topics related to a specific location based on that information.
[1706] For example: If the user is near a tourist attraction, choose a topic related to the history and attractions of that place.
[1707] User Emotion Recognition and Topic Adjustment
[1708] The server receives the user's voice or facial expression analysis data sent from the terminal, evaluates the user's emotional state using an emotion engine, and adjusts the topic based on this evaluation to provide conversation content appropriate to the user's emotional state.
[1709] Example: If the emotion engine detects that the user is tired, it will provide relaxing music or topics.
[1710] 2. Role of the terminal
[1711] Acquisition and transmission of GPS location information
[1712] The device periodically receives GPS signals, acquires its current location information, and sends it to the server.
[1713] Emotion recognition by emotion engine
[1714] The device is equipped with a facial recognition camera and microphone to analyze the user's voice and facial expressions in real time, and the analysis results are sent to a server.
[1715] Conducting a voice interaction
[1716] The device uses speech synthesis technology to converse with the user based on the topic data received from the server. When the user responds by voice, the device analyzes the voice and sends the analysis results back to the server.
[1717] For example, you could say, "Hello. There is a historic building nearby. Do you know about this place?"
[1718] 3. Linkage with car navigation devices
[1719] The terminal works in conjunction with the car navigation system, automatically pausing the conversation with the user when a voice command is given, and resuming the conversation after the command is completed.
[1720] Prompt Sentence Examples
[1721] "There's a historic building nearby. Would you like to know about this place? And if you're interested in music, would you like to tell me about a new hit song?"
[1722] This system can select appropriate topics based on the user's past activity information, current location information, and emotional information, supporting a fun and comfortable drive. By providing interesting topics based on location and emotional information, the user can stay relaxed and alert while driving. This invention achieves more personalized interactions by using a generative AI model and an emotion recognition engine.
[1723] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1724] Step 1:
[1725] Consent verification and user data collection
[1726] Input: User ID
[1727] How it works: When a user launches the app, the server sends a consent prompt to the device.
[1728] When the user presses the consent button, the information is sent back to the server via the terminal.
[1729] Output: User consent information
[1730] Data processing: The server collects and analyzes users' past search history and purchase history.
[1731] Output: User's interests and a list of topics based on them
[1732] Step 2:
[1733] Acquiring and sending location information
[1734] Input: None
[1735] How it works: The device periodically receives GPS signals and obtains its current location.
[1736] Output: Location information
[1737] Data processing: The device sends the acquired location information to the server.
[1738] Output: Location data
[1739] Step 3:
[1740] Topic selection
[1741] Input: User's interest list, location
[1742] How it works: The server matches the location information with the user's interest list.
[1743] The server selects interesting topics related to the location and formats them into conversational format.
[1744] Output: topic list
[1745] Data processing: Selecting topics based on user interests and location information
[1746] Output: A formatted list of conversation topics
[1747] Step 4:
[1748] emotion recognition
[1749] Input: Voice data, facial expression data
[1750] Operation:
[1751] The device collects the user's voice and facial expressions using a microphone and a facial recognition camera.
[1752] The device uses an emotion engine to analyze emotions in real time.
[1753] Output: Emotion data
[1754] Data processing: Emotion evaluation as a result of analyzing voice data and facial expression data
[1755] Output: User's emotional state
[1756] Step 5:
[1757] Adjusting the topic based on emotions
[1758] Input: Topic list, sentiment data
[1759] Operation:
[1760] The server evaluates the user's current emotional state based on the emotion engine's evaluation.
[1761] The server adjusts the topic list according to the emotional state.
[1762] Output: Adjusted topic list
[1763] Data processing: Selecting topics optimized for emotions
[1764] Output: Adjusted topic list
[1765] Step 6:
[1766] Voice Interactions
[1767] Input: adjusted topic list
[1768] Operation:
[1769] The device uses speech synthesis technology to speak aloud based on topic data.
[1770] When the user responds by voice, the terminal analyzes the voice and sends the analysis results to the server.
[1771] Output: Audio data, analysis results
[1772] Data processing: Selecting the next topic based on the results of voice data analysis
[1773] Output: Next topic or response
[1774] Step 7:
[1775] Linking with car navigation systems
[1776] Input: Car navigation voice instructions
[1777] Operation:
[1778] The terminal constantly monitors the voice instructions of the car navigation device.
[1779] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[1780] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[1781] Output: Trigger for the end of voice command
[1782] Data processing: Conversation control based on voice instructions for car navigation
[1783] Output: The resumed conversation
[1784] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1785] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1786] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1787] [Fourth embodiment]
[1788] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1789] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1790] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1791] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1792] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1793] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1794] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1795] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1796] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1797] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1798] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1799] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1800] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1801] System Overview
[1802] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to allow them to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities and provides appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, thereby ensuring driving safety.
[1803] Program processing overview and specific examples
[1804] Initial Setup
[1805] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[1806] 1. The server sends a prompt to the device asking the user for consent to data collection.
[1807] 2. When the user presses the consent button, the information is sent back to the server.
[1808] 3. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[1809] Example: If a user has made a lot of travel-related searches and purchases in the past, use that information to prioritize travel topics.
[1810] Obtaining location information
[1811] The device obtains its current location and sends that information to the server.
[1812] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[1813] 2. The device sends the acquired location information to the server.
[1814] Example: If the user is near a tourist attraction in Tokyo, the current location information is sent to the server that the user is near Tokyo Tower.
[1815] Topic selection
[1816] The server selects topics based on interests and location.
[1817] 1. The server matches the received location information with the user's interest list.
[1818] 2. The server picks up topics related to the user's interests and location information and sends them to the device.
[1819] Example: If the current location is near Tokyo Tower and the user is interested in architecture, select the topic "The history of the architecture of Tokyo Tower."
[1820] Start a conversation
[1821] The device will interact via voice based on the selected topic.
[1822] 1. The device uses speech synthesis technology to speak to the user based on the topic data received from the server.
[1823] Example: The device says, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[1824] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[1825] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1826] Linking with car navigation systems
[1827] The device synchronizes with the car navigation system's speech.
[1828] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1829] 2. When the car navigation system issues voice instructions, the device automatically interrupts interaction with the user.
[1830] Example: The device pauses voice conversation while the car navigation system says, "Turn right at the next traffic light."
[1831] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1832] The system takes into account the user's individual interests and current geographical environment to provide a safe and comfortable drive while preventing drowsiness while driving.
[1833] The processing flow will be explained below.
[1834] Step 1:
[1835] The server confirms the user's consent
[1836] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[1837] 2. The device displays a consent confirmation screen to the user.
[1838] 3. When the user presses the consent button, the information is sent back to the server via the device.
[1839] Step 2:
[1840] The server collects and analyzes the user's past activity data.
[1841] 1. Once the server receives the user's consent, it will collect usage history data for LINE Yahoo! services (e.g., search history, purchase history).
[1842] 2. The server analyzes the collected data and identifies the user's areas of interest.
[1843] Example: If a user searches for and purchases a lot of automotive-related items, tag their interests as automotive-related.
[1844] 3. The server generates a priority list of interests based on the analysis results.
[1845] Step 3:
[1846] The device acquires location information using GPS
[1847] 1. The device periodically receives GPS signals and obtains its current location information.
[1848] 2. The device sends the acquired location information to the server.
[1849] Example: If your current location is a tourist spot in a city, send that information.
[1850] Step 4:
[1851] The server selects topics based on interests and location information
[1852] 1. The server matches the location information with the user's interest list.
[1853] 2. The server selects interesting topics related to the location.
[1854] Example: If the location is near Tokyo Tower and the user is interested in architecture, select topics related to the history of architecture.
[1855] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[1856] Step 5:
[1857] The device interacts with the user through voice
[1858] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[1859] Example: Say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[1860] 2. When the user responds by voice, the device analyzes the voice.
[1861] 3. The device sends the analysis results to the server and requests the next conversation content.
[1862] Step 6:
[1863] The terminal monitors and controls the voice instructions of the car navigation device
[1864] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[1865] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[1866] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[1867] For example: "Let's talk about the construction of Tokyo Tower."
[1868] This allows users to avoid drowsiness while driving and enjoy conversations about interesting topics. In addition, safe driving is ensured by linking with car navigation systems.
[1869] Example 1
[1870] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1871] Conventional long-distance driving assistance systems have had problems with drivers feeling drowsy, reducing driving safety. They also lacked entertainment and information provision while driving, making it impossible to provide topics that matched the driver's interests. Furthermore, if voice utterances were not properly synchronized with the car navigation system, the driver's attention could be distracted, potentially leading to an accident.
[1872] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1873] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, means for collecting data with the user's consent, and means for analyzing the user's responses and generating the next topic or response based on the analysis results. This allows for a safe and comfortable drive while preventing drowsiness while taking into account the user's individual interests and current geographical environment.
[1874] "Means of analyzing a user's interests and concerns based on their past activities and providing appropriate topics" refers to a system that collects past behavioral data such as a user's search history and purchase history, analyzes this data to identify the user's interests and concerns, and automatically selects and provides related topics based on that.
[1875] "Means for obtaining location information and selecting topics related to that information" refers to a mechanism in which the user's device obtains current geographical location information using GPS or other means, and selects interesting related topics based on this location information.
[1876] "Means for interacting with the user via voice using a selected topic" refers to a system that uses voice synthesis technology to provide information to the user via voice based on a selected topic, and allows the user to respond via voice, thereby achieving two-way communication.
[1877] The "means for controlling synchronization of voice utterances with the car navigation device" is a mechanism for starting or interrupting voice interaction with the user at appropriate timing in synchronization with voice instructions from the car navigation device.
[1878] "Means for collecting data with user consent" refers to a mechanism by which the system sends a prompt to the user requesting consent for data collection, obtains consent from the user, and collects the necessary data based on that consent.
[1879] "Means for analyzing the user's response and generating the next topic or response based on the analysis results" refers to a mechanism that analyzes the content of the user's voice response and generates the next topic or an appropriate response based on the analysis results.
[1880] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. It uses the selected topics to interact with the user through voice and further controls the synchronization of voice utterances with the car navigation system to ensure driving safety.
[1881] Initial Setup
[1882] The server confirms the user's consent and collects the necessary data. First, the server sends a prompt to the device to obtain the user's consent. This prompt includes a message such as "Please consent to data collection." When the user presses the consent button, the information is returned to the server, which stores it in a database. The server then collects the user's past search history, purchase history, etc., and uses this data to analyze areas of interest and generate a topic list. A natural language processing algorithm is preferably used in this process. Specific techniques that can be used include data analysis tools such as Python and R.
[1883] Obtaining location information
[1884] The device obtains its current location and sends that information to the server. The device (car navigation system or smartphone) periodically obtains its current location information using a GPS module. For example, it can be set to update its location information every 5 minutes. The obtained location information is sent to the server via an HTTP POST request, and the server stores this information in a database.
[1885] Topic selection
[1886] The server compares the received location information with the user's interest list and selects a topic based on that interest. The server uses database searches to compare the location information with the user's areas of interest. The selected topic is sent to the device in JSON format or similar. For example, if the current location is near Tokyo Tower and the user is interested in architecture, the topic "The architectural history of Tokyo Tower" will be selected.
[1887] Start a conversation
[1888] The device uses speech synthesis technology to speak to the user based on the selected topic. Specifically, it uses the Google Text-to-Speech API to say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" When the user responds verbally, the device analyzes this speech and converts it into text data using the Google Speech-to-Text API or similar. The analyzed data is sent to the server, which generates the next topic and response based on the analysis results. At this stage, a generative AI model (such as GPT-4) is used.
[1889] Linking with car navigation systems
[1890] The device constantly monitors voice instructions from the car navigation system. This is achieved through API integration between the car navigation system and the device. When the car navigation system issues a voice instruction, the device automatically suspends interaction with the user. For example, when the car navigation system says, "Turn right at the next traffic light," the voice conversation is paused. When the car navigation system finishes issuing the instruction, the device resumes the conversation with the user.
[1891] Specific examples and examples of prompts for generative AI models
[1892] As a concrete example, if a user asks, "How tall is Tokyo Tower?", the server will generate a response based on the analysis results and reply, "Tokyo Tower is approximately 333 meters tall."
[1893] An example of a prompt for a generative AI model is:
[1894] "Based on your past search history, what are some travel topics that have recently interested you?"
[1895] This system takes into account the user's individual interests and geographical environment, making it possible to provide a safe and comfortable drive while preventing drowsiness while driving.
[1896] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1897] Step 1: Confirm user agreement
[1898] The server sends a prompt to the device to obtain the user's consent. This uses an HTTP request and includes the message "Please consent to data collection." The input is a consent confirmation message, and the output is a prompt display. Specifically, the server sends an HTTP request to the user's device, and the consent confirmation message is displayed on the device.
[1899] Step 2: Submit consent data
[1900] When the device receives the user's consent, it sends it back to the server. Specifically, it uses an HTTP POST request to send the user's consent data. The input is the user's consent, and the output is the transmission of the consent data to the server. Specifically, the device recognizes the user's "Agree" operation and sends the consent data to the server.
[1901] Step 3: Collect and analyze user data
[1902] The server collects the user's past search history and purchase history and analyzes their areas of interest based on this data. The collected data includes search history, purchase history, etc., and analyzes the text data using a natural language processing algorithm. The input is the user's past search history and purchase history, and the output is the user's areas of interest and a topic list. Specifically, the server retrieves the user's history from the database and inputs it into the interest analysis algorithm.
[1903] Step 4: Obtaining location information
[1904] Devices (car navigation systems and smartphones) periodically obtain their current location information using a GPS module. The input is GPS location data, and the output is sending the location information to a server. Specifically, the device updates its location information every five minutes and sends that data to the server via an HTTP POST request.
[1905] Step 5: Send location information
[1906] The device sends the acquired location information to the server. This information is sent to the server in JSON format or similar. The acquired location information is used as input, and the location information is sent to the server as output. Specifically, the device sends coordinate data from the GPS to the server.
[1907] Step 6: Matching and selecting topics
[1908] The server compares the received location information with the user's interest list and selects relevant topics. This process uses a database search. The location information and interest list are input, and the selected topics are output and sent to the device. Specifically, the server searches the database using an SQL query to extract relevant topics.
[1909] Step 7: Submit topic data
[1910] The server sends the selected topic to the terminal using a data format such as JSON. The selected topic is used as input, and the topic data is sent to the terminal as output.
[1911] Step 8: Uttering topic data
[1912] The device uses speech synthesis technology to speak to the user based on the topic data received from the server. The received topic data is used as input, and the voice speech is used as output. Specifically, using the Google Text-to-Speech API, the device speaks, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?"
[1913] Step 9: Analyze user responses
[1914] The user responds by voice, and the device analyzes the content. The user's voice response is the input, and the analysis results are sent to the server as the output. Specifically, the device uses the Google Speech-to-Text API to convert the voice into text and send it to the server.
[1915] Step 10: Generate response data
[1916] The server generates the next topic and response based on the analysis results and sends them to the device. The input is the analysis result of the user's response, and the output is the generated response. Specifically, a generative AI model (e.g., GPT-4) is used to generate an appropriate response.
[1917] Step 11: Monitor car navigation voice instructions
[1918] The device constantly monitors the voice instructions from the car navigation system, and when these are heard, the device pauses the interaction. The input is the voice instructions from the car navigation system, and the output is pausing the interaction. Specifically, the device monitors the voice instructions through the car navigation API, and pauses the interaction when it says "Turn right at the next traffic light."
[1919] Step 12: Restart the conversation after receiving instructions from the car navigation system
[1920] When the car navigation system's voice instructions end, the device resumes the conversation with the user. The input is the end of the car navigation system's voice instructions, and the output is the resumption of the interaction. Specifically, the device confirms that the car navigation system's instructions have ended, and resumes the voice interaction based on the previous topic and the next topic.
[1921] (Application example 1)
[1922] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1923] During long-distance driving, drivers can become bored or drowsy, increasing the risk of accidents. Even in self-driving vehicles, passengers can lose interest and become bored. It is essential to prevent this and ensure a safe and comfortable drive.
[1924] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1925] In this invention, the server includes means for analyzing the user's interests based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to the information, means for engaging in voice interaction with the user using the selected topics, means for controlling synchronization of voice utterances with the car navigation device, and means for synchronizing the interaction with safety instructions from the automated driving system and pausing the interaction under certain conditions, thereby enabling drivers and passengers to maintain their interest even during long-distance driving, ensuring a comfortable and safe drive.
[1926] "User's past activities" is a general term for behavioral data such as a user's past search history, purchase history, and browsing history.
[1927] "Interests and concerns" refers to the tendency or degree of interest a user has in a particular theme or topic.
[1928] "Location information" refers to information indicating a user's current location and past travel routes obtained using GPS or other location measurement technologies.
[1929] "Topics" refer to information and content of conversations provided to users, and are selected based on the user's interests and location information.
[1930] "Voice interaction" refers to communication between a user and a system through voice, including the user asking questions and responding by voice.
[1931] A "car navigation device" is a device that displays map information for a vehicle and provides route guidance to a destination.
[1932] "Synchronization of voice utterances" refers to the timely coordination of voice guidance from the car navigation device and voice interaction with the user.
[1933] An "automated driving system" is a system that enables a vehicle to drive automatically, controlling the vehicle without the need for driver intervention.
[1934] "Safety instructions" refers to guidance and warning information provided by the automated driving system to ensure operational safety.
[1935] "Means for temporarily suspending interaction under certain conditions" refers to a control method for temporarily halting voice dialogue with the user when a safety instruction is issued by a car navigation device or an automated driving system.
[1936] This invention provides a system for preventing drowsiness during long-distance driving and allowing drivers to enjoy a pleasant and comfortable drive. This system analyzes the user's interests based on their past activities, provides appropriate topics, acquires location information and selects topics related to that information, and interacts with the user through voice using the selected topics. It also includes a function for controlling the synchronization of voice utterances with a car navigation system and synchronizing interactions with safety instructions from an automated driving system.
[1937] System configuration
[1938] The system is realized with the following components:
[1939] 1. Server: Collects and analyzes data such as your past search and purchase history to identify your interests.
[1940] 2. Terminal: Includes the dashboard, GPS module, and speech synthesis and recognition system installed in the vehicle.
[1941] 3. Car navigation device: Displays map information for the vehicle and provides guidance on the route to the destination.
[1942] 4. Autonomous driving system: A system in which a vehicle drives itself.
[1943] 5. Speech synthesis and recognition engine: An engine for realizing voice interaction with the user. Specifically, Google Cloud Text-to-Speech and Speech-to-Text APIs are used.
[1944] Program processing overview
[1945] Below is a description of how each component works together to process and calculate data:
[1946] Initial Setup
[1947] The server verifies the user's consent and collects the necessary data. The specific steps are as follows:
[1948] 1. The device prompts the user to consent to data collection.
[1949] 2. The user consents and the information is returned to the server.
[1950] 3. The server collects and analyzes the user's past activity data, identifies areas of interest, and generates a list of topics.
[1951] Obtaining location information
[1952] The device obtains its current location and sends that information to the server. The specific steps are as follows:
[1953] 1. The device (car navigation system or smartphone) periodically obtains its current location information using GPS.
[1954] 2. The acquired location information is sent to the server.
[1955] Topic selection
[1956] The server selects topics based on interests and location, as follows:
[1957] 1. The server matches the received location information with the user's interest list.
[1958] 2. Pick up topics related to the user's interests and location information and send them to the device.
[1959] Start a conversation
[1960] The device will then interact with the user through voice based on the selected topic. The specific steps are as follows:
[1961] 1. The device uses speech synthesis technology to provide topics to the user.
[1962] 2. The user responds verbally, and the device analyzes the speech.
[1963] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[1964] Cooperation with autonomous driving systems
[1965] The terminal synchronizes the safety instructions and interactions of the automated driving system. The specific steps are as follows:
[1966] 1. The device constantly monitors the voice instructions of the autonomous driving system.
[1967] 2. When an autonomous driving safety instruction is issued, the system pauses interaction with the user.
[1968] 3. When the instruction is completed, the interaction resumes.
[1969] Specific examples
[1970] For example, if a user has been interested in travel in the past and is currently near Tokyo Tower, the server will select a topic about the "history of Tokyo Tower's architecture." The device will then say, "Hello. Tokyo Tower is nearby. Do you know about the architecture of Tokyo Tower?" If the autonomous driving system issues the instruction "Turn right at the next traffic light," the infotainment system will pause and resume the conversation after the instruction is completed.
[1971] Prompt Sentence Examples
[1972] Here are some example prompts to input to a generative AI model:
[1973] "Design an infotainment system for autonomous vehicles that can provide topical information related to the driver's current location and enable voice interaction based on the driver's interests."
[1974] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1975] Step 1:
[1976] The server sends a prompt to the terminal asking the user for consent to data collection.
[1977] Input: The server receives the data needed to generate the consent prompt (e.g., the consent message).
[1978] Data calculation: The server generates a prompt message.
[1979] Output: The generated prompt message is sent to the terminal.
[1980] Step 2:
[1981] The terminal displays the consent prompt received from the server to the user.
[1982] Input: The terminal receives the prompt message.
[1983] Data processing: The terminal displays a prompt message on the user interface.
[1984] Output: User's choice to accept or decline.
[1985] Step 3:
[1986] The user presses the consent button and the information is sent back to the server.
[1987] Input: User's consent or denial of information.
[1988] Data calculation: The terminal analyzes the user's selection and generates consent information.
[1989] Output: The consent information is sent to the server.
[1990] Step 4:
[1991] The server collects data such as the user's past search and purchase history to identify areas of interest.
[1992] Input: The server receives the user's activity data along with consent information.
[1993] Data Calculation: The server collects the user's past activity data from a database and uses analysis algorithms to identify areas of interest.
[1994] Output: A list of areas of interest is generated.
[1995] Step 5:
[1996] The device obtains its current location and sends that information to the server.
[1997] Input: The device receives current location information from the GPS module.
[1998] Data processing: The acquired location information is converted into a format suitable for the server.
[1999] Output: The converted current location information is sent to the server.
[2000] Step 6:
[2001] The server compares the received location information with the interest list, selects appropriate topics, and sends them to the device.
[2002] Input: The server receives the current location and a list of interests.
[2003] Data calculation: The server matches the location information with the list of areas of interest and generates related topics.
[2004] Output: A list of relevant topics is generated and sent to the terminal.
[2005] Step 7:
[2006] The terminal uses speech synthesis technology to speak to the user based on the topic data received from the server.
[2007] Input: The terminal receives topic data.
[2008] Data processing: Using voice synthesis technology, the topic data is converted into a voice message.
[2009] Output: The audio message is spoken to the user.
[2010] Step 8:
[2011] The user responds with voice, which is then analyzed by the device.
[2012] Input: The user's spoken response.
[2013] Data Computing: A speech recognition engine converts the user's response into text and analyzes it.
[2014] Output: The analysis results are sent to the server.
[2015] Step 9:
[2016] The server generates the next topic and response based on the analysis results and sends them back to the terminal.
[2017] Input: The server receives the user's voice analysis results.
[2018] Data calculation: The server generates an appropriate response based on the analysis results.
[2019] Output: The generated response is sent to the terminal.
[2020] Step 10:
[2021] When the autonomous driving system issues safety instructions, the device temporarily stops interacting with the user.
[2022] Input: The terminal receives safety instructions from the automated driving system.
[2023] Data processing: Pause the interaction.
[2024] Output: Voice interaction with the user stops during the safety instruction period.
[2025] Step 11:
[2026] Once the safety instructions from the automated driving system are complete, the device resumes interaction with the user.
[2027] Input: Notification of end of safety instruction.
[2028] Data processing: Restarting a stopped interaction.
[2029] Output: Voice interaction with the user is resumed.
[2030] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2031] System Overview
[2032] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[2033] Program processing overview and specific examples
[2034] Initial Setup
[2035] The server confirms the user's consent and collects the necessary data. The initial setup of this system is as follows:
[2036] 1. The server sends a prompt to the user's device asking for consent to data collection.
[2037] 2. The device displays a consent confirmation screen to the user.
[2038] 3. When the user presses the consent button, the information is sent back to the server via the device.
[2039] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[2040] Example: If a user has made many music-related searches and purchases in the past, tag their interests as music-related and add music topics to the priority list.
[2041] Obtaining location information
[2042] The device obtains location information using GPS and sends that information to the server.
[2043] 1. The device periodically receives GPS signals and obtains its current location information.
[2044] 2. The device sends the acquired location information to the server.
[2045] Example: If your current location is near a tourist attraction, send that information.
[2046] Topic selection
[2047] The server selects topics based on interests and location.
[2048] 1. The server matches the location information with the user's interest list.
[2049] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[2050] Example: If the location is near a historical place and the user is interested in history, select topics about the history of that place.
[2051] Emotion recognition by emotion engine
[2052] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[2053] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[2054] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[2055] Example: If the user is tired or stressed, the emotion engine will recognize that emotional state.
[2056] Adjusting the topic based on emotions
[2057] The server adjusts the topic based on the user's emotional state.
[2058] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[2059] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[2060] For example, if the user is tired, select relaxing music or topics that will help them relax.
[2061] Start a conversation
[2062] The device will interact via voice based on the selected topic.
[2063] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[2064] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[2065] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[2066] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[2067] Linking with car navigation systems
[2068] The terminal monitors and controls the voice instructions of the car navigation device.
[2069] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[2070] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[2071] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[2072] For example: "Let me tell you about the history of this area."
[2073] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[2074] The processing flow will be explained below.
[2075] Step 1:
[2076] The server confirms the user's consent
[2077] 1. The server sends a prompt to the device requesting the user's consent to data collection.
[2078] 2. The device displays a consent confirmation screen to the user.
[2079] 3. When the user presses the consent button, the information is sent back to the server via the device.
[2080] Step 2:
[2081] The server collects and analyzes the user's past activity data.
[2082] 1. Once the server receives the user's consent, it collects usage history data (e.g., search history, purchase history) of the relevant online service.
[2083] 2. The server analyzes the collected data and identifies the user's areas of interest.
[2084] Example: If a user does a lot of movie-related searches and purchases, tag their interests as movie-related.
[2085] 3. The server generates a priority list of interests based on the analysis results.
[2086] Step 3:
[2087] The device acquires location information using GPS
[2088] 1. The device periodically receives GPS signals and obtains its current location information.
[2089] 2. The device sends the acquired location information to the server.
[2090] Example: If your current location is near a tourist attraction, send that information.
[2091] Step 4:
[2092] The server selects topics based on interests and location information
[2093] 1. The server matches the location information with the user's interest list.
[2094] 2. The server selects interesting topics related to the location.
[2095] Example: If the location is near a historical landmark and the user is interested in history, select topics related to the history of that place.
[2096] 3. The server formats the selected topic into a conversational format and sends it to the terminal.
[2097] Step 5:
[2098] The device recognizes the user's emotions
[2099] 1. The device uses an emotion engine to analyze the user's voice and facial expressions.
[2100] 2. The device assesses the user's emotional state (e.g., happy, sad, tired, etc.).
[2101] 3. The device sends the evaluation results to the server.
[2102] Step 6:
[2103] The server adjusts the topic to match the emotional state
[2104] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[2105] 2. The server selects topics according to the user's emotional state and provides conversation content appropriate to the user's emotions.
[2106] For example, if the user is tired, select a topic or music that has a relaxing effect.
[2107] Step 7:
[2108] The device interacts with the user through voice
[2109] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[2110] For example, say, "Hello. There's a historic building nearby. Do you know about it?"
[2111] 2. When the user responds by voice, the device analyzes the voice.
[2112] 3. The device sends the analysis results to the server and requests the next conversation content.
[2113] Step 8:
[2114] The terminal monitors and controls the voice instructions of the car navigation device
[2115] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[2116] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[2117] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[2118] For example: "Let me tell you about the history of this area."
[2119] In this way, the system of the present invention can support a comfortable and safe drive by considering the user's emotional state and interests and providing the user with the most appropriate topics to talk about while driving.
[2120] Example 2
[2121] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2122] It is difficult for drivers to enjoy a comfortable long-distance drive without feeling drowsy. It is also not easy to provide appropriate interactions based on the user's interests, current location, and emotional state. In particular, current technology is not sufficient to synchronize the voice instructions of the car navigation system with the conversation between the user and the car.
[2123] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2124] In this invention, the server includes means for analyzing the user's interests and concerns based on past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for engaging in voice interaction with the user using the selected topics, and means for acquiring the user's emotional data using an emotion engine and adjusting the topics according to the emotional state. This makes it possible to provide appropriate interactions according to the user's interests, location information, and emotional state, prevent drowsiness during long-distance driving, and support a comfortable drive.
[2125] "User" means a person who operates or uses a particular system or device.
[2126] "Past activity" refers to behavioral data such as a user's past search history and purchase history.
[2127] "Means for analyzing interests and concerns" refers to technology or devices for analyzing a user's past activity data to identify what the user is interested in.
[2128] "Means for providing topics" refers to technologies and devices for presenting appropriate topics to users based on the analyzed data.
[2129] "Location Information" refers to a user's current physical location data obtained through technologies such as GPS.
[2130] "Means for obtaining location information" refers to technologies and devices that use GPS or other location measurement technologies to determine a user's current location.
[2131] "Means for selecting a topic" refers to a technology or device for selecting a topic appropriate for a user based on the acquired location information.
[2132] "Means for voice interaction" refers to technology and devices for using voice synthesis technology to have a voice dialogue with a user.
[2133] "Car navigation device" refers to an electronic device that displays the current location of a vehicle and provides directions to a destination.
[2134] The "means for controlling synchronization of voice utterances" refers to a technique or device for appropriately adjusting the voice instructions of the car navigation device and the voice dialogue with the user.
[2135] An "emotion engine" refers to technology or a device that analyzes a user's voice and facial expressions to assess their emotional state.
[2136] "Means for acquiring emotional data" refers to technology or devices for acquiring the user's emotional state in real time using an emotion engine.
[2137] "Means for adjusting the topic" refers to technology or devices for changing or adapting the topic provided to the user based on the acquired emotional data.
[2138] System Overview
[2139] The system of the present invention aims to prevent drivers from becoming drowsy during long-distance driving and to enable them to enjoy a pleasant and comfortable drive. This system has the function of analyzing the user's interests based on their past activities and providing appropriate topics. It also acquires location information and selects topics related to that information. Furthermore, it uses the selected topics to interact with the user via voice and controls the synchronization of voice utterances with the car navigation system, ensuring safe driving. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it is possible to adjust the topics provided according to the user's emotional state.
[2140] Program processing overview and specific examples
[2141] Initial Setup
[2142] The server verifies the user's consent and collects the necessary data.
[2143] 1. The server sends a prompt to the user's device asking for consent to data collection.
[2144] 2. The device displays a consent confirmation screen to the user.
[2145] 3. When the user presses the consent button, the information is sent back to the server via the device.
[2146] 4. The server collects the user's past search history and purchase history, analyzes their areas of interest based on this, and generates a topic list.
[2147] Examples:
[2148] If a user has made many music-related searches or purchases in the past, we tag their interests as music-related and add music topics to the priority list.
[2149] Obtaining location information
[2150] The device obtains location information using GPS and sends that information to the server.
[2151] 1. The device periodically receives GPS signals and obtains its current location information.
[2152] 2. The device sends the acquired location information to the server.
[2153] Examples:
[2154] If the location information is near a tourist attraction, the information is sent.
[2155] Topic selection
[2156] The server selects topics based on interests and location.
[2157] 1. The server matches the location information with the user's interest list.
[2158] 2. The server selects interesting topics related to the location information, formats them into a conversational format, and sends them to the device.
[2159] Examples:
[2160] If the location information is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[2161] Emotion recognition by emotion engine
[2162] The device recognizes the user's emotions and adjusts the topic according to their emotional state.
[2163] 1. The device is equipped with an emotion engine that analyzes the user's voice and facial expressions.
[2164] 2. The device evaluates the user's emotions in real time and sends that information to the server.
[2165] Examples:
[2166] If the user is tired or stressed, the emotion engine will recognize that emotional state.
[2167] Adjusting the topic based on emotions
[2168] The server adjusts the topic based on the user's emotional state.
[2169] 1. The server receives data from the emotion engine and evaluates the user's current emotional state.
[2170] 2. The server selects topics according to the user's emotional state and provides relaxing or interesting topics.
[2171] Examples:
[2172] If the user is tired, select relaxing music or topics that will help them relax.
[2173] Start a conversation
[2174] The device will interact via voice based on the selected topic.
[2175] 1. The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[2176] Examples:
[2177] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[2178] 2. When the user responds by voice, the device analyzes the voice and sends the analysis results to the server.
[2179] 3. The server generates the next topic and response based on the analysis results and sends it back to the terminal.
[2180] Linking with car navigation systems
[2181] The terminal monitors and controls the voice instructions of the car navigation device.
[2182] 1. The terminal constantly monitors the voice instructions of the car navigation device.
[2183] 2. When the car navigation system issues a voice command (e.g., "Turn right at the next traffic light"), the device automatically pauses the conversation with the user.
[2184] 3. When the car navigation system finishes its instructions, the device resumes the conversation with the user.
[2185] Examples:
[2186] "Let me tell you about the history of this area," he resumes.
[2187] Prompt Sentence Examples
[2188] Examples:
[2189] A user is on a long drive and the car navigation system tells them to "turn left at the next traffic light."
[2190] Example prompt for a generative AI model:
[2191] "You're interested in history and are currently driving near a historical landmark. After the navigation system finishes giving you current directions, you can start a conversation about the history of that landmark."
[2192] This system can provide appropriate topics according to the user's emotional state and interests while driving, supporting a safe and comfortable drive.
[2193] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2194] Step 1:
[2195] The server sends a user consent prompt to the terminal.
[2196] Specific behavior:
[2197] The server creates a prompt to obtain the user's consent to data collection and sends the prompt to the terminal.
[2198] Input: Consent prompt template information, User ID.
[2199] Output: The user consent prompt sent to the device.
[2200] Step 2:
[2201] The terminal displays a consent confirmation screen to the user.
[2202] Specific behavior:
[2203] The terminal generates a consent confirmation screen based on the received consent prompt and displays it to the user.
[2204] Input: The consent prompt received from the server.
[2205] Output: The consent screen shown to the user.
[2206] Step 3:
[2207] The user presses the consent button and the information is sent back to the server via the terminal.
[2208] Specific behavior:
[2209] When the user presses the accept button, their input is captured on the device.
[2210] The terminal returns the user's consent information to the server.
[2211] Input: User consent.
[2212] Output: User consent information sent to the server.
[2213] Step 4:
[2214] The server collects the user's past search and purchase history, analyzes their areas of interest, and generates a topic list.
[2215] Specific behavior:
[2216] The server collects the user's past search history and purchase history from a database.
[2217] The server analyzes the collected data to identify the user's areas of interest.
[2218] The server generates a topic list appropriate for the user based on the identified areas of interest.
[2219] Input: User's past search history, purchase history.
[2220] Output: Analyzed areas of interest, generated topic list.
[2221] Examples:
[2222] If a user has made many music-related searches or purchases in the past, the server will use that information to add music-related topics to the list.
[2223] Step 5:
[2224] The device receives a GPS signal and sends the location information to the server.
[2225] Specific behavior:
[2226] The device periodically receives GPS signals to obtain its current location.
[2227] The terminal transmits the acquired location information to the server.
[2228] Input: GPS signal.
[2229] Output: Current location information sent to the server.
[2230] Examples:
[2231] If the location information is near a tourist attraction, the information is sent to the server.
[2232] Step 6:
[2233] The server compares the location information with the user's interest list and selects topics.
[2234] Specific behavior:
[2235] The server matches the received location information with the user's interest list.
[2236] The server selects interesting topics related to the location.
[2237] The topics are formatted into a conversational format and sent to the terminal.
[2238] Input: Location, user interest list.
[2239] Output: Conversational topic data sent to the device.
[2240] Examples:
[2241] If the current location is near a historical place and the user is interested in history, a topic about the history of that place is selected.
[2242] Step 7:
[2243] The terminal recognizes the user's emotions and transmits the emotional state to the server.
[2244] Specific behavior:
[2245] The device acquires emotional data using an emotion engine that analyzes the user's voice and facial expressions.
[2246] The device evaluates the acquired emotion data and transmits the information to the server.
[2247] Input: User's voice and facial expression data.
[2248] Output: Emotion data sent to the server.
[2249] Examples:
[2250] Recognizes the user's emotional state when they are tired or stressed.
[2251] Step 8:
[2252] The server adjusts the topic based on the emotional state and sends it to the terminal.
[2253] Specific behavior:
[2254] The server evaluates the user's current emotional state based on the data received from the emotion engine.
[2255] The server selects a topic according to the emotional state and sends it to the terminal.
[2256] Input: Sentiment data, topic list.
[2257] Output: The adjusted topic data sent to the device.
[2258] Examples:
[2259] If the user is tired, select relaxing music or topics that will help them relax.
[2260] Step 9:
[2261] The device speaks the topic and engages in interaction.
[2262] Specific behavior:
[2263] The device uses speech synthesis technology to speak aloud based on the topic data received from the server.
[2264] When the user responds by voice, the voice is analyzed and the analysis results are sent to the server.
[2265] Input: Topic data from the server, user voice input.
[2266] Output: Audio utterance, analysis results sent to the server.
[2267] Examples:
[2268] Say, "Hello. There is a historic building nearby. Do you know about this place?"
[2269] Step 10:
[2270] The device monitors the car navigation system's voice instructions and controls the interaction.
[2271] Specific behavior:
[2272] The terminal constantly monitors the voice instructions of the car navigation device.
[2273] When the car navigation system issues a voice command, the device automatically pauses the conversation with the user.
[2274] When the car navigation system finishes giving instructions, the terminal resumes conversation with the user.
[2275] Input: Car navigation voice instructions.
[2276] Output: Pause and resume conversation.
[2277] Examples:
[2278] When the car navigation system says, "Turn right at the next traffic light," it pauses the conversation with the user, and then resumes after the instruction is completed, saying, "Let me tell you about the history of this area."
[2279] (Application example 2)
[2280] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2281] Currently, there are not enough systems proposed to prevent drowsiness during long-distance driving and ensure a comfortable and safe drive. Furthermore, no systems have been found that combine a variety of elements, such as providing topics based on the driver's emotional state and interests, selecting topics based on location information, and synchronizing voice output with a car navigation system. The present invention aims to solve these problems and provide a system that prevents drowsiness during long-distance driving and enables drivers to enjoy a fun and comfortable drive.
[2282] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's interests based on their past activities and providing appropriate topics, means for acquiring location information and selecting topics related to that information, means for recognizing the user's emotions and adjusting the topics according to that state, and means for controlling synchronization of voice utterances with a car navigation device. This makes it possible to provide topics according to the interests and concerns of drivers who are driving long distances, select appropriate topics based on location information, adjust the content of the conversation according to their emotional state, and synchronize voice utterances with the navigation device to ensure safe driving.
[2283] "Means for analyzing a user's interests based on their past activities and providing appropriate topics" refers to a device or program that collects data such as a user's past search history and purchase history, analyzes it to identify areas and topics that the user is likely to be interested in, and generates and provides topics based on that data.
[2284] "Means for acquiring location information and selecting topics related to that information" refers to a device or program that identifies the user's current location using a location information acquisition means such as GPS, and selects appropriate topics based on historical background, tourist attractions, local characteristics, etc. related to that location information.
[2285] "Means for recognizing a user's emotions and adjusting the topic according to that state" refers to a device or program that uses voice and facial expression analysis technology to analyze the content of a user's speech and facial expressions in real time, recognizes the user's emotional state (e.g., joy, surprise, fatigue, stress, etc.), and selects the topic that best suits the recognized emotion.
[2286] The "means for controlling synchronization of voice utterances with a car navigation device" refers to a control device or program that monitors voice instructions issued by a car navigation system in real time, automatically stops voice interaction with the user when a navigation instruction is given, and resumes conversation after the instruction is completed.
[2287] "Control means for automatically stopping interaction with the user when the car navigation device gives a voice instruction" refers to a device or program that has the function of detecting when the car navigation system is about to give the next voice instruction and temporarily stopping the conversation with the user during that time.
[2288] The present invention is a system that aims to prevent drivers from getting drowsy during long-distance driving and to allow them to enjoy a fun and comfortable drive. The system is configured using the following hardware and software.
[2289] System Overview
[2290] 1. Server Role
[2291] Analyze and recommend topics based on user's past activities
[2292] The server collects the user's past search history and purchase history, and by analyzing this data, identifies the areas the user is interested in. Based on this data, it generates an appropriate topic list.
[2293] Example: If users are doing a lot of music-related searches and purchases, prioritize music topics in your list.
[2294] Obtain location information and select related topics
[2295] The server receives GPS location information sent from the device and selects topics related to a specific location based on that information.
[2296] For example: If the user is near a tourist attraction, choose a topic related to the history and attractions of that place.
[2297] User Emotion Recognition and Topic Adjustment
[2298] The server receives the user's voice or facial expression analysis data sent from the terminal, evaluates the user's emotional state using an emotion engine, and adjusts the topic based on this evaluation to provide conversation content appropriate to the user's emotional state.
[2299] Example: If the emotion engine detects that the user is tired, it will provide relaxing music or topics.
[2300] 2. Role of the terminal
[2301] Acquisition and transmission of GPS location information
[2302] The device periodically receives GPS signals, acquires its current location information, and sends it to the server.
[2303] Emotion recognition by emotion engine
[2304] The device is equipped with a facial recognition camera and microphone to analyze the user's voice and facial expressions in real time, and the analysis results are sent to a server.
[2305] Conducting a voice interaction
[2306] The device uses speech synthesis technology to converse with the user based on the topic data received from the server. When the user responds by voice, the device analyzes the voice and sends the analysis results back to the server.
[2307] For example, you could say, "Hello. There is a historic building nearby. Do you know about this place?"
[2308] 3. Linkage with car navigation devices
[2309] The terminal works in conjunction with the car navigation system, automatically pausing the conversation with the user when a voice command is given, and resuming the conversation after the command is completed.
[2310] Prompt Sentence Examples
[2311] "There's a historic building nearby. Would you like to know about this place? And if you're interested in music, would you like to tell me about a new hit song?"
[2312] This system can select appropriate topics based on the user's past activity information, current location information, and emotional information, supporting a fun and comfortable drive. By providing interesting topics based on location and emotional information, the user can stay relaxed and alert while driving. This invention achieves more personalized interactions by using a generative AI model and an emotion recognition engine.
[2313] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2314] Step 1:
[2315] Consent verification and user data collection
[2316] Input: User ID
[2317] How it works: When a user launches the app, the server sends a consent prompt to the device.
[2318] When the user presses the consent button, the information is sent back to the server via the terminal.
[2319] Output: User consent information
[2320] Data processing: The server collects and analyzes users' past search history and purchase history.
[2321] Output: User's interests and a list of topics based on them
[2322] Step 2:
[2323] Acquiring and sending location information
[2324] Input: None
[2325] How it works: The device periodically receives GPS signals and obtains its current location.
[2326] Output: Location information
[2327] Data processing: The device sends the acquired location information to the server.
[2328] Output: Location data
[2329] Step 3:
[2330] Topic selection
[2331] In...
Claims
1. A means of analyzing users' interests and concerns based on their past activities and providing them with relevant topics; A means for acquiring location information and selecting topics related to the information; means for vocally interacting with the user using the selected topic; means for controlling synchronization of voice utterance with a car navigation device; A system including:
2. 2. The system according to claim 1, further comprising control means for automatically stopping interaction with the user when the car navigation device issues voice instructions.
3. 2. The system of claim 1, further comprising means for generating topics likely to be of interest to a user by collecting the user's past search history and purchase history and analyzing the data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A