system
The system addresses the lack of personalized support in vehicles by authenticating users, utilizing generative AI for schedule integration and real-time IoT data, enhancing user experience and driving efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
Smart Images

Figure 2026063836000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional information systems provided in vehicles have difficulty providing personalized support for individual users, and the user experience has not been sufficient. Also, real-time data collection and analysis have not been performed, and it has not been possible to provide the driver with the latest and optimal information. As a result, drivers and passengers often felt stressed during the journey to the destination.
Means for Solving the Problems
[0005] The present invention solves the above problems by providing a system comprising means for authenticating the user's identity, means for receiving user commands via speech recognition, means for acquiring the user's schedule information and suggesting a destination using a generative AI, means for collecting IoT data and calculating the optimal route in real time, and means for providing information to the user via speech synthesis. This system authenticates the user's face using a camera and performs identity authentication by comparing the authentication result with the cloud. Furthermore, the speech recognition system converts speech into text data, and the generative AI retrieves the schedule from the cloud calendar to suggest a personalized destination to the user. In addition, the collection of IoT data enables real-time optimal route calculation, allowing the provision of the latest and most optimal information.
[0006] "Personal authentication" is the process of identifying a specific individual and verifying their identity.
[0007] "Speech recognition" is a technology that analyzes speech and converts it into text or commands.
[0008] "Generative AI" is an artificial intelligence technology that uses machine learning models to generate new data and perform analysis and predictions.
[0009] "Schedule information" refers to information about activities and events that the user has planned.
[0010] "IoT data" refers to data collected from objects connected to the internet.
[0011] "Real-time" means that processing and responses are performed instantly based on the current time.
[0012] "Speech synthesis" is a technology that converts text data into speech and artificially generates voice.
[0013] "Cloud" refers to a network of remote servers provided via the internet, offering infrastructure for data storage and processing.
[0014] A "navigation system" is a system that calculates and guides you along the optimal route to your destination.
[0015] "Real-time support" is a function that provides users with information and support immediately.
[0016] "Route calculation" is the process of determining the optimal route from a starting point to a destination.
[0017] "Personalization" means providing experiences and services that are customized to the specific needs and preferences of an individual. [Brief explanation of the drawing]
[0018] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which a plurality of emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0020] First, the terms used in the following description will be described.
[0021] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0022] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is programmed according to the following procedure.
[0040] User authentication
[0041] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[0042] Setting a destination
[0043] The user asks the system by voice, "What's on my schedule today?" This voice command is converted into text data by the terminal's voice recognition system. The text data is sent to a generating AI, and the server retrieves the user's schedule from the cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0044] Navigation and real-time support
[0045] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[0046] Transaction processing
[0047] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[0048] summary
[0049] Through the above processes, users can efficiently navigate their journey to their destination while receiving personalized support. This system utilizes voice recognition, generative AI, and real-time data collection and analysis to provide the latest and most relevant information, ensuring a comfortable user experience for both drivers and passengers.
[0050] The following describes the processing flow.
[0051] Specific processing steps for carrying out the invention
[0052] User authentication
[0053] Step 1:
[0054] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[0055] Step 2:
[0056] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[0057] Step 3:
[0058] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[0059] Step 4:
[0060] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[0061] Step 5:
[0062] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[0063] Setting a destination
[0064] Step 1:
[0065] User: "What's on the schedule for today?" asks the system by voice.
[0066] Step 2:
[0067] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[0068] Step 3:
[0069] Terminal: Sends the converted text data to the generating AI.
[0070] Step 4:
[0071] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[0072] Step 5:
[0073] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[0074] Step 6:
[0075] Terminal: The generated message is read aloud to the user using speech synthesis.
[0076] Step 7:
[0077] User: "Yes, I'm going to the office," they reply, confirming the destination.
[0078] Navigation and real-time support
[0079] Step 1:
[0080] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[0081] Step 2:
[0082] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[0083] Step 3:
[0084] Terminal: Sends collected data to the server in real time.
[0085] Step 4:
[0086] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[0087] Step 5:
[0088] Server: Sends updated routes and additional information to the terminal.
[0089] Step 6:
[0090] Terminal: Notifies the user of updated information using speech synthesis.
[0091] Step 7:
[0092] User: Drive according to instructions based on traffic congestion and accident information.
[0093] Transaction processing
[0094] Step 1:
[0095] User: "Find a nearby parking lot," gives a voice command.
[0096] Step 2:
[0097] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[0098] Step 3:
[0099] Terminal: Sends the converted text data to the server.
[0100] Step 4:
[0101] Server: Searches for the nearest parking lot and checks its availability.
[0102] Step 5:
[0103] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[0104] Step 6:
[0105] Terminal: The suggested message is read aloud to the user using speech synthesis.
[0106] Step 7:
[0107] User: Responds with "Yes" and instructs to reserve a parking space.
[0108] Step 8:
[0109] Server: Executes parking reservations through the system.
[0110] Step 9:
[0111] Server: Sends a reservation completion notification to the device.
[0112] Step 10:
[0113] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[0114] In this way, users can arrive at their destination comfortably and efficiently through a series of processes. This allows them to spend their time in the car more meaningfully.
[0115] (Example 1)
[0116] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0117] Conventional vehicle driver assistance systems are increasingly expected to offer advanced functions beyond user authentication and voice recognition-based command acceptance, such as real-time information provision, schedule management, and parking reservation. However, few systems provide these functions in an integrated manner, resulting in limited improvements to the user experience. Furthermore, real-time data analysis and navigation optimization have been insufficient, making it difficult to provide a safe and efficient driving environment.
[0118] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0119] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands via voice recognition, means for acquiring user schedule information and suggesting destinations using generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user via speech synthesis, means for activating a navigation system and recalculating and providing the route in real time, means for collecting and analyzing speed, fuel level, and location information from vehicle sensors, and means for collecting parking information and executing reservations. This enables an improved user experience and the realization of a safe and efficient driving environment.
[0120] "User authentication" is a process that uses cameras mounted on the vehicle to capture the user's face and compares the encrypted facial data with a database in the cloud.
[0121] "Voice recognition" is a technology that converts the voice spoken by a user inside a vehicle into text data, which the system then accepts as a command it can understand.
[0122] "Generative AI" is an algorithm that uses artificial intelligence technology to analyze a user's schedule information and suggest appropriate destinations and actions.
[0123] "IoT data" refers to real-time data such as speed, fuel level, and location information collected from vehicle sensors.
[0124] "Calculating the optimal route in real time" is a process that instantly analyzes the most efficient travel route based on collected IoT data, external traffic information, and weather information.
[0125] "Speech synthesis" is a technology that converts text data generated by a system into speech and provides information to the user.
[0126] "Activating the navigation system" means activating the vehicle's navigation system to begin providing optimal route guidance to the set destination.
[0127] "Recalculating and providing routes" means re-analyzing the optimal route based on real-time data during travel and providing the user with the latest route information.
[0128] "Vehicle sensors" are devices used to measure various conditions inside and around the vehicle, collecting information such as speed, fuel level, and location.
[0129] "Gathering parking information" is the process of finding available parking spaces and their availability near your current location.
[0130] "Executing a reservation" refers to the process of reserving a suitable parking space for the user based on the collected parking information.
[0131] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is implemented using the following hardware and software:
[0132] hardware
[0133] Camera: Mounted in the vehicle and used to capture the user's face.
[0134] Vehicle sensors: Used to collect speed, fuel level, and location information in real time.
[0135] Terminal: A computer device installed inside a vehicle that runs voice recognition systems, generative AI, and navigation systems.
[0136] Server: Located in a cloud environment, it performs facial recognition data matching, IoT data analysis, and acquires schedule information and calculates routes using generated AI.
[0137] software
[0138] Speech recognition system: Converts user speech into text data.
[0139] Generative AI model: Analyzes user schedule information and suggests appropriate destinations and actions.
[0140] Navigation system: Calculates the optimal route and guides the user.
[0141] Encryption technology: Used to securely transmit user facial data to the cloud.
[0142] This system is implemented in the following steps:
[0143] 1. User authentication:
[0144] The terminal uses a camera mounted on the vehicle to capture the user's face. The authentication process starts automatically when the engine is started. The acquired facial data is encrypted and sent to a server. The server compares the facial data with a database in the cloud and sends the authentication result back to the terminal. The terminal then notifies the user by voice, "Hello, [username]. Where are you going today?"
[0145] 2. Setting the destination:
[0146] The user asks aloud, "What's on my schedule today?" A speech recognition system converts the speech into text data, which is then sent to a generating AI. The server retrieves the user's schedule from the cloud calendar and generates a message saying, "You have a meeting at the office at 9am today. Shall we head to the office?" The device reads this message aloud. When the user responds, "Yes, I'll go to the office," the destination is set.
[0147] 3. Navigation and real-time support:
[0148] The terminal activates the navigation system and begins calculating the optimal route. IoT data such as speed, fuel level, and location information collected from the vehicle's sensors is transmitted to the server in real time. The server analyzes this data and recalculates the optimal route by comparing it with traffic and weather information. The latest route information is sent to the terminal and provided to the user via speech synthesis. For example, instructions such as, "We have calculated the optimal route to avoid traffic congestion. Please turn right," are provided.
[0149] 4. Transaction processing:
[0150] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." This is converted into text by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message, "The nearest parking lot is XX. Would you like to make a reservation?" which the device reads aloud. If the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. Afterward, reservation confirmation information is sent to the device and notified by voice.
[0151] Examples of prompt statements
[0152] The following are specific examples of prompt statements to be input to a generative AI model:
[0153] 1. "Tell me your plans for today."
[0154] 2. "Calculate the optimal route."
[0155] 3. "Please find a nearby parking lot."
[0156] By using these prompts, users can utilize the system more smoothly and intuitively.
[0157] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0158] Processing steps
[0159] Step 1:
[0160] Input: User's face data (image)
[0161] Specific operation: The device uses the vehicle's onboard camera to automatically capture the user's face when the engine starts. This image data is then stored in the device.
[0162] Data processing: Acquired facial data is encrypted on the device.
[0163] Output: Encrypted facial data is generated.
[0164] Step 2:
[0165] Input: Encrypted facial data
[0166] Specific operation: The device sends encrypted facial data to the server via a secure communication channel.
[0167] Data processing: The server compares the received facial data with user information stored in a database on the cloud.
[0168] Output: An authentication result (success or failure) is generated.
[0169] Step 3:
[0170] Input: Authentication result
[0171] Specific action: The server sends the authentication result back to the terminal.
[0172] Data processing: If successful, corresponding user information will be attached.
[0173] Output: The authentication result (and user information) is sent to the terminal.
[0174] Step 4:
[0175] Input: Authentication result
[0176] Specific operation: The device receives the authentication result, and if it contains user information, it notifies the user via voice, "Hello, [username]. Where are you going today?"
[0177] Data processing: Convert text to speech using speech synthesis technology.
[0178] Output: An audio notification is generated for the user.
[0179] Step 5:
[0180] Input: User voice input ("What are my plans for today?")
[0181] Specific action: The user asks aloud, "What's on the schedule for today?"
[0182] Data processing: The speech recognition system converts the audio into text data.
[0183] Output: Text data ("What are your plans for today?") is generated.
[0184] Step 6:
[0185] Input: Text data ("What are your plans for today?")
[0186] Specific action: This text data is sent to the generating AI.
[0187] Data processing: The server uses generative AI to retrieve user appointments from cloud calendars and scheduling information.
[0188] Output: User schedule information is generated.
[0189] Step 7:
[0190] Input: User's schedule information
[0191] Specific operation: Based on the acquired information, the server generates a suggestion message such as, "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[0192] Data calculation: Text message generation
[0193] Output: A suggestion message is generated.
[0194] Step 8:
[0195] Input: Suggestion message
[0196] Specific action: The terminal reads out the suggestion message.
[0197] Data processing: Text is converted to speech using speech synthesis technology.
[0198] Output: An audio notification is generated for the user.
[0199] Step 9:
[0200] Input: User voice input ("Yes, I'm going to the office.")
[0201] Specific action: The user responds, "Yes, I will go to the office."
[0202] Data processing: The speech recognition system converts the audio into text data.
[0203] Output: Text data ("Yes, I will go to the office") is generated.
[0204] Step 10:
[0205] Input: Text data ("Yes, I will go to the office")
[0206] Specific action: The terminal receives this text data and sets the office as the destination.
[0207] Data calculation: Setting destination information
[0208] Output: Destination setting complete.
[0209] Step 11:
[0210] Input: Destination information
[0211] Specific action: The terminal activates the navigation system and begins calculating the optimal route.
[0212] Data calculation: Calculation of the optimal route
[0213] Output: A navigation route is generated.
[0214] Step 12:
[0215] Input: Vehicle sensor data (speed, fuel level, location information)
[0216] Specific operation: The terminal collects data from the vehicle's sensors in real time and sends it to the server.
[0217] Data processing: The server analyzes the collected IoT data and compares it with traffic and weather information.
[0218] Output: Analysis results are generated.
[0219] Step 13:
[0220] Input: Analysis results
[0221] Specific operation: The server recalculates the optimal route based on the analysis results and sends it to the terminal.
[0222] Data calculation: Recalculating navigation routes
[0223] Output: The recalculated navigation route is sent to the terminal.
[0224] Step 14:
[0225] Input: Recalculated navigation route
[0226] Specific operation: The terminal provides the user with the latest recalculated route information through speech synthesis.
[0227] Data processing: Text is converted to speech using speech synthesis technology.
[0228] Output: An audio notification is generated for the user.
[0229] Step 15:
[0230] Input: User voice input ("Find a nearby parking lot")
[0231] Specific action: The user reaches near their destination and gives a voice command saying, "Find a nearby parking lot."
[0232] Data processing: The speech recognition system converts the audio into text data.
[0233] Output: Text data ("Find a nearby parking lot") is generated.
[0234] Step 16:
[0235] Input: Text data ("Find nearby parking")
[0236] Specific operation: The server receives text data and collects information about the nearest parking lot.
[0237] Data processing: The server checks the availability of parking spaces and selects a suitable one.
[0238] Output: A message recommending parking is generated.
[0239] Step 17:
[0240] Input: Recommended parking message
[0241] Specific action: The device reads out a recommendation message and notifies the user, "The nearest parking lot is XX. Would you like to make a reservation?"
[0242] Data processing: Text is converted to speech using speech synthesis technology.
[0243] Output: An audio notification is generated for the user.
[0244] Step 18:
[0245] Input: User voice input ("Yes")
[0246] Specific action: When the user responds with "yes," the server accesses the parking reservation system and completes the reservation.
[0247] Data processing: Generating confirmation information for parking reservations
[0248] Output: Reservation confirmation information is generated and sent to the terminal.
[0249] Step 19:
[0250] Input: Reservation confirmation information
[0251] Specific operation: The device receives reservation confirmation information and notifies the user via voice.
[0252] Data processing: Text is converted to speech using speech synthesis technology.
[0253] Output: An audio notification is generated for the user.
[0254] Through these steps, users can travel to their destination comfortably and efficiently while receiving personalized support. The system as a whole utilizes speech recognition, generative AI, and real-time data collection and analysis technologies.
[0255] (Application Example 1)
[0256] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0257] Current autonomous vehicle systems face numerous challenges in improving the user experience, particularly in areas such as user authentication, destination setting, and real-time route optimization. For example, issues include difficulties with smooth facial recognition and cumbersome destination setting. Furthermore, a lack of proper integration of real-time optimal route calculations during driving and parking reservation information near the destination significantly reduces user convenience. Addressing these challenges requires advanced speech recognition, generative AI, and the integration of IoT data.
[0258] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0259] In this invention, the server includes means for capturing the user's face and performing facial recognition, means for transmitting real-time vehicle sensor data to the server and analyzing it, and means for collecting parking information near the destination, checking availability, and proposing and completing reservations. This makes it possible to smoothly perform everything from user authentication to destination setting, navigation, real-time route optimization, and parking reservation.
[0260] "User authentication" refers to the authentication process used when a user logs into a system, utilizing personal characteristics such as their face or voice.
[0261] "Speech recognition" is a technology that allows a machine to understand what a user says and convert it into text data.
[0262] "Generative AI" is an artificial intelligence technology that generates new data or answers based on pre-trained data.
[0263] "IoT data" refers to real-time information collected from various sensors and devices.
[0264] "Calculating the optimal route" means determining the most efficient route to the destination, taking into account current traffic conditions and road congestion.
[0265] "Speech synthesis" is a technology that converts text data into speech and provides information to the user as audio.
[0266] A "smartphone" is an electronic device that combines the functions of a mobile phone and a computer, enabling internet connectivity and application usage.
[0267] "Vehicle sensor data" refers to data such as speed, fuel level, and location information measured by sensors installed in the vehicle.
[0268] "Parking information" refers to information such as the location, availability, and fees of the parking lot.
[0269] A "reservation" is a procedure aimed at securing goods or services in advance.
[0270] "Authentication" refers to the process of verifying that someone is a person.
[0271] A "proposal" is an expression of an idea or plan regarding a particular subject.
[0272] "Analysis" refers to the process of analyzing collected data in detail.
[0273] This invention is a system that provides personalized support for autonomous vehicles, including user facial recognition, voice recognition, acquisition of schedule information using generative AI, real-time optimal route calculation, and parking reservation.
[0274] Hardware and software configuration
[0275] Hardware: Smartphones, cameras, in-vehicle sensors (speed sensors, fuel sensors, GPS, etc.)
[0276] Software: OpenCV, face_recognition, speech_recognition, cloud server, generative AI model
[0277] This system uses the user's smartphone camera to perform facial recognition before the user enters the vehicle. The facial data captured by the camera is encrypted and sent to a cloud server for facial recognition. If facial recognition is successful, the server retrieves the corresponding user's information and sends it back to the device.
[0278] Users set destinations and give other instructions using voice commands. These voice commands are received through the smartphone's microphone and converted into text data by speech_recognition software. The text data is sent to a generative AI model, which suggests appropriate destinations based on the user's schedule information. The generative AI model retrieves the user's schedule from sources such as cloud calendars and generates the optimal route and action plan.
[0279] Once a destination is set, data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors and sent to a cloud server. Based on the collected data, the server calculates the optimal route in real time and sends the result back to the terminal. The terminal uses speech synthesis to provide the user with the latest route information in real time. For example, it may provide specific instructions such as, "We have calculated the optimal route to avoid traffic. Please turn right."
[0280] When the user arrives near the destination, the terminal is instructed verbally to "search for nearby parking lots". This instruction is converted into text data by the speech recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects an appropriate parking lot and generates a proposal message that says, "The nearest parking lot is ○○. Do you want to make a reservation?" If the user responds with "yes", the server accesses the parking lot reservation system to complete the reservation. After that, the reservation confirmation information is sent to the terminal and notified to the user verbally.
[0281] Specific examples of prompt sentences
[0282] "Please propose a destination based on the user's schedule."
[0283] "Please calculate the optimal route to avoid traffic congestion."
[0284] "Please provide the availability and reservation information of nearby parking lots."
[0285] By operating this system, the user can enjoy an efficient and safe drive while receiving individually personalized support.
[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0287] Step 1:
[0288] The user approaches the vehicle using a smartphone. The terminal (smartphone) activates the camera and captures the user's face. This image data is processed using a face authentication framework (e.g., OpenCV and face_recognition). Facial feature points are extracted by face recognition and then encrypted. The encrypted data is sent to the cloud server.
[0289] Input: Camera image of smartphone
[0290] Data processing: Extraction and encryption of facial feature points.
[0291] Output: Encrypted facial data
[0292] Step 2:
[0293] The server authenticates the user by comparing the received encrypted facial data with a database in the cloud. If authentication is successful, the user's individual information is sent back from the server to the device. This completes the user authentication process.
[0294] Input: Encrypted facial data
[0295] Data processing: Database matching
[0296] Output: Authentication results and user information
[0297] Step 3:
[0298] The user gives voice commands to the device. The device receives the user's voice via its microphone and converts it into text data using the speech_recognition library. This text data is then sent to a cloud server.
[0299] Input: Audio data
[0300] Data processing: Speech-to-text conversion
[0301] Output: Text data of voice commands
[0302] Step 4:
[0303] The server analyzes the received text data based on a generative AI model. The generative AI model retrieves the user's schedule information from a cloud calendar and other sources, and suggests destinations. These suggestions are sent to the terminal and notified to the user via speech synthesis.
[0304] Input: Text data of voice commands
[0305] Data processing: Obtaining a schedule and proposing a destination by a generative AI model
[0306] Output: Proposed destination message
[0307] Step 5:
[0308] The user accepts the proposed destination by voice. For example, respond by voice with "Yes, I will go to the office". This response is converted back to text data by speech_recognition and sent to the server.
[0309] Input: Voice response
[0310] Data processing: Text conversion of voice
[0311] Output: Text data of the response
[0312] Step 6:
[0313] The server analyzes the received text data and sets the destination. Additionally, real-time data (speed, fuel level, location information) from the vehicle sensors is collected and sent to the server. The server calculates the optimal route based on this data and sends the result to the terminal.
[0314] Input: Text data of the response and sensor data
[0315] Data processing: Calculation of the optimal route
[0316] Output: Optimal route information
[0317] Step 7:
[0318] The terminal uses speech synthesis to guide the user to the optimal route. For example, it might provide instructions such as, "The optimal route has been set. Please turn right at the next intersection."
[0319] Input: Optimal route information
[0320] Data processing: Speech synthesis
[0321] Output: Voice guidance
[0322] Step 8:
[0323] When the user approaches their destination, they will voice-instruct the device to "find a nearby parking lot." This instruction is also converted into text data via speech_recognition and sent to the server.
[0324] Input: Audio data for parking instructions
[0325] Data processing: Speech-to-text conversion
[0326] Output: Text data of parking instructions
[0327] Step 9:
[0328] The server analyzes the text data of the parking instructions and collects information about nearby parking lots. It checks availability and sends a suggestion message to the terminal indicating a suitable parking lot. A message such as "The nearest parking lot is XX. Would you like to make a reservation?" is generated.
[0329] Input: Text data for parking instructions
[0330] Data processing: Obtaining parking information and checking availability.
[0331] Output: Parking suggestion message
[0332] Step 10:
[0333] When the user responds with "yes," the terminal converts the response into text using its speech recognition system and sends it to the server. The server accesses the parking reservation system and completes the reservation. Reservation confirmation information is sent to the terminal and notified to the user via voice.
[0334] Input: Audio data of parking reservation response
[0335] Data processing: Text conversion and reservation processing of responses
[0336] Output: Reservation confirmation information
[0337] Through the processing steps described above, this system provides consistent support from user facial recognition to parking reservation, offering a safe and efficient driving environment.
[0338] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0339] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[0340] User authentication
[0341] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[0342] Setting a destination
[0343] The user asks the system by voice, "What's on my schedule today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to a generating AI, which retrieves the user's schedule from a cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0344] Navigation and real-time support
[0345] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[0346] How the emotion engine works
[0347] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[0348] Transaction processing
[0349] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[0350] summary
[0351] Through the above processes, users can enjoy an efficient and comfortable journey to their destination while receiving personalized support. This system can further enhance the user experience by recognizing and appropriately responding to emotional states, in addition to voice recognition, generative AI, and real-time data collection and analysis.
[0352] The following describes the processing flow.
[0353] Specific processing steps for carrying out the invention
[0354] User authentication
[0355] Step 1:
[0356] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[0357] Step 2:
[0358] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[0359] Step 3:
[0360] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[0361] Step 4:
[0362] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[0363] Step 5:
[0364] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[0365] Setting a destination
[0366] Step 1:
[0367] User: "What's on the schedule for today?" asks the system by voice.
[0368] Step 2:
[0369] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[0370] Step 3:
[0371] Terminal: Sends the converted text data to the generating AI.
[0372] Step 4:
[0373] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[0374] Step 5:
[0375] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[0376] Step 6:
[0377] Terminal: The generated message is read aloud to the user using speech synthesis.
[0378] Step 7:
[0379] User: "Yes, I'm going to the office," they reply, confirming the destination.
[0380] Navigation and real-time support
[0381] Step 1:
[0382] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[0383] Step 2:
[0384] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[0385] Step 3:
[0386] Terminal: Sends collected data to the server in real time.
[0387] Step 4:
[0388] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[0389] Step 5:
[0390] Server: Sends updated routes and additional information to the terminal.
[0391] Step 6:
[0392] Terminal: Notifies the user of updated information using speech synthesis.
[0393] Step 7:
[0394] User: Drive according to instructions based on traffic congestion and accident information.
[0395] How the emotion engine works
[0396] Step 1:
[0397] Terminal: Uses the vehicle's camera and microphone to capture the user's facial expressions and voice tone in real time.
[0398] Step 2:
[0399] Terminal: Sends captured data to the emotion engine.
[0400] Step 3:
[0401] Server: The emotion engine analyzes the data and determines the user's emotional state.
[0402] Step 4:
[0403] Server: Based on the user's emotional state, the generating AI proposes appropriate countermeasures.
[0404] Step 5:
[0405] Device: The generating AI reads aloud messages and suggestions tailored to the user's emotions using speech synthesis. For example, if the user is feeling stressed, it might suggest, "Would you like to play some music to relax?"
[0406] Transaction processing
[0407] Step 1:
[0408] User: "Find a nearby parking lot," gives a voice command.
[0409] Step 2:
[0410] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[0411] Step 3:
[0412] Terminal: Sends the converted text data to the server.
[0413] Step 4:
[0414] Server: Searches for the nearest parking lot and checks its availability.
[0415] Step 5:
[0416] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[0417] Step 6:
[0418] Terminal: The suggested message is read aloud to the user using speech synthesis.
[0419] Step 7:
[0420] User: Responds with "Yes" and instructs to reserve a parking space.
[0421] Step 8:
[0422] Server: Executes parking reservations through the system.
[0423] Step 9:
[0424] Server: Sends a reservation completion notification to the device.
[0425] Step 10:
[0426] Terminal: Notifies the user via voice when the reservation is complete and displays the reservation information.
[0427] Through the processing flow described above, users can enjoy a comfortable driving environment that takes their emotional state into consideration, while also receiving individually personalized support.
[0428] (Example 2)
[0429] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0430] Conventional driver assistance systems typically only offer navigation and voice recognition functions, lacking features to address the individual needs and emotional states of users. As a result, while safe and efficient driver assistance may be achievable, providing emotional satisfaction and a personalized experience for users remains a challenge.
[0431] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0432] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, means using an emotion engine to analyze the user's emotional state, and means for the generative AI to make suggestions based on the emotional state. This makes it possible to respond to the user's individual needs and improve emotional satisfaction.
[0433] "Means of authenticating a user's identity" refers to a function that uses sensors such as cameras to acquire the user's biometric information and verifies the user's identity based on that information.
[0434] "A means of receiving user commands through voice recognition" refers to a technology that uses acoustic devices such as microphones to capture user voice instructions, analyzes them, and converts them into text data or commands.
[0435] "A method for acquiring user schedule information and suggesting destinations using generative AI" refers to a system that uses artificial intelligence to access user schedule data and suggests the optimal destination based on that data.
[0436] "A method for collecting IoT data and calculating the optimal route in real time" refers to a technology that analyzes traffic conditions and road information based on data collected from various sensors in a vehicle, and calculates the optimal route in real time.
[0437] "A means of providing information to users through speech synthesis" refers to a technology that converts text data into speech output and notifies users of necessary information via voice.
[0438] "Methods using an emotion engine to analyze the user's emotional state" refers to technologies that analyze the user's voice and facial expressions to estimate their emotional state and take appropriate action.
[0439] "A means by which AI generates suggestions based on emotional state" refers to a function in which the generating AI takes into account the user's emotional state and provides the user with appropriate responses and suggestions.
[0440] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[0441] First, as a means of user authentication, the terminal uses a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares it with user information stored in the cloud and performs authentication. If authentication is successful, the server sends the authentication result, including the corresponding user's individual information, back to the terminal, and the terminal notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[0442] Next, regarding setting the destination, the user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?", which the terminal reads aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0443] Once a destination is set, the terminal activates its navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, upon receiving traffic congestion information, it might instruct the user, "We have calculated the optimal route to avoid congestion. Please turn right."
[0444] Regarding the operation of the emotion engine, the agent system collects the user's voice and facial expressions using the camera and microphone and sends them to the emotion engine. The emotion engine analyzes this data to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to make suggestions. Specifically, if the user is expressing stress, the device will make suggestions via voice, such as, "Shall I play some music to help you relax?" or "Shall I find a nearby resting place?"
[0445] Regarding transaction processing, when a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The terminal reads this message aloud. If the user responds with "yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and notified by voice.
[0446] As a concrete example of its operation, when a user gets into a car and asks, "What are my plans for today?", the system suggests, "I have a meeting at the office at 9:00," and navigation begins. Also, when near the destination, if the user says, "Find a nearby parking lot," the terminal converts the voice into text data, and the server searches for and reserves the nearest parking lot.
[0447] As an example of a prompt statement,
[0448] 1. "Please tell me about your appointments and plans for today."
[0449] 2. "Where are you planning to go today?"
[0450] There is.
[0451] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0452] Step 1:
[0453] The device uses a camera mounted on the vehicle to capture the user's face. When the user starts the engine, facial recognition begins automatically. The input data is the facial image captured by the camera, which is processed by a facial recognition algorithm to generate encrypted facial data. The output is the encrypted facial data.
[0454] Step 2:
[0455] The device sends encrypted facial data to the server. The input is encrypted facial data, and the output is the data sent to the server. Specifically, it uses an encrypted communication protocol to send the data.
[0456] Step 3:
[0457] The server performs authentication by comparing the received encrypted data with user information stored in the cloud. The input consists of encrypted facial data and facial data stored in the cloud, and a match is verified through a database search operation. The output is the authentication result.
[0458] Step 4:
[0459] If authentication is successful, the server returns an authentication result to the terminal, which includes the individual user information of the corresponding user. The input consists of the authentication result and user information, which are encrypted and sent to the terminal. The output is the authentication information received by the terminal.
[0460] Step 5:
[0461] Upon successful authentication, the device uses speech synthesis to notify the user, "Hello, Mr. / Ms. Yamada. Where are you going today?" The input is authentication information, which is converted from text to speech via a speech synthesis engine. The output is a voice notification.
[0462] Step 6:
[0463] The user asks the system, "What's on the schedule today?" using voice. The input is the user's voice command, which the terminal collects. The output is voice data.
[0464] Step 7:
[0465] The terminal analyzes voice commands using a speech recognition system and converts them into text data. The input is voice data, and the speech recognition engine converts the voice to text. The output is text data.
[0466] Step 8:
[0467] The device sends text data to the generating AI, and the server retrieves the user's schedule from a cloud calendar or other scheduling information. The input is text data, which is fed into the generating AI model. The output is the user's schedule information.
[0468] Step 9:
[0469] The server generates a suggestion message based on the acquired information: "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is schedule information, and a generation AI is used to generate the suggestion message. The output is the suggestion message.
[0470] Step 10:
[0471] The device notifies the user of this suggestion message via voice. The input is the suggestion message, which is converted from text to speech using a speech synthesis engine. The output is a voice notification.
[0472] Step 11:
[0473] The user responds by voice, "Yes, I'm going to the office," which sets the destination. The input is the user's voice response, which the terminal collects. The output is voice data.
[0474] Step 12:
[0475] The terminal uses a speech recognition system to convert the user's voice responses into text data and confirm the destination setting. The input is voice data, which is converted to text using the speech recognition engine. The output is text data.
[0476] Step 13:
[0477] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. The input is destination information, and the system uses a route calculation algorithm to generate the best route. The output is the optimal route information.
[0478] Step 14:
[0479] The terminal collects IoT data such as speed, fuel level, and location information from vehicle sensors in real time. The input is data from vehicle sensors, which is collected and analyzed. The output is real-time data.
[0480] Step 15:
[0481] The collected IoT data is sent to the server. The input is real-time data, and the output is data transmission to the server.
[0482] Step 16:
[0483] The server analyzes IoT data by cross-referencing it with traffic and weather information, and recalculates the optimal route. Inputs are real-time data and external information, and the optimal route is generated through data analysis. The output is the recalculated route information.
[0484] Step 17:
[0485] This function sends the latest route information to the terminal. The input is the recalculated route information, and the output is the data transmission to the terminal.
[0486] Step 18:
[0487] The device provides the user with the latest route information through speech synthesis. The input is the latest route information, which is converted into speech using the speech synthesis engine. The output is a voice notification.
[0488] (Application Example 2)
[0489] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0490] While conventional autonomous driving systems supported user authentication and destination setting, they did not provide driving support based on the user's emotional state. Furthermore, although they offered real-time optimal route calculations and voice-based information, they did not address the user's emotional needs. As a result, while they provided a safe and efficient driving environment, they were insufficient to fully deliver a comfortable driving experience for the user.
[0491] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for authenticating the user's personal information, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generation AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, and means for recognizing the user's emotional state using an emotion engine and providing support based on this. This makes it possible to respond to the user's emotional needs and to provide an even more comfortable and personalized driving experience.
[0492] "User personal authentication" is a method of recognizing the user's face using a camera installed in the vehicle and matching it with data in the cloud.
[0493] "Speech recognition" is a method of converting a user's voice into text data, allowing the system to accept user commands.
[0494] "Generative AI" is a method that, based on the user's voice instructions, retrieves schedule information from cloud calendars and other sources and suggests destinations.
[0495] "IoT data" refers to data such as speed, fuel level, and location information collected from vehicle sensors.
[0496] "A method for calculating the optimal route in real time" refers to a method that uses IoT data to calculate the optimal route by comparing it with traffic and weather information.
[0497] "Speech synthesis" is a method of providing users with calculation results or suggested messages in voice.
[0498] An "emotion engine" is a method of recognizing a user's emotional state by analyzing their voice and facial expressions.
[0499] "Means of providing support" refers to methods of proposing solutions that meet the user's emotional needs based on the analysis results of the emotion engine.
[0500] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This system not only realizes a safe and efficient driving environment but also addresses the user's emotional needs, providing a more comfortable and personalized driving experience.
[0501] User authentication
[0502] The server captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the cloud server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the vehicle terminal. The terminal notifies the user by voice, "Hello, user. Where are you going today?"
[0503] Setting a destination
[0504] The user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the vehicle terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9:00 today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0505] Navigation and real-time support
[0506] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[0507] How the emotion engine works
[0508] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[0509] Parking reservation support
[0510] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[0511] Examples of specific cases and prompt statements
[0512] For example, if a user asks, "What's on my schedule today?", the system retrieves information from a cloud calendar and suggests, "You have a meeting at the office at 9:00." Also, if the user is feeling stressed, the system might suggest, "Would you like some relaxing music?"
[0513] Example of a prompt:
[0514] What are your plans for today? Please tell me your schedule.
[0515]
[0516] Please suggest ways to address situations where users are experiencing stress.
[0517] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0518] Step 1:
[0519] The camera is activated and the user's face is captured. For authentication, the facial image acquired from the camera is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. The input is facial image data, and the output is the authentication result. Specifically, facial recognition is performed using OpenCV, and if authentication is successful, the corresponding user information is sent back to the terminal.
[0520] Step 2:
[0521] The device notifies the user via speech synthesis, "Hello, user. Where are you going today?" The user's voice response is captured by the microphone and converted into text data using a speech recognition system. This converted text data is the input, and the speech recognition result is the output. Specifically, the Google® Cloud Speech-to-Text API is used to convert speech to text.
[0522] Step 3:
[0523] When a user asks "What's on my schedule today?", a generative AI retrieves the user's schedule information. The generative AI obtains the user's schedule from cloud calendars and other schedule information and generates destination suggestions. The input to this process is the user's voice command "What's on my schedule today?", and the output is destination suggestions. Specifically, it uses the OpenAI® GPT model to obtain schedule information based on the prompt and generates suggestion messages.
[0524] Step 4:
[0525] The device synthesizes suggested messages into speech and notifies the user. For example, it might read aloud the message, "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is a suggested message from the generation AI, and the output is a synthesized speech notification.
[0526] Step 5:
[0527] Once the user confirms the destination setting, the navigation system activates and begins calculating the optimal route. The device collects IoT data such as speed, fuel level, and location information in real time from sensors and sends it to the server. The server compares this data with traffic and weather information and recalculates the optimal route. The input is IoT data and real-time traffic and weather information, and the output is instructions for the optimal route.
[0528] Step 6:
[0529] The device uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotional state. The input is the user's voice and facial expression data, and the output is the result of the emotional state analysis. Specifically, it uses voice and video analysis technologies to analyze the user's emotions.
[0530] Step 7:
[0531] The server uses generative AI to suggest countermeasures based on the analysis results from the emotion engine. For example, if the user is feeling stressed, it might suggest, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?" The input is the analysis results from the emotion engine, and the output is the suggested countermeasures. Specifically, the generative AI makes suggestions based on the user's emotional needs.
[0532] Step 8:
[0533] When a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This is converted into text data by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It then selects a suitable parking lot and prompts the user to make a reservation. The input is the user's voice command, and the output is parking availability information and reservation confirmation.
[0534] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0535] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0536] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0537] [Second Embodiment]
[0538] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0539] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0540] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0541] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0542] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0543] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0544] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0545] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0546] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0547] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0548] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0549] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0550] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is programmed according to the following procedure.
[0551] User authentication
[0552] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[0553] Setting a destination
[0554] The user asks the system by voice, "What's on my schedule today?" This voice command is converted into text data by the terminal's voice recognition system. The text data is sent to a generating AI, and the server retrieves the user's schedule from the cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0555] Navigation and real-time support
[0556] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[0557] Transaction processing
[0558] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[0559] summary
[0560] Through the above processes, users can efficiently navigate their journey to their destination while receiving personalized support. This system utilizes voice recognition, generative AI, and real-time data collection and analysis to provide the latest and most relevant information, ensuring a comfortable user experience for both drivers and passengers.
[0561] The following describes the processing flow.
[0562] Specific processing steps for carrying out the invention
[0563] User authentication
[0564] Step 1:
[0565] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[0566] Step 2:
[0567] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[0568] Step 3:
[0569] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[0570] Step 4:
[0571] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[0572] Step 5:
[0573] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[0574] Setting a destination
[0575] Step 1:
[0576] User: "What's on the schedule for today?" asks the system by voice.
[0577] Step 2:
[0578] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[0579] Step 3:
[0580] Terminal: Sends the converted text data to the generating AI.
[0581] Step 4:
[0582] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[0583] Step 5:
[0584] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[0585] Step 6:
[0586] Terminal: The generated message is read aloud to the user using speech synthesis.
[0587] Step 7:
[0588] User: "Yes, I'm going to the office," they reply, confirming the destination.
[0589] Navigation and real-time support
[0590] Step 1:
[0591] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[0592] Step 2:
[0593] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[0594] Step 3:
[0595] Terminal: Sends collected data to the server in real time.
[0596] Step 4:
[0597] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[0598] Step 5:
[0599] Server: Sends updated routes and additional information to the terminal.
[0600] Step 6:
[0601] Terminal: Notifies the user of updated information using speech synthesis.
[0602] Step 7:
[0603] User: Drive according to instructions based on traffic congestion and accident information.
[0604] Transaction processing
[0605] Step 1:
[0606] User: "Find a nearby parking lot," gives a voice command.
[0607] Step 2:
[0608] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[0609] Step 3:
[0610] Terminal: Sends the converted text data to the server.
[0611] Step 4:
[0612] Server: Searches for the nearest parking lot and checks its availability.
[0613] Step 5:
[0614] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[0615] Step 6:
[0616] Terminal: The suggested message is read aloud to the user using speech synthesis.
[0617] Step 7:
[0618] User: Responds with "Yes" and instructs to reserve a parking space.
[0619] Step 8:
[0620] Server: Executes parking reservations through the system.
[0621] Step 9:
[0622] Server: Sends a reservation completion notification to the device.
[0623] Step 10:
[0624] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[0625] In this way, users can arrive at their destination comfortably and efficiently through a series of processes. This allows them to spend their time in the car more meaningfully.
[0626] (Example 1)
[0627] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0628] Conventional vehicle driver assistance systems are increasingly expected to offer advanced functions beyond user authentication and voice recognition-based command acceptance, such as real-time information provision, schedule management, and parking reservation. However, few systems provide these functions in an integrated manner, resulting in limited improvements to the user experience. Furthermore, real-time data analysis and navigation optimization have been insufficient, making it difficult to provide a safe and efficient driving environment.
[0629] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0630] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands via voice recognition, means for acquiring user schedule information and suggesting destinations using generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user via speech synthesis, means for activating a navigation system and recalculating and providing the route in real time, means for collecting and analyzing speed, fuel level, and location information from vehicle sensors, and means for collecting parking information and executing reservations. This enables an improved user experience and the realization of a safe and efficient driving environment.
[0631] "User authentication" is a process that uses cameras mounted on the vehicle to capture the user's face and compares the encrypted facial data with a database in the cloud.
[0632] "Voice recognition" is a technology that converts the voice spoken by a user inside a vehicle into text data, which the system then accepts as a command it can understand.
[0633] "Generative AI" is an algorithm that uses artificial intelligence technology to analyze a user's schedule information and suggest appropriate destinations and actions.
[0634] "IoT data" refers to real-time data such as speed, fuel level, and location information collected from vehicle sensors.
[0635] "Calculating the optimal route in real time" is a process that instantly analyzes the most efficient travel route based on collected IoT data, external traffic information, and weather information.
[0636] "Speech synthesis" is a technology that converts text data generated by a system into speech and provides information to the user.
[0637] "Activating the navigation system" means activating the vehicle's navigation system to begin providing optimal route guidance to the set destination.
[0638] "Recalculating and providing routes" means re-analyzing the optimal route based on real-time data during travel and providing the user with the latest route information.
[0639] "Vehicle sensors" are devices used to measure various conditions inside and around the vehicle, collecting information such as speed, fuel level, and location.
[0640] "Gathering parking information" is the process of finding available parking spaces and their availability near your current location.
[0641] "Executing a reservation" refers to the process of reserving a suitable parking space for the user based on the collected parking information.
[0642] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is implemented using the following hardware and software:
[0643] hardware
[0644] Camera: Mounted in the vehicle and used to capture the user's face.
[0645] Vehicle sensors: Used to collect speed, fuel level, and location information in real time.
[0646] Terminal: A computer device installed inside a vehicle that runs voice recognition systems, generative AI, and navigation systems.
[0647] Server: Located in a cloud environment, it performs facial recognition data matching, IoT data analysis, and acquires schedule information and calculates routes using generated AI.
[0648] software
[0649] Speech recognition system: Converts user speech into text data.
[0650] Generative AI model: Analyzes user schedule information and suggests appropriate destinations and actions.
[0651] Navigation system: Calculates the optimal route and guides the user.
[0652] Encryption technology: Used to securely transmit user facial data to the cloud.
[0653] This system is implemented in the following steps:
[0654] 1. User authentication:
[0655] The terminal uses a camera mounted on the vehicle to capture the user's face. The authentication process starts automatically when the engine is started. The acquired facial data is encrypted and sent to a server. The server compares the facial data with a database in the cloud and sends the authentication result back to the terminal. The terminal then notifies the user by voice, "Hello, [username]. Where are you going today?"
[0656] 2. Setting the destination:
[0657] The user asks aloud, "What's on my schedule today?" A speech recognition system converts the speech into text data, which is then sent to a generating AI. The server retrieves the user's schedule from the cloud calendar and generates a message saying, "You have a meeting at the office at 9am today. Shall we head to the office?" The device reads this message aloud. When the user responds, "Yes, I'll go to the office," the destination is set.
[0658] 3. Navigation and real-time support:
[0659] The terminal activates the navigation system and begins calculating the optimal route. IoT data such as speed, fuel level, and location information collected from the vehicle's sensors is transmitted to the server in real time. The server analyzes this data and recalculates the optimal route by comparing it with traffic and weather information. The latest route information is sent to the terminal and provided to the user via speech synthesis. For example, instructions such as, "We have calculated the optimal route to avoid traffic congestion. Please turn right," are provided.
[0660] 4. Transaction processing:
[0661] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." This is converted into text by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message, "The nearest parking lot is XX. Would you like to make a reservation?" which the device reads aloud. If the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. Afterward, reservation confirmation information is sent to the device and notified by voice.
[0662] Examples of prompt statements
[0663] The following are specific examples of prompt statements to be input to a generative AI model:
[0664] 1. "Tell me your plans for today."
[0665] 2. "Calculate the optimal route."
[0666] 3. "Please find a nearby parking lot."
[0667] By using these prompts, users can utilize the system more smoothly and intuitively.
[0668] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0669] Processing steps
[0670] Step 1:
[0671] Input: User's face data (image)
[0672] Specific operation: The device uses the vehicle's onboard camera to automatically capture the user's face when the engine starts. This image data is then stored in the device.
[0673] Data processing: Acquired facial data is encrypted on the device.
[0674] Output: Encrypted facial data is generated.
[0675] Step 2:
[0676] Input: Encrypted facial data
[0677] Specific operation: The device sends encrypted facial data to the server via a secure communication channel.
[0678] Data processing: The server compares the received facial data with user information stored in a database on the cloud.
[0679] Output: An authentication result (success or failure) is generated.
[0680] Step 3:
[0681] Input: Authentication result
[0682] Specific action: The server sends the authentication result back to the terminal.
[0683] Data processing: If successful, corresponding user information will be attached.
[0684] Output: The authentication result (and user information) is sent to the terminal.
[0685] Step 4:
[0686] Input: Authentication result
[0687] Specific operation: The device receives the authentication result, and if it contains user information, it notifies the user via voice, "Hello, [username]. Where are you going today?"
[0688] Data processing: Convert text to speech using speech synthesis technology.
[0689] Output: An audio notification is generated for the user.
[0690] Step 5:
[0691] Input: User voice input ("What are my plans for today?")
[0692] Specific action: The user asks aloud, "What's on the schedule for today?"
[0693] Data processing: The speech recognition system converts the audio into text data.
[0694] Output: Text data ("What are your plans for today?") is generated.
[0695] Step 6:
[0696] Input: Text data ("What are your plans for today?")
[0697] Specific action: This text data is sent to the generating AI.
[0698] Data processing: The server uses generative AI to retrieve user appointments from cloud calendars and scheduling information.
[0699] Output: User schedule information is generated.
[0700] Step 7:
[0701] Input: User's schedule information
[0702] Specific operation: Based on the acquired information, the server generates a suggestion message such as, "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[0703] Data calculation: Text message generation
[0704] Output: A suggestion message is generated.
[0705] Step 8:
[0706] Input: Suggestion message
[0707] Specific action: The terminal reads out the suggestion message.
[0708] Data processing: Text is converted to speech using speech synthesis technology.
[0709] Output: An audio notification is generated for the user.
[0710] Step 9:
[0711] Input: User voice input ("Yes, I'm going to the office.")
[0712] Specific action: The user responds, "Yes, I will go to the office."
[0713] Data processing: The speech recognition system converts the audio into text data.
[0714] Output: Text data ("Yes, I will go to the office") is generated.
[0715] Step 10:
[0716] Input: Text data ("Yes, I will go to the office")
[0717] Specific action: The terminal receives this text data and sets the office as the destination.
[0718] Data calculation: Setting destination information
[0719] Output: Destination setting complete.
[0720] Step 11:
[0721] Input: Destination information
[0722] Specific action: The terminal activates the navigation system and begins calculating the optimal route.
[0723] Data calculation: Calculation of the optimal route
[0724] Output: A navigation route is generated.
[0725] Step 12:
[0726] Input: Vehicle sensor data (speed, fuel level, location information)
[0727] Specific operation: The terminal collects data from the vehicle's sensors in real time and sends it to the server.
[0728] Data processing: The server analyzes the collected IoT data and compares it with traffic and weather information.
[0729] Output: Analysis results are generated.
[0730] Step 13:
[0731] Input: Analysis results
[0732] Specific operation: The server recalculates the optimal route based on the analysis results and sends it to the terminal.
[0733] Data calculation: Recalculating navigation routes
[0734] Output: The recalculated navigation route is sent to the terminal.
[0735] Step 14:
[0736] Input: Recalculated navigation route
[0737] Specific operation: The terminal provides the user with the latest recalculated route information through speech synthesis.
[0738] Data processing: Text is converted to speech using speech synthesis technology.
[0739] Output: An audio notification is generated for the user.
[0740] Step 15:
[0741] Input: User voice input ("Find a nearby parking lot")
[0742] Specific action: The user reaches near their destination and gives a voice command saying, "Find a nearby parking lot."
[0743] Data processing: The speech recognition system converts the audio into text data.
[0744] Output: Text data ("Find a nearby parking lot") is generated.
[0745] Step 16:
[0746] Input: Text data ("Find nearby parking")
[0747] Specific operation: The server receives text data and collects information about the nearest parking lot.
[0748] Data processing: The server checks the availability of parking spaces and selects a suitable one.
[0749] Output: A message recommending parking is generated.
[0750] Step 17:
[0751] Input: Recommended parking message
[0752] Specific action: The device reads out a recommendation message and notifies the user, "The nearest parking lot is XX. Would you like to make a reservation?"
[0753] Data processing: Text is converted to speech using speech synthesis technology.
[0754] Output: An audio notification is generated for the user.
[0755] Step 18:
[0756] Input: User voice input ("Yes")
[0757] Specific action: When the user responds with "yes," the server accesses the parking reservation system and completes the reservation.
[0758] Data processing: Generating confirmation information for parking reservations
[0759] Output: Reservation confirmation information is generated and sent to the terminal.
[0760] Step 19:
[0761] Input: Reservation confirmation information
[0762] Specific operation: The device receives reservation confirmation information and notifies the user via voice.
[0763] Data processing: Text is converted to speech using speech synthesis technology.
[0764] Output: An audio notification is generated for the user.
[0765] Through these steps, users can travel to their destination comfortably and efficiently while receiving personalized support. The system as a whole utilizes speech recognition, generative AI, and real-time data collection and analysis technologies.
[0766] (Application Example 1)
[0767] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0768] Current autonomous vehicle systems face numerous challenges in improving the user experience, particularly in areas such as user authentication, destination setting, and real-time route optimization. For example, issues include difficulties with smooth facial recognition and cumbersome destination setting. Furthermore, a lack of proper integration of real-time optimal route calculations during driving and parking reservation information near the destination significantly reduces user convenience. Addressing these challenges requires advanced speech recognition, generative AI, and the integration of IoT data.
[0769] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0770] In this invention, the server includes means for capturing the user's face and performing facial recognition, means for transmitting real-time vehicle sensor data to the server and analyzing it, and means for collecting parking information near the destination, checking availability, and proposing and completing reservations. This makes it possible to smoothly perform everything from user authentication to destination setting, navigation, real-time route optimization, and parking reservation.
[0771] "User authentication" refers to the authentication process used when a user logs into a system, utilizing personal characteristics such as their face or voice.
[0772] "Speech recognition" is a technology that allows a machine to understand what a user says and convert it into text data.
[0773] "Generative AI" is an artificial intelligence technology that generates new data or answers based on pre-trained data.
[0774] "IoT data" refers to real-time information collected from various sensors and devices.
[0775] "Calculating the optimal route" means determining the most efficient route to the destination, taking into account current traffic conditions and road congestion.
[0776] "Speech synthesis" is a technology that converts text data into speech and provides information to the user as audio.
[0777] A "smartphone" is an electronic device that combines the functions of a mobile phone and a computer, enabling internet connectivity and application usage.
[0778] "Vehicle sensor data" refers to data such as speed, fuel level, and location information measured by sensors installed in the vehicle.
[0779] "Parking information" refers to information such as the location, availability, and fees of the parking lot.
[0780] A "reservation" is a procedure aimed at securing goods or services in advance.
[0781] "Authentication" refers to the process of verifying that someone is a person.
[0782] A "proposal" is an expression of an idea or plan regarding a particular subject.
[0783] "Analysis" refers to the process of analyzing collected data in detail.
[0784] This invention is a system that provides personalized support for autonomous vehicles, including user facial recognition, voice recognition, acquisition of schedule information using generative AI, real-time optimal route calculation, and parking reservation.
[0785] Hardware and software configuration
[0786] Hardware: Smartphones, cameras, in-vehicle sensors (speed sensors, fuel sensors, GPS, etc.)
[0787] Software: OpenCV, face_recognition, speech_recognition, cloud server, generative AI model
[0788] This system uses the user's smartphone camera to perform facial recognition before the user enters the vehicle. The facial data captured by the camera is encrypted and sent to a cloud server for facial recognition. If facial recognition is successful, the server retrieves the corresponding user's information and sends it back to the device.
[0789] Users set destinations and give other instructions using voice commands. These voice commands are received through the smartphone's microphone and converted into text data by speech_recognition software. The text data is sent to a generative AI model, which suggests appropriate destinations based on the user's schedule information. The generative AI model retrieves the user's schedule from sources such as cloud calendars and generates the optimal route and action plan.
[0790] Once a destination is set, data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors and sent to a cloud server. Based on the collected data, the server calculates the optimal route in real time and sends the result back to the terminal. The terminal uses speech synthesis to provide the user with the latest route information in real time. For example, it may provide specific instructions such as, "We have calculated the optimal route to avoid traffic. Please turn right."
[0791] When the user arrives near their destination, they instruct the terminal by voice saying, "Find a nearby parking lot." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a suggestion message such as, "The nearest parking lot is XX. Would you like to make a reservation?" If the user responds with "Yes," the server accesses the parking lot reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and the user is notified by voice.
[0792] Examples of prompt statements
[0793] "Please suggest destinations based on the user's schedule."
[0794] "Please calculate the optimal route to avoid traffic congestion."
[0795] "Please provide information on the availability and reservation status of nearby parking lots."
[0796] This system allows users to enjoy efficient and safe driving while receiving personalized support.
[0797] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0798] Step 1:
[0799] The user approaches the vehicle using their smartphone. The device (smartphone) activates its camera and captures the user's face. This image data is processed using a facial recognition framework (e.g., OpenCV and face_recognition). Facial recognition extracts facial feature points, which are then encrypted. The encrypted data is sent to a cloud server.
[0800] Input: Smartphone camera image
[0801] Data processing: Extraction and encryption of facial feature points.
[0802] Output: Encrypted facial data
[0803] Step 2:
[0804] The server authenticates the user by comparing the received encrypted facial data with a database in the cloud. If authentication is successful, the user's individual information is sent back from the server to the device. This completes the user authentication process.
[0805] Input: Encrypted facial data
[0806] Data processing: Database matching
[0807] Output: Authentication results and user information
[0808] Step 3:
[0809] The user gives voice commands to the device. The device receives the user's voice via its microphone and converts it into text data using the speech_recognition library. This text data is then sent to a cloud server.
[0810] Input: Audio data
[0811] Data processing: Speech-to-text conversion
[0812] Output: Text data of voice commands
[0813] Step 4:
[0814] The server analyzes the received text data based on a generative AI model. The generative AI model retrieves the user's schedule information from a cloud calendar and other sources, and suggests destinations. These suggestions are sent to the terminal and notified to the user via speech synthesis.
[0815] Input: Text data of voice commands
[0816] Data processing: Schedule acquisition and destination suggestion using generative AI models.
[0817] Output: Destination suggestion message
[0818] Step 5:
[0819] The user accepts the destination suggested by voice. For example, they might respond by saying, "Yes, I'm going to the office." This response is then converted back into text data by speech_recognition and sent to the server.
[0820] Input: Voice response
[0821] Data processing: Speech-to-text conversion
[0822] Output: Text data of the response
[0823] Step 6:
[0824] The server analyzes the received text data and sets the destination. Furthermore, real-time data from vehicle sensors (speed, fuel level, location) is collected and sent to the server. Based on this data, the server calculates the optimal route and sends the result to the terminal.
[0825] Input: Response text data and sensor data
[0826] Data processing: Calculation of the optimal route
[0827] Output: Optimal route information
[0828] Step 7:
[0829] The terminal uses speech synthesis to guide the user to the optimal route. For example, it might provide instructions such as, "The optimal route has been set. Please turn right at the next intersection."
[0830] Input: Optimal route information
[0831] Data processing: Speech synthesis
[0832] Output: Voice guidance
[0833] Step 8:
[0834] When the user approaches their destination, they will voice-instruct the device to "find a nearby parking lot." This instruction is also converted into text data via speech_recognition and sent to the server.
[0835] Input: Audio data for parking instructions
[0836] Data processing: Speech-to-text conversion
[0837] Output: Text data of parking instructions
[0838] Step 9:
[0839] The server analyzes the text data of the parking instructions and collects information about nearby parking lots. It checks availability and sends a suggestion message to the terminal indicating a suitable parking lot. A message such as "The nearest parking lot is XX. Would you like to make a reservation?" is generated.
[0840] Input: Text data for parking instructions
[0841] Data processing: Obtaining parking information and checking availability.
[0842] Output: Parking suggestion message
[0843] Step 10:
[0844] When the user responds with "yes," the terminal converts the response into text using its speech recognition system and sends it to the server. The server accesses the parking reservation system and completes the reservation. Reservation confirmation information is sent to the terminal and notified to the user via voice.
[0845] Input: Audio data of parking reservation response
[0846] Data processing: Text conversion and reservation processing of responses
[0847] Output: Reservation confirmation information
[0848] Through the processing steps described above, this system provides consistent support from user facial recognition to parking reservation, offering a safe and efficient driving environment.
[0849] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0850] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[0851] User authentication
[0852] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[0853] Setting a destination
[0854] The user asks the system by voice, "What's on my schedule today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to a generating AI, which retrieves the user's schedule from a cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0855] Navigation and real-time support
[0856] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[0857] How the emotion engine works
[0858] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[0859] Transaction processing
[0860] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[0861] summary
[0862] Through the above processes, users can enjoy an efficient and comfortable journey to their destination while receiving personalized support. This system can further enhance the user experience by recognizing and appropriately responding to emotional states, in addition to voice recognition, generative AI, and real-time data collection and analysis.
[0863] The following describes the processing flow.
[0864] Specific processing steps for carrying out the invention
[0865] User authentication
[0866] Step 1:
[0867] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[0868] Step 2:
[0869] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[0870] Step 3:
[0871] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[0872] Step 4:
[0873] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[0874] Step 5:
[0875] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[0876] Setting a destination
[0877] Step 1:
[0878] User: "What's on the schedule for today?" asks the system by voice.
[0879] Step 2:
[0880] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[0881] Step 3:
[0882] Terminal: Sends the converted text data to the generating AI.
[0883] Step 4:
[0884] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[0885] Step 5:
[0886] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[0887] Step 6:
[0888] Terminal: The generated message is read aloud to the user using speech synthesis.
[0889] Step 7:
[0890] User: "Yes, I'm going to the office," they reply, confirming the destination.
[0891] Navigation and real-time support
[0892] Step 1:
[0893] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[0894] Step 2:
[0895] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[0896] Step 3:
[0897] Terminal: Sends collected data to the server in real time.
[0898] Step 4:
[0899] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[0900] Step 5:
[0901] Server: Sends updated routes and additional information to the terminal.
[0902] Step 6:
[0903] Terminal: Notifies the user of updated information using speech synthesis.
[0904] Step 7:
[0905] User: Drive according to instructions based on traffic congestion and accident information.
[0906] How the emotion engine works
[0907] Step 1:
[0908] Terminal: Uses the vehicle's camera and microphone to capture the user's facial expressions and voice tone in real time.
[0909] Step 2:
[0910] Terminal: Sends captured data to the emotion engine.
[0911] Step 3:
[0912] Server: The emotion engine analyzes the data and determines the user's emotional state.
[0913] Step 4:
[0914] Server: Based on the user's emotional state, the generating AI proposes appropriate countermeasures.
[0915] Step 5:
[0916] Device: The generating AI reads aloud messages and suggestions tailored to the user's emotions using speech synthesis. For example, if the user is feeling stressed, it might suggest, "Would you like to play some music to relax?"
[0917] Transaction processing
[0918] Step 1:
[0919] User: "Find a nearby parking lot," gives a voice command.
[0920] Step 2:
[0921] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[0922] Step 3:
[0923] Terminal: Sends the converted text data to the server.
[0924] Step 4:
[0925] Server: Searches for the nearest parking lot and checks its availability.
[0926] Step 5:
[0927] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[0928] Step 6:
[0929] Terminal: The suggested message is read aloud to the user using speech synthesis.
[0930] Step 7:
[0931] User: Responds with "Yes" and instructs to reserve a parking space.
[0932] Step 8:
[0933] Server: Executes parking reservations through the system.
[0934] Step 9:
[0935] Server: Sends a reservation completion notification to the device.
[0936] Step 10:
[0937] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[0938] Through the processing flow described above, users can enjoy a comfortable driving environment that takes their emotional state into consideration, while also receiving individually personalized support.
[0939] (Example 2)
[0940] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0941] Conventional driver assistance systems typically only offer navigation and voice recognition functions, lacking features to address the individual needs and emotional states of users. As a result, while safe and efficient driver assistance may be achievable, providing emotional satisfaction and a personalized experience for users remains a challenge.
[0942] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0943] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, means using an emotion engine to analyze the user's emotional state, and means for the generative AI to make suggestions based on the emotional state. This makes it possible to respond to the user's individual needs and improve emotional satisfaction.
[0944] "Means of authenticating a user's identity" refers to a function that uses sensors such as cameras to acquire the user's biometric information and verifies the user's identity based on that information.
[0945] "A means of receiving user commands through voice recognition" refers to a technology that uses acoustic devices such as microphones to capture user voice instructions, analyzes them, and converts them into text data or commands.
[0946] "A method for acquiring user schedule information and suggesting destinations using generative AI" refers to a system that uses artificial intelligence to access user schedule data and suggests the optimal destination based on that data.
[0947] "A method for collecting IoT data and calculating the optimal route in real time" refers to a technology that analyzes traffic conditions and road information based on data collected from various sensors in a vehicle, and calculates the optimal route in real time.
[0948] "A means of providing information to users through speech synthesis" refers to a technology that converts text data into speech output and notifies users of necessary information via voice.
[0949] "Methods using an emotion engine to analyze the user's emotional state" refers to technologies that analyze the user's voice and facial expressions to estimate their emotional state and take appropriate action.
[0950] "A means by which AI generates suggestions based on emotional state" refers to a function in which the generating AI takes into account the user's emotional state and provides the user with appropriate responses and suggestions.
[0951] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[0952] First, as a means of user authentication, the terminal uses a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares it with user information stored in the cloud and performs authentication. If authentication is successful, the server sends the authentication result, including the corresponding user's individual information, back to the terminal, and the terminal notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[0953] Next, regarding setting the destination, the user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?", which the terminal reads aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[0954] Once a destination is set, the terminal activates its navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, upon receiving traffic congestion information, it might instruct the user, "We have calculated the optimal route to avoid congestion. Please turn right."
[0955] Regarding the operation of the emotion engine, the agent system collects the user's voice and facial expressions using the camera and microphone and sends them to the emotion engine. The emotion engine analyzes this data to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to make suggestions. Specifically, if the user is expressing stress, the device will make suggestions via voice, such as, "Shall I play some music to help you relax?" or "Shall I find a nearby resting place?"
[0956] Regarding transaction processing, when a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The terminal reads this message aloud. If the user responds with "yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and notified by voice.
[0957] As a concrete example of its operation, when a user gets into a car and asks, "What are my plans for today?", the system suggests, "I have a meeting at the office at 9:00," and navigation begins. Also, when near the destination, if the user says, "Find a nearby parking lot," the terminal converts the voice into text data, and the server searches for and reserves the nearest parking lot.
[0958] As an example of a prompt statement,
[0959] 1. "Please tell me about your appointments and plans for today."
[0960] 2. "Where are you planning to go today?"
[0961] There is.
[0962] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0963] Step 1:
[0964] The device uses a camera mounted on the vehicle to capture the user's face. When the user starts the engine, facial recognition begins automatically. The input data is the facial image captured by the camera, which is processed by a facial recognition algorithm to generate encrypted facial data. The output is the encrypted facial data.
[0965] Step 2:
[0966] The device sends encrypted facial data to the server. The input is encrypted facial data, and the output is the data sent to the server. Specifically, it uses an encrypted communication protocol to send the data.
[0967] Step 3:
[0968] The server performs authentication by comparing the received encrypted data with user information stored in the cloud. The input consists of encrypted facial data and facial data stored in the cloud, and a match is verified through a database search operation. The output is the authentication result.
[0969] Step 4:
[0970] If authentication is successful, the server returns an authentication result to the terminal, which includes the individual user information of the corresponding user. The input consists of the authentication result and user information, which are encrypted and sent to the terminal. The output is the authentication information received by the terminal.
[0971] Step 5:
[0972] Upon successful authentication, the device uses speech synthesis to notify the user, "Hello, Mr. / Ms. Yamada. Where are you going today?" The input is authentication information, which is converted from text to speech via a speech synthesis engine. The output is a voice notification.
[0973] Step 6:
[0974] The user asks the system, "What's on the schedule today?" using voice. The input is the user's voice command, which the terminal collects. The output is voice data.
[0975] Step 7:
[0976] The terminal analyzes voice commands using a speech recognition system and converts them into text data. The input is voice data, and the speech recognition engine converts the voice to text. The output is text data.
[0977] Step 8:
[0978] The device sends text data to the generating AI, and the server retrieves the user's schedule from a cloud calendar or other scheduling information. The input is text data, which is fed into the generating AI model. The output is the user's schedule information.
[0979] Step 9:
[0980] The server generates a suggestion message based on the acquired information: "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is schedule information, and a generation AI is used to generate the suggestion message. The output is the suggestion message.
[0981] Step 10:
[0982] The device notifies the user of this suggestion message via voice. The input is the suggestion message, which is converted from text to speech using a speech synthesis engine. The output is a voice notification.
[0983] Step 11:
[0984] The user responds by voice, "Yes, I'm going to the office," which sets the destination. The input is the user's voice response, which the terminal collects. The output is voice data.
[0985] Step 12:
[0986] The terminal uses a speech recognition system to convert the user's voice responses into text data and confirm the destination setting. The input is voice data, which is converted to text using the speech recognition engine. The output is text data.
[0987] Step 13:
[0988] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. The input is destination information, and the system uses a route calculation algorithm to generate the best route. The output is the optimal route information.
[0989] Step 14:
[0990] The terminal collects IoT data such as speed, fuel level, and location information from vehicle sensors in real time. The input is data from vehicle sensors, which is collected and analyzed. The output is real-time data.
[0991] Step 15:
[0992] The collected IoT data is sent to the server. The input is real-time data, and the output is data transmission to the server.
[0993] Step 16:
[0994] The server analyzes IoT data by cross-referencing it with traffic and weather information, and recalculates the optimal route. Inputs are real-time data and external information, and the optimal route is generated through data analysis. The output is the recalculated route information.
[0995] Step 17:
[0996] This function sends the latest route information to the terminal. The input is the recalculated route information, and the output is the data transmission to the terminal.
[0997] Step 18:
[0998] The device provides the user with the latest route information through speech synthesis. The input is the latest route information, which is converted into speech using the speech synthesis engine. The output is a voice notification.
[0999] (Application Example 2)
[1000] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1001] While conventional autonomous driving systems supported user authentication and destination setting, they did not provide driving support based on the user's emotional state. Furthermore, although they offered real-time optimal route calculations and voice-based information, they did not address the user's emotional needs. As a result, while they provided a safe and efficient driving environment, they were insufficient to fully deliver a comfortable driving experience for the user.
[1002] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for authenticating the user's personal information, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generation AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, and means for recognizing the user's emotional state using an emotion engine and providing support based on this. This makes it possible to respond to the user's emotional needs and to provide an even more comfortable and personalized driving experience.
[1003] "User personal authentication" is a method of recognizing the user's face using a camera installed in the vehicle and matching it with data in the cloud.
[1004] "Speech recognition" is a method of converting a user's voice into text data, allowing the system to accept user commands.
[1005] "Generative AI" is a method that, based on the user's voice instructions, retrieves schedule information from cloud calendars and other sources and suggests destinations.
[1006] "IoT data" refers to data such as speed, fuel level, and location information collected from vehicle sensors.
[1007] "A method for calculating the optimal route in real time" refers to a method that uses IoT data to calculate the optimal route by comparing it with traffic and weather information.
[1008] "Speech synthesis" is a method of providing users with calculation results or suggested messages in voice.
[1009] An "emotion engine" is a method of recognizing a user's emotional state by analyzing their voice and facial expressions.
[1010] "Means of providing support" refers to methods of proposing solutions that meet the user's emotional needs based on the analysis results of the emotion engine.
[1011] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This system not only realizes a safe and efficient driving environment but also addresses the user's emotional needs, providing a more comfortable and personalized driving experience.
[1012] User authentication
[1013] The server captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the cloud server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the vehicle terminal. The terminal notifies the user by voice, "Hello, user. Where are you going today?"
[1014] Setting a destination
[1015] The user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the vehicle terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9:00 today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1016] Navigation and real-time support
[1017] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[1018] How the emotion engine works
[1019] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[1020] Parking reservation support
[1021] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[1022] Examples of specific cases and prompt statements
[1023] For example, if a user asks, "What's on my schedule today?", the system retrieves information from a cloud calendar and suggests, "You have a meeting at the office at 9:00." Also, if the user is feeling stressed, the system might suggest, "Would you like some relaxing music?"
[1024] Example of a prompt:
[1025] What are your plans for today? Please tell me your schedule.
[1026]
[1027] Please suggest ways to address situations where users are experiencing stress.
[1028] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1029] Step 1:
[1030] The camera is activated and the user's face is captured. For authentication, the facial image acquired from the camera is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. The input is facial image data, and the output is the authentication result. Specifically, facial recognition is performed using OpenCV, and if authentication is successful, the corresponding user information is sent back to the terminal.
[1031] Step 2:
[1032] The device notifies the user via speech synthesis, "Hello, user. Where are you going today?" The user's voice response is captured by the microphone and converted into text data using a speech recognition system. This converted text data is the input, and the speech recognition result is the output. Specifically, the Google Cloud Speech-to-Text API is used to convert speech to text.
[1033] Step 3:
[1034] When a user asks "What's on my schedule today?", a generative AI retrieves the user's schedule information. The generative AI obtains the user's schedule from a cloud calendar and other schedule information and generates destination suggestions. The input to this process is the user's voice command "What's on my schedule today?", and the output is destination suggestions. Specifically, it uses the OpenAI GPT model to obtain schedule information based on the prompt sentence and generates suggestion messages.
[1035] Step 4:
[1036] The device synthesizes suggested messages into speech and notifies the user. For example, it might read aloud the message, "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is a suggested message from the generation AI, and the output is a synthesized speech notification.
[1037] Step 5:
[1038] Once the user confirms the destination setting, the navigation system activates and begins calculating the optimal route. The device collects IoT data such as speed, fuel level, and location information in real time from sensors and sends it to the server. The server compares this data with traffic and weather information and recalculates the optimal route. The input is IoT data and real-time traffic and weather information, and the output is instructions for the optimal route.
[1039] Step 6:
[1040] The device uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotional state. The input is the user's voice and facial expression data, and the output is the result of the emotional state analysis. Specifically, it uses voice and video analysis technologies to analyze the user's emotions.
[1041] Step 7:
[1042] The server uses generative AI to suggest countermeasures based on the analysis results from the emotion engine. For example, if the user is feeling stressed, it might suggest, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?" The input is the analysis results from the emotion engine, and the output is the suggested countermeasures. Specifically, the generative AI makes suggestions based on the user's emotional needs.
[1043] Step 8:
[1044] When a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This is converted into text data by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It then selects a suitable parking lot and prompts the user to make a reservation. The input is the user's voice command, and the output is parking availability information and reservation confirmation.
[1045] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1046] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1047] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1048] [Third Embodiment]
[1049] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1050] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1051] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1052] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1053] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1054] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1055] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1056] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1057] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1058] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1059] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1060] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1061] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is programmed according to the following procedure.
[1062] User authentication
[1063] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[1064] Setting a destination
[1065] The user asks the system by voice, "What's on my schedule today?" This voice command is converted into text data by the terminal's voice recognition system. The text data is sent to a generating AI, and the server retrieves the user's schedule from the cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1066] Navigation and real-time support
[1067] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[1068] Transaction processing
[1069] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[1070] summary
[1071] Through the above processes, users can efficiently navigate their journey to their destination while receiving personalized support. This system utilizes voice recognition, generative AI, and real-time data collection and analysis to provide the latest and most relevant information, ensuring a comfortable user experience for both drivers and passengers.
[1072] The following describes the processing flow.
[1073] Specific processing steps for carrying out the invention
[1074] User authentication
[1075] Step 1:
[1076] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[1077] Step 2:
[1078] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[1079] Step 3:
[1080] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[1081] Step 4:
[1082] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[1083] Step 5:
[1084] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[1085] Setting a destination
[1086] Step 1:
[1087] User: "What's on the schedule for today?" asks the system by voice.
[1088] Step 2:
[1089] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[1090] Step 3:
[1091] Terminal: Sends the converted text data to the generating AI.
[1092] Step 4:
[1093] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[1094] Step 5:
[1095] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[1096] Step 6:
[1097] Terminal: The generated message is read aloud to the user using speech synthesis.
[1098] Step 7:
[1099] User: "Yes, I'm going to the office," they reply, confirming the destination.
[1100] Navigation and real-time support
[1101] Step 1:
[1102] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[1103] Step 2:
[1104] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[1105] Step 3:
[1106] Terminal: Sends collected data to the server in real time.
[1107] Step 4:
[1108] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[1109] Step 5:
[1110] Server: Sends updated routes and additional information to the terminal.
[1111] Step 6:
[1112] Terminal: Notifies the user of updated information using speech synthesis.
[1113] Step 7:
[1114] User: Drive according to instructions based on traffic congestion and accident information.
[1115] Transaction processing
[1116] Step 1:
[1117] User: "Find a nearby parking lot," gives a voice command.
[1118] Step 2:
[1119] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[1120] Step 3:
[1121] Terminal: Sends the converted text data to the server.
[1122] Step 4:
[1123] Server: Searches for the nearest parking lot and checks its availability.
[1124] Step 5:
[1125] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[1126] Step 6:
[1127] Terminal: The suggested message is read aloud to the user using speech synthesis.
[1128] Step 7:
[1129] User: Responds with "Yes" and instructs to reserve a parking space.
[1130] Step 8:
[1131] Server: Executes parking reservations through the system.
[1132] Step 9:
[1133] Server: Sends a reservation completion notification to the device.
[1134] Step 10:
[1135] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[1136] In this way, users can arrive at their destination comfortably and efficiently through a series of processes. This allows them to spend their time in the car more meaningfully.
[1137] (Example 1)
[1138] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1139] Conventional vehicle driver assistance systems are increasingly expected to offer advanced functions beyond user authentication and voice recognition-based command acceptance, such as real-time information provision, schedule management, and parking reservation. However, few systems provide these functions in an integrated manner, resulting in limited improvements to the user experience. Furthermore, real-time data analysis and navigation optimization have been insufficient, making it difficult to provide a safe and efficient driving environment.
[1140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1141] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands via voice recognition, means for acquiring user schedule information and suggesting destinations using generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user via speech synthesis, means for activating a navigation system and recalculating and providing the route in real time, means for collecting and analyzing speed, fuel level, and location information from vehicle sensors, and means for collecting parking information and executing reservations. This enables an improved user experience and the realization of a safe and efficient driving environment.
[1142] "User authentication" is a process that uses cameras mounted on the vehicle to capture the user's face and compares the encrypted facial data with a database in the cloud.
[1143] "Voice recognition" is a technology that converts the voice spoken by a user inside a vehicle into text data, which the system then accepts as a command it can understand.
[1144] "Generative AI" is an algorithm that uses artificial intelligence technology to analyze a user's schedule information and suggest appropriate destinations and actions.
[1145] "IoT data" refers to real-time data such as speed, fuel level, and location information collected from vehicle sensors.
[1146] "Calculating the optimal route in real time" is a process that instantly analyzes the most efficient travel route based on collected IoT data, external traffic information, and weather information.
[1147] "Speech synthesis" is a technology that converts text data generated by a system into speech and provides information to the user.
[1148] "Activating the navigation system" means activating the vehicle's navigation system to begin providing optimal route guidance to the set destination.
[1149] "Recalculating and providing routes" means re-analyzing the optimal route based on real-time data during travel and providing the user with the latest route information.
[1150] "Vehicle sensors" are devices used to measure various conditions inside and around the vehicle, collecting information such as speed, fuel level, and location.
[1151] "Gathering parking information" is the process of finding available parking spaces and their availability near your current location.
[1152] "Executing a reservation" refers to the process of reserving a suitable parking space for the user based on the collected parking information.
[1153] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is implemented using the following hardware and software:
[1154] hardware
[1155] Camera: Mounted in the vehicle and used to capture the user's face.
[1156] Vehicle sensors: Used to collect speed, fuel level, and location information in real time.
[1157] Terminal: A computer device installed inside a vehicle that runs voice recognition systems, generative AI, and navigation systems.
[1158] Server: Located in a cloud environment, it performs facial recognition data matching, IoT data analysis, and acquires schedule information and calculates routes using generated AI.
[1159] software
[1160] Speech recognition system: Converts user speech into text data.
[1161] Generative AI model: Analyzes user schedule information and suggests appropriate destinations and actions.
[1162] Navigation system: Calculates the optimal route and guides the user.
[1163] Encryption technology: Used to securely transmit user facial data to the cloud.
[1164] This system is implemented in the following steps:
[1165] 1. User authentication:
[1166] The terminal uses a camera mounted on the vehicle to capture the user's face. The authentication process starts automatically when the engine is started. The acquired facial data is encrypted and sent to a server. The server compares the facial data with a database in the cloud and sends the authentication result back to the terminal. The terminal then notifies the user by voice, "Hello, [username]. Where are you going today?"
[1167] 2. Setting the destination:
[1168] The user asks aloud, "What's on my schedule today?" A speech recognition system converts the speech into text data, which is then sent to a generating AI. The server retrieves the user's schedule from the cloud calendar and generates a message saying, "You have a meeting at the office at 9am today. Shall we head to the office?" The device reads this message aloud. When the user responds, "Yes, I'll go to the office," the destination is set.
[1169] 3. Navigation and real-time support:
[1170] The terminal activates the navigation system and begins calculating the optimal route. IoT data such as speed, fuel level, and location information collected from the vehicle's sensors is transmitted to the server in real time. The server analyzes this data and recalculates the optimal route by comparing it with traffic and weather information. The latest route information is sent to the terminal and provided to the user via speech synthesis. For example, instructions such as, "We have calculated the optimal route to avoid traffic congestion. Please turn right," are provided.
[1171] 4. Transaction processing:
[1172] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." This is converted into text by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message, "The nearest parking lot is XX. Would you like to make a reservation?" which the device reads aloud. If the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. Afterward, reservation confirmation information is sent to the device and notified by voice.
[1173] Examples of prompt statements
[1174] The following are specific examples of prompt statements to be input to a generative AI model:
[1175] 1. "Tell me your plans for today."
[1176] 2. "Calculate the optimal route."
[1177] 3. "Please find a nearby parking lot."
[1178] By using these prompts, users can utilize the system more smoothly and intuitively.
[1179] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1180] Processing steps
[1181] Step 1:
[1182] Input: User's face data (image)
[1183] Specific operation: The device uses the vehicle's onboard camera to automatically capture the user's face when the engine starts. This image data is then stored in the device.
[1184] Data processing: Acquired facial data is encrypted on the device.
[1185] Output: Encrypted facial data is generated.
[1186] Step 2:
[1187] Input: Encrypted facial data
[1188] Specific operation: The device sends encrypted facial data to the server via a secure communication channel.
[1189] Data processing: The server compares the received facial data with user information stored in a database on the cloud.
[1190] Output: An authentication result (success or failure) is generated.
[1191] Step 3:
[1192] Input: Authentication result
[1193] Specific action: The server sends the authentication result back to the terminal.
[1194] Data processing: If successful, corresponding user information will be attached.
[1195] Output: The authentication result (and user information) is sent to the terminal.
[1196] Step 4:
[1197] Input: Authentication result
[1198] Specific operation: The device receives the authentication result, and if it contains user information, it notifies the user via voice, "Hello, [username]. Where are you going today?"
[1199] Data processing: Convert text to speech using speech synthesis technology.
[1200] Output: An audio notification is generated for the user.
[1201] Step 5:
[1202] Input: User voice input ("What are my plans for today?")
[1203] Specific action: The user asks aloud, "What's on the schedule for today?"
[1204] Data processing: The speech recognition system converts the audio into text data.
[1205] Output: Text data ("What are your plans for today?") is generated.
[1206] Step 6:
[1207] Input: Text data ("What are your plans for today?")
[1208] Specific action: This text data is sent to the generating AI.
[1209] Data processing: The server uses generative AI to retrieve user appointments from cloud calendars and scheduling information.
[1210] Output: User schedule information is generated.
[1211] Step 7:
[1212] Input: User's schedule information
[1213] Specific operation: Based on the acquired information, the server generates a suggestion message such as, "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[1214] Data calculation: Text message generation
[1215] Output: A suggestion message is generated.
[1216] Step 8:
[1217] Input: Suggestion message
[1218] Specific action: The terminal reads out the suggestion message.
[1219] Data processing: Text is converted to speech using speech synthesis technology.
[1220] Output: An audio notification is generated for the user.
[1221] Step 9:
[1222] Input: User voice input ("Yes, I'm going to the office.")
[1223] Specific action: The user responds, "Yes, I will go to the office."
[1224] Data processing: The speech recognition system converts the audio into text data.
[1225] Output: Text data ("Yes, I will go to the office") is generated.
[1226] Step 10:
[1227] Input: Text data ("Yes, I will go to the office")
[1228] Specific action: The terminal receives this text data and sets the office as the destination.
[1229] Data calculation: Setting destination information
[1230] Output: Destination setting complete.
[1231] Step 11:
[1232] Input: Destination information
[1233] Specific action: The terminal activates the navigation system and begins calculating the optimal route.
[1234] Data calculation: Calculation of the optimal route
[1235] Output: A navigation route is generated.
[1236] Step 12:
[1237] Input: Vehicle sensor data (speed, fuel level, location information)
[1238] Specific operation: The terminal collects data from the vehicle's sensors in real time and sends it to the server.
[1239] Data processing: The server analyzes the collected IoT data and compares it with traffic and weather information.
[1240] Output: Analysis results are generated.
[1241] Step 13:
[1242] Input: Analysis results
[1243] Specific operation: The server recalculates the optimal route based on the analysis results and sends it to the terminal.
[1244] Data calculation: Recalculating navigation routes
[1245] Output: The recalculated navigation route is sent to the terminal.
[1246] Step 14:
[1247] Input: Recalculated navigation route
[1248] Specific operation: The terminal provides the user with the latest recalculated route information through speech synthesis.
[1249] Data processing: Text is converted to speech using speech synthesis technology.
[1250] Output: An audio notification is generated for the user.
[1251] Step 15:
[1252] Input: User voice input ("Find a nearby parking lot")
[1253] Specific action: The user reaches near their destination and gives a voice command saying, "Find a nearby parking lot."
[1254] Data processing: The speech recognition system converts the audio into text data.
[1255] Output: Text data ("Find a nearby parking lot") is generated.
[1256] Step 16:
[1257] Input: Text data ("Find nearby parking")
[1258] Specific operation: The server receives text data and collects information about the nearest parking lot.
[1259] Data processing: The server checks the availability of parking spaces and selects a suitable one.
[1260] Output: A message recommending parking is generated.
[1261] Step 17:
[1262] Input: Recommended parking message
[1263] Specific action: The device reads out a recommendation message and notifies the user, "The nearest parking lot is XX. Would you like to make a reservation?"
[1264] Data processing: Text is converted to speech using speech synthesis technology.
[1265] Output: An audio notification is generated for the user.
[1266] Step 18:
[1267] Input: User voice input ("Yes")
[1268] Specific action: When the user responds with "yes," the server accesses the parking reservation system and completes the reservation.
[1269] Data processing: Generating confirmation information for parking reservations
[1270] Output: Reservation confirmation information is generated and sent to the terminal.
[1271] Step 19:
[1272] Input: Reservation confirmation information
[1273] Specific operation: The device receives reservation confirmation information and notifies the user via voice.
[1274] Data processing: Text is converted to speech using speech synthesis technology.
[1275] Output: An audio notification is generated for the user.
[1276] Through these steps, users can travel to their destination comfortably and efficiently while receiving personalized support. The system as a whole utilizes speech recognition, generative AI, and real-time data collection and analysis technologies.
[1277] (Application Example 1)
[1278] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1279] Current autonomous vehicle systems face numerous challenges in improving the user experience, particularly in areas such as user authentication, destination setting, and real-time route optimization. For example, issues include difficulties with smooth facial recognition and cumbersome destination setting. Furthermore, a lack of proper integration of real-time optimal route calculations during driving and parking reservation information near the destination significantly reduces user convenience. Addressing these challenges requires advanced speech recognition, generative AI, and the integration of IoT data.
[1280] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1281] In this invention, the server includes means for capturing the user's face and performing facial recognition, means for transmitting real-time vehicle sensor data to the server and analyzing it, and means for collecting parking information near the destination, checking availability, and proposing and completing reservations. This makes it possible to smoothly perform everything from user authentication to destination setting, navigation, real-time route optimization, and parking reservation.
[1282] "User authentication" refers to the authentication process used when a user logs into a system, utilizing personal characteristics such as their face or voice.
[1283] "Speech recognition" is a technology that allows a machine to understand what a user says and convert it into text data.
[1284] "Generative AI" is an artificial intelligence technology that generates new data or answers based on pre-trained data.
[1285] "IoT data" refers to real-time information collected from various sensors and devices.
[1286] "Calculating the optimal route" means determining the most efficient route to the destination, taking into account current traffic conditions and road congestion.
[1287] "Speech synthesis" is a technology that converts text data into speech and provides information to the user as audio.
[1288] A "smartphone" is an electronic device that combines the functions of a mobile phone and a computer, enabling internet connectivity and application usage.
[1289] "Vehicle sensor data" refers to data such as speed, fuel level, and location information measured by sensors installed in the vehicle.
[1290] "Parking information" refers to information such as the location, availability, and fees of the parking lot.
[1291] A "reservation" is a procedure aimed at securing goods or services in advance.
[1292] "Authentication" refers to the process of verifying that someone is a person.
[1293] A "proposal" is an expression of an idea or plan regarding a particular subject.
[1294] "Analysis" refers to the process of analyzing collected data in detail.
[1295] This invention is a system that provides personalized support for autonomous vehicles, including user facial recognition, voice recognition, acquisition of schedule information using generative AI, real-time optimal route calculation, and parking reservation.
[1296] Hardware and software configuration
[1297] Hardware: Smartphones, cameras, in-vehicle sensors (speed sensors, fuel sensors, GPS, etc.)
[1298] Software: OpenCV, face_recognition, speech_recognition, cloud server, generative AI model
[1299] This system uses the user's smartphone camera to perform facial recognition before the user enters the vehicle. The facial data captured by the camera is encrypted and sent to a cloud server for facial recognition. If facial recognition is successful, the server retrieves the corresponding user's information and sends it back to the device.
[1300] Users set destinations and give other instructions using voice commands. These voice commands are received through the smartphone's microphone and converted into text data by speech_recognition software. The text data is sent to a generative AI model, which suggests appropriate destinations based on the user's schedule information. The generative AI model retrieves the user's schedule from sources such as cloud calendars and generates the optimal route and action plan.
[1301] Once a destination is set, data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors and sent to a cloud server. Based on the collected data, the server calculates the optimal route in real time and sends the result back to the terminal. The terminal uses speech synthesis to provide the user with the latest route information in real time. For example, it may provide specific instructions such as, "We have calculated the optimal route to avoid traffic. Please turn right."
[1302] When the user arrives near their destination, they instruct the terminal by voice saying, "Find a nearby parking lot." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a suggestion message such as, "The nearest parking lot is XX. Would you like to make a reservation?" If the user responds with "Yes," the server accesses the parking lot reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and the user is notified by voice.
[1303] Examples of prompt statements
[1304] "Please suggest destinations based on the user's schedule."
[1305] "Please calculate the optimal route to avoid traffic congestion."
[1306] "Please provide information on the availability and reservation status of nearby parking lots."
[1307] This system allows users to enjoy efficient and safe driving while receiving personalized support.
[1308] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1309] Step 1:
[1310] The user approaches the vehicle using their smartphone. The device (smartphone) activates its camera and captures the user's face. This image data is processed using a facial recognition framework (e.g., OpenCV and face_recognition). Facial recognition extracts facial feature points, which are then encrypted. The encrypted data is sent to a cloud server.
[1311] Input: Smartphone camera image
[1312] Data processing: Extraction and encryption of facial feature points.
[1313] Output: Encrypted facial data
[1314] Step 2:
[1315] The server authenticates the user by comparing the received encrypted facial data with a database in the cloud. If authentication is successful, the user's individual information is sent back from the server to the device. This completes the user authentication process.
[1316] Input: Encrypted facial data
[1317] Data processing: Database matching
[1318] Output: Authentication results and user information
[1319] Step 3:
[1320] The user gives voice commands to the device. The device receives the user's voice via its microphone and converts it into text data using the speech_recognition library. This text data is then sent to a cloud server.
[1321] Input: Audio data
[1322] Data processing: Speech-to-text conversion
[1323] Output: Text data of voice commands
[1324] Step 4:
[1325] The server analyzes the received text data based on a generative AI model. The generative AI model retrieves the user's schedule information from a cloud calendar and other sources, and suggests destinations. These suggestions are sent to the terminal and notified to the user via speech synthesis.
[1326] Input: Text data of voice commands
[1327] Data processing: Schedule acquisition and destination suggestion using generative AI models.
[1328] Output: Destination suggestion message
[1329] Step 5:
[1330] The user accepts the destination suggested by voice. For example, they might respond by saying, "Yes, I'm going to the office." This response is then converted back into text data by speech_recognition and sent to the server.
[1331] Input: Voice response
[1332] Data processing: Speech-to-text conversion
[1333] Output: Text data of the response
[1334] Step 6:
[1335] The server analyzes the received text data and sets the destination. Furthermore, real-time data from vehicle sensors (speed, fuel level, location) is collected and sent to the server. Based on this data, the server calculates the optimal route and sends the result to the terminal.
[1336] Input: Response text data and sensor data
[1337] Data processing: Calculation of the optimal route
[1338] Output: Optimal route information
[1339] Step 7:
[1340] The terminal uses speech synthesis to guide the user to the optimal route. For example, it might provide instructions such as, "The optimal route has been set. Please turn right at the next intersection."
[1341] Input: Optimal route information
[1342] Data processing: Speech synthesis
[1343] Output: Voice guidance
[1344] Step 8:
[1345] When the user approaches their destination, they will voice-instruct the device to "find a nearby parking lot." This instruction is also converted into text data via speech_recognition and sent to the server.
[1346] Input: Audio data for parking instructions
[1347] Data processing: Speech-to-text conversion
[1348] Output: Text data of parking instructions
[1349] Step 9:
[1350] The server analyzes the text data of the parking instructions and collects information about nearby parking lots. It checks availability and sends a suggestion message to the terminal indicating a suitable parking lot. A message such as "The nearest parking lot is XX. Would you like to make a reservation?" is generated.
[1351] Input: Text data for parking instructions
[1352] Data processing: Obtaining parking information and checking availability.
[1353] Output: Parking suggestion message
[1354] Step 10:
[1355] When the user responds with "yes," the terminal converts the response into text using its speech recognition system and sends it to the server. The server accesses the parking reservation system and completes the reservation. Reservation confirmation information is sent to the terminal and notified to the user via voice.
[1356] Input: Audio data of parking reservation response
[1357] Data processing: Text conversion and reservation processing of responses
[1358] Output: Reservation confirmation information
[1359] Through the processing steps described above, this system provides consistent support from user facial recognition to parking reservation, offering a safe and efficient driving environment.
[1360] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1361] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[1362] User authentication
[1363] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[1364] Setting a destination
[1365] The user asks the system by voice, "What's on my schedule today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to a generating AI, which retrieves the user's schedule from a cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1366] Navigation and real-time support
[1367] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[1368] How the emotion engine works
[1369] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[1370] Transaction processing
[1371] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[1372] summary
[1373] Through the above processes, users can enjoy an efficient and comfortable journey to their destination while receiving personalized support. This system can further enhance the user experience by recognizing and appropriately responding to emotional states, in addition to voice recognition, generative AI, and real-time data collection and analysis.
[1374] The following describes the processing flow.
[1375] Specific processing steps for carrying out the invention
[1376] User authentication
[1377] Step 1:
[1378] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[1379] Step 2:
[1380] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[1381] Step 3:
[1382] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[1383] Step 4:
[1384] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[1385] Step 5:
[1386] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[1387] Setting a destination
[1388] Step 1:
[1389] User: "What's on the schedule for today?" asks the system by voice.
[1390] Step 2:
[1391] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[1392] Step 3:
[1393] Terminal: Sends the converted text data to the generating AI.
[1394] Step 4:
[1395] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[1396] Step 5:
[1397] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[1398] Step 6:
[1399] Terminal: The generated message is read aloud to the user using speech synthesis.
[1400] Step 7:
[1401] User: "Yes, I'm going to the office," they reply, confirming the destination.
[1402] Navigation and real-time support
[1403] Step 1:
[1404] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[1405] Step 2:
[1406] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[1407] Step 3:
[1408] Terminal: Sends collected data to the server in real time.
[1409] Step 4:
[1410] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[1411] Step 5:
[1412] Server: Sends updated routes and additional information to the terminal.
[1413] Step 6:
[1414] Terminal: Notifies the user of updated information using speech synthesis.
[1415] Step 7:
[1416] User: Drive according to instructions based on traffic congestion and accident information.
[1417] How the emotion engine works
[1418] Step 1:
[1419] Terminal: Uses the vehicle's camera and microphone to capture the user's facial expressions and voice tone in real time.
[1420] Step 2:
[1421] Terminal: Sends captured data to the emotion engine.
[1422] Step 3:
[1423] Server: The emotion engine analyzes the data and determines the user's emotional state.
[1424] Step 4:
[1425] Server: Based on the user's emotional state, the generating AI proposes appropriate countermeasures.
[1426] Step 5:
[1427] Device: The generating AI reads aloud messages and suggestions tailored to the user's emotions using speech synthesis. For example, if the user is feeling stressed, it might suggest, "Would you like to play some music to relax?"
[1428] Transaction processing
[1429] Step 1:
[1430] User: "Find a nearby parking lot," gives a voice command.
[1431] Step 2:
[1432] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[1433] Step 3:
[1434] Terminal: Sends the converted text data to the server.
[1435] Step 4:
[1436] Server: Searches for the nearest parking lot and checks its availability.
[1437] Step 5:
[1438] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[1439] Step 6:
[1440] Terminal: The suggested message is read aloud to the user using speech synthesis.
[1441] Step 7:
[1442] User: Responds with "Yes" and instructs to reserve a parking space.
[1443] Step 8:
[1444] Server: Executes parking reservations through the system.
[1445] Step 9:
[1446] Server: Sends a reservation completion notification to the device.
[1447] Step 10:
[1448] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[1449] Through the processing flow described above, users can enjoy a comfortable driving environment that takes their emotional state into consideration, while also receiving individually personalized support.
[1450] (Example 2)
[1451] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1452] Conventional driver assistance systems typically only offer navigation and voice recognition functions, lacking features to address the individual needs and emotional states of users. As a result, while safe and efficient driver assistance may be achievable, providing emotional satisfaction and a personalized experience for users remains a challenge.
[1453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1454] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, means using an emotion engine to analyze the user's emotional state, and means for the generative AI to make suggestions based on the emotional state. This makes it possible to respond to the user's individual needs and improve emotional satisfaction.
[1455] "Means of authenticating a user's identity" refers to a function that uses sensors such as cameras to acquire the user's biometric information and verifies the user's identity based on that information.
[1456] "A means of receiving user commands through voice recognition" refers to a technology that uses acoustic devices such as microphones to capture user voice instructions, analyzes them, and converts them into text data or commands.
[1457] "A method for acquiring user schedule information and suggesting destinations using generative AI" refers to a system that uses artificial intelligence to access user schedule data and suggests the optimal destination based on that data.
[1458] "A method for collecting IoT data and calculating the optimal route in real time" refers to a technology that analyzes traffic conditions and road information based on data collected from various sensors in a vehicle, and calculates the optimal route in real time.
[1459] "A means of providing information to users through speech synthesis" refers to a technology that converts text data into speech output and notifies users of necessary information via voice.
[1460] "Methods using an emotion engine to analyze the user's emotional state" refers to technologies that analyze the user's voice and facial expressions to estimate their emotional state and take appropriate action.
[1461] "A means by which AI generates suggestions based on emotional state" refers to a function in which the generating AI takes into account the user's emotional state and provides the user with appropriate responses and suggestions.
[1462] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[1463] First, as a means of user authentication, the terminal uses a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares it with user information stored in the cloud and performs authentication. If authentication is successful, the server sends the authentication result, including the corresponding user's individual information, back to the terminal, and the terminal notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[1464] Next, regarding setting the destination, the user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?", which the terminal reads aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1465] Once a destination is set, the terminal activates its navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, upon receiving traffic congestion information, it might instruct the user, "We have calculated the optimal route to avoid congestion. Please turn right."
[1466] Regarding the operation of the emotion engine, the agent system collects the user's voice and facial expressions using the camera and microphone and sends them to the emotion engine. The emotion engine analyzes this data to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to make suggestions. Specifically, if the user is expressing stress, the device will make suggestions via voice, such as, "Shall I play some music to help you relax?" or "Shall I find a nearby resting place?"
[1467] Regarding transaction processing, when a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The terminal reads this message aloud. If the user responds with "yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and notified by voice.
[1468] As a concrete example of its operation, when a user gets into a car and asks, "What are my plans for today?", the system suggests, "I have a meeting at the office at 9:00," and navigation begins. Also, when near the destination, if the user says, "Find a nearby parking lot," the terminal converts the voice into text data, and the server searches for and reserves the nearest parking lot.
[1469] As an example of a prompt statement,
[1470] 1. "Please tell me about your appointments and plans for today."
[1471] 2. "Where are you planning to go today?"
[1472] There is.
[1473] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1474] Step 1:
[1475] The device uses a camera mounted on the vehicle to capture the user's face. When the user starts the engine, facial recognition begins automatically. The input data is the facial image captured by the camera, which is processed by a facial recognition algorithm to generate encrypted facial data. The output is the encrypted facial data.
[1476] Step 2:
[1477] The device sends encrypted facial data to the server. The input is encrypted facial data, and the output is the data sent to the server. Specifically, it uses an encrypted communication protocol to send the data.
[1478] Step 3:
[1479] The server performs authentication by comparing the received encrypted data with user information stored in the cloud. The input consists of encrypted facial data and facial data stored in the cloud, and a match is verified through a database search operation. The output is the authentication result.
[1480] Step 4:
[1481] If authentication is successful, the server returns an authentication result to the terminal, which includes the individual user information of the corresponding user. The input consists of the authentication result and user information, which are encrypted and sent to the terminal. The output is the authentication information received by the terminal.
[1482] Step 5:
[1483] Upon successful authentication, the device uses speech synthesis to notify the user, "Hello, Mr. / Ms. Yamada. Where are you going today?" The input is authentication information, which is converted from text to speech via a speech synthesis engine. The output is a voice notification.
[1484] Step 6:
[1485] The user asks the system, "What's on the schedule today?" using voice. The input is the user's voice command, which the terminal collects. The output is voice data.
[1486] Step 7:
[1487] The terminal analyzes voice commands using a speech recognition system and converts them into text data. The input is voice data, and the speech recognition engine converts the voice to text. The output is text data.
[1488] Step 8:
[1489] The device sends text data to the generating AI, and the server retrieves the user's schedule from a cloud calendar or other scheduling information. The input is text data, which is fed into the generating AI model. The output is the user's schedule information.
[1490] Step 9:
[1491] The server generates a suggestion message based on the acquired information: "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is schedule information, and a generation AI is used to generate the suggestion message. The output is the suggestion message.
[1492] Step 10:
[1493] The device notifies the user of this suggestion message via voice. The input is the suggestion message, which is converted from text to speech using a speech synthesis engine. The output is a voice notification.
[1494] Step 11:
[1495] The user responds by voice, "Yes, I'm going to the office," which sets the destination. The input is the user's voice response, which the terminal collects. The output is voice data.
[1496] Step 12:
[1497] The terminal uses a speech recognition system to convert the user's voice responses into text data and confirm the destination setting. The input is voice data, which is converted to text using the speech recognition engine. The output is text data.
[1498] Step 13:
[1499] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. The input is destination information, and the system uses a route calculation algorithm to generate the best route. The output is the optimal route information.
[1500] Step 14:
[1501] The terminal collects IoT data such as speed, fuel level, and location information from vehicle sensors in real time. The input is data from vehicle sensors, which is collected and analyzed. The output is real-time data.
[1502] Step 15:
[1503] The collected IoT data is sent to the server. The input is real-time data, and the output is data transmission to the server.
[1504] Step 16:
[1505] The server analyzes IoT data by cross-referencing it with traffic and weather information, and recalculates the optimal route. Inputs are real-time data and external information, and the optimal route is generated through data analysis. The output is the recalculated route information.
[1506] Step 17:
[1507] This function sends the latest route information to the terminal. The input is the recalculated route information, and the output is the data transmission to the terminal.
[1508] Step 18:
[1509] The device provides the user with the latest route information through speech synthesis. The input is the latest route information, which is converted into speech using the speech synthesis engine. The output is a voice notification.
[1510] (Application Example 2)
[1511] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1512] While conventional autonomous driving systems supported user authentication and destination setting, they did not provide driving support based on the user's emotional state. Furthermore, although they offered real-time optimal route calculations and voice-based information, they did not address the user's emotional needs. As a result, while they provided a safe and efficient driving environment, they were insufficient to fully deliver a comfortable driving experience for the user.
[1513] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for authenticating the user's personal information, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generation AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, and means for recognizing the user's emotional state using an emotion engine and providing support based on this. This makes it possible to respond to the user's emotional needs and to provide an even more comfortable and personalized driving experience.
[1514] "User personal authentication" is a method of recognizing the user's face using a camera installed in the vehicle and matching it with data in the cloud.
[1515] "Speech recognition" is a method of converting a user's voice into text data, allowing the system to accept user commands.
[1516] "Generative AI" is a method that, based on the user's voice instructions, retrieves schedule information from cloud calendars and other sources and suggests destinations.
[1517] "IoT data" refers to data such as speed, fuel level, and location information collected from vehicle sensors.
[1518] "A method for calculating the optimal route in real time" refers to a method that uses IoT data to calculate the optimal route by comparing it with traffic and weather information.
[1519] "Speech synthesis" is a method of providing users with calculation results or suggested messages in voice.
[1520] An "emotion engine" is a method of recognizing a user's emotional state by analyzing their voice and facial expressions.
[1521] "Means of providing support" refers to methods of proposing solutions that meet the user's emotional needs based on the analysis results of the emotion engine.
[1522] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This system not only realizes a safe and efficient driving environment but also addresses the user's emotional needs, providing a more comfortable and personalized driving experience.
[1523] User authentication
[1524] The server captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the cloud server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the vehicle terminal. The terminal notifies the user by voice, "Hello, user. Where are you going today?"
[1525] Setting a destination
[1526] The user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the vehicle terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9:00 today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1527] Navigation and real-time support
[1528] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[1529] How the emotion engine works
[1530] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[1531] Parking reservation support
[1532] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[1533] Examples of specific cases and prompt statements
[1534] For example, if a user asks, "What's on my schedule today?", the system retrieves information from a cloud calendar and suggests, "You have a meeting at the office at 9:00." Also, if the user is feeling stressed, the system might suggest, "Would you like some relaxing music?"
[1535] Example of a prompt:
[1536] What are your plans for today? Please tell me your schedule.
[1537]
[1538] Please suggest ways to address situations where users are experiencing stress.
[1539] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1540] Step 1:
[1541] The camera is activated and the user's face is captured. For authentication, the facial image acquired from the camera is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. The input is facial image data, and the output is the authentication result. Specifically, facial recognition is performed using OpenCV, and if authentication is successful, the corresponding user information is sent back to the terminal.
[1542] Step 2:
[1543] The device notifies the user via speech synthesis, "Hello, user. Where are you going today?" The user's voice response is captured by the microphone and converted into text data using a speech recognition system. This converted text data is the input, and the speech recognition result is the output. Specifically, the Google Cloud Speech-to-Text API is used to convert speech to text.
[1544] Step 3:
[1545] When a user asks "What's on my schedule today?", a generative AI retrieves the user's schedule information. The generative AI obtains the user's schedule from a cloud calendar and other schedule information and generates destination suggestions. The input to this process is the user's voice command "What's on my schedule today?", and the output is destination suggestions. Specifically, it uses the OpenAI GPT model to obtain schedule information based on the prompt sentence and generates suggestion messages.
[1546] Step 4:
[1547] The device synthesizes suggested messages into speech and notifies the user. For example, it might read aloud the message, "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is a suggested message from the generation AI, and the output is a synthesized speech notification.
[1548] Step 5:
[1549] Once the user confirms the destination setting, the navigation system activates and begins calculating the optimal route. The device collects IoT data such as speed, fuel level, and location information in real time from sensors and sends it to the server. The server compares this data with traffic and weather information and recalculates the optimal route. The input is IoT data and real-time traffic and weather information, and the output is instructions for the optimal route.
[1550] Step 6:
[1551] The device uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotional state. The input is the user's voice and facial expression data, and the output is the result of the emotional state analysis. Specifically, it uses voice and video analysis technologies to analyze the user's emotions.
[1552] Step 7:
[1553] The server uses generative AI to suggest countermeasures based on the analysis results from the emotion engine. For example, if the user is feeling stressed, it might suggest, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?" The input is the analysis results from the emotion engine, and the output is the suggested countermeasures. Specifically, the generative AI makes suggestions based on the user's emotional needs.
[1554] Step 8:
[1555] When a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This is converted into text data by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It then selects a suitable parking lot and prompts the user to make a reservation. The input is the user's voice command, and the output is parking availability information and reservation confirmation.
[1556] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1557] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1558] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1559] [Fourth Embodiment]
[1560] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1561] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1562] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1563] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1564] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1565] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1566] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1567] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1568] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1569] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1570] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1571] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1572] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1573] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is programmed according to the following procedure.
[1574] User authentication
[1575] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[1576] Setting a destination
[1577] The user asks the system by voice, "What's on my schedule today?" This voice command is converted into text data by the terminal's voice recognition system. The text data is sent to a generating AI, and the server retrieves the user's schedule from the cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1578] Navigation and real-time support
[1579] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[1580] Transaction processing
[1581] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[1582] summary
[1583] Through the above processes, users can efficiently navigate their journey to their destination while receiving personalized support. This system utilizes voice recognition, generative AI, and real-time data collection and analysis to provide the latest and most relevant information, ensuring a comfortable user experience for both drivers and passengers.
[1584] The following describes the processing flow.
[1585] Specific processing steps for carrying out the invention
[1586] User authentication
[1587] Step 1:
[1588] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[1589] Step 2:
[1590] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[1591] Step 3:
[1592] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[1593] Step 4:
[1594] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[1595] Step 5:
[1596] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[1597] Setting a destination
[1598] Step 1:
[1599] User: "What's on the schedule for today?" asks the system by voice.
[1600] Step 2:
[1601] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[1602] Step 3:
[1603] Terminal: Sends the converted text data to the generating AI.
[1604] Step 4:
[1605] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[1606] Step 5:
[1607] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[1608] Step 6:
[1609] Terminal: The generated message is read aloud to the user using speech synthesis.
[1610] Step 7:
[1611] User: "Yes, I'm going to the office," they reply, confirming the destination.
[1612] Navigation and real-time support
[1613] Step 1:
[1614] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[1615] Step 2:
[1616] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[1617] Step 3:
[1618] Terminal: Sends collected data to the server in real time.
[1619] Step 4:
[1620] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[1621] Step 5:
[1622] Server: Sends updated routes and additional information to the terminal.
[1623] Step 6:
[1624] Terminal: Notifies the user of updated information using speech synthesis.
[1625] Step 7:
[1626] User: Drive according to instructions based on traffic congestion and accident information.
[1627] Transaction processing
[1628] Step 1:
[1629] User: "Find a nearby parking lot," gives a voice command.
[1630] Step 2:
[1631] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[1632] Step 3:
[1633] Terminal: Sends the converted text data to the server.
[1634] Step 4:
[1635] Server: Searches for the nearest parking lot and checks its availability.
[1636] Step 5:
[1637] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[1638] Step 6:
[1639] Terminal: The suggested message is read aloud to the user using speech synthesis.
[1640] Step 7:
[1641] User: Responds with "Yes" and instructs to reserve a parking space.
[1642] Step 8:
[1643] Server: Executes parking reservations through the system.
[1644] Step 9:
[1645] Server: Sends a reservation completion notification to the device.
[1646] Step 10:
[1647] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[1648] In this way, users can arrive at their destination comfortably and efficiently through a series of processes. This allows them to spend their time in the car more meaningfully.
[1649] (Example 1)
[1650] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1651] Conventional vehicle driver assistance systems are increasingly expected to offer advanced functions beyond user authentication and voice recognition-based command acceptance, such as real-time information provision, schedule management, and parking reservation. However, few systems provide these functions in an integrated manner, resulting in limited improvements to the user experience. Furthermore, real-time data analysis and navigation optimization have been insufficient, making it difficult to provide a safe and efficient driving environment.
[1652] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1653] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands via voice recognition, means for acquiring user schedule information and suggesting destinations using generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user via speech synthesis, means for activating a navigation system and recalculating and providing the route in real time, means for collecting and analyzing speed, fuel level, and location information from vehicle sensors, and means for collecting parking information and executing reservations. This enables an improved user experience and the realization of a safe and efficient driving environment.
[1654] "User authentication" is a process that uses cameras mounted on the vehicle to capture the user's face and compares the encrypted facial data with a database in the cloud.
[1655] "Voice recognition" is a technology that converts the voice spoken by a user inside a vehicle into text data, which the system then accepts as a command it can understand.
[1656] "Generative AI" is an algorithm that uses artificial intelligence technology to analyze a user's schedule information and suggest appropriate destinations and actions.
[1657] "IoT data" refers to real-time data such as speed, fuel level, and location information collected from vehicle sensors.
[1658] "Calculating the optimal route in real time" is a process that instantly analyzes the most efficient travel route based on collected IoT data, external traffic information, and weather information.
[1659] "Speech synthesis" is a technology that converts text data generated by a system into speech and provides information to the user.
[1660] "Activating the navigation system" means activating the vehicle's navigation system to begin providing optimal route guidance to the set destination.
[1661] "Recalculating and providing routes" means re-analyzing the optimal route based on real-time data during travel and providing the user with the latest route information.
[1662] "Vehicle sensors" are devices used to measure various conditions inside and around the vehicle, collecting information such as speed, fuel level, and location.
[1663] "Gathering parking information" is the process of finding available parking spaces and their availability near your current location.
[1664] "Executing a reservation" refers to the process of reserving a suitable parking space for the user based on the collected parking information.
[1665] This invention is a system that authenticates users and provides personalized support to drivers and passengers using voice recognition, generative AI, IoT data collection, and real-time optimal route calculation. This enables a safer and more efficient driving environment and improves the user experience. The system is implemented using the following hardware and software:
[1666] hardware
[1667] Camera: Mounted in the vehicle and used to capture the user's face.
[1668] Vehicle sensors: Used to collect speed, fuel level, and location information in real time.
[1669] Terminal: A computer device installed inside a vehicle that runs voice recognition systems, generative AI, and navigation systems.
[1670] Server: Located in a cloud environment, it performs facial recognition data matching, IoT data analysis, and acquires schedule information and calculates routes using generated AI.
[1671] software
[1672] Speech recognition system: Converts user speech into text data.
[1673] Generative AI model: Analyzes user schedule information and suggests appropriate destinations and actions.
[1674] Navigation system: Calculates the optimal route and guides the user.
[1675] Encryption technology: Used to securely transmit user facial data to the cloud.
[1676] This system is implemented in the following steps:
[1677] 1. User authentication:
[1678] The terminal uses a camera mounted on the vehicle to capture the user's face. The authentication process starts automatically when the engine is started. The acquired facial data is encrypted and sent to a server. The server compares the facial data with a database in the cloud and sends the authentication result back to the terminal. The terminal then notifies the user by voice, "Hello, [username]. Where are you going today?"
[1679] 2. Setting the destination:
[1680] The user asks aloud, "What's on my schedule today?" A speech recognition system converts the speech into text data, which is then sent to a generating AI. The server retrieves the user's schedule from the cloud calendar and generates a message saying, "You have a meeting at the office at 9am today. Shall we head to the office?" The device reads this message aloud. When the user responds, "Yes, I'll go to the office," the destination is set.
[1681] 3. Navigation and real-time support:
[1682] The terminal activates the navigation system and begins calculating the optimal route. IoT data such as speed, fuel level, and location information collected from the vehicle's sensors is transmitted to the server in real time. The server analyzes this data and recalculates the optimal route by comparing it with traffic and weather information. The latest route information is sent to the terminal and provided to the user via speech synthesis. For example, instructions such as, "We have calculated the optimal route to avoid traffic congestion. Please turn right," are provided.
[1683] 4. Transaction processing:
[1684] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." This is converted into text by a voice recognition system and sent to a server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message, "The nearest parking lot is XX. Would you like to make a reservation?" which the device reads aloud. If the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. Afterward, reservation confirmation information is sent to the device and notified by voice.
[1685] Examples of prompt statements
[1686] The following are specific examples of prompt statements to be input to a generative AI model:
[1687] 1. "Tell me your plans for today."
[1688] 2. "Calculate the optimal route."
[1689] 3. "Please find a nearby parking lot."
[1690] By using these prompts, users can utilize the system more smoothly and intuitively.
[1691] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1692] Processing steps
[1693] Step 1:
[1694] Input: User's face data (image)
[1695] Specific operation: The device uses the vehicle's onboard camera to automatically capture the user's face when the engine starts. This image data is then stored in the device.
[1696] Data processing: Acquired facial data is encrypted on the device.
[1697] Output: Encrypted facial data is generated.
[1698] Step 2:
[1699] Input: Encrypted facial data
[1700] Specific operation: The device sends encrypted facial data to the server via a secure communication channel.
[1701] Data processing: The server compares the received facial data with user information stored in a database on the cloud.
[1702] Output: An authentication result (success or failure) is generated.
[1703] Step 3:
[1704] Input: Authentication result
[1705] Specific action: The server sends the authentication result back to the terminal.
[1706] Data processing: If successful, corresponding user information will be attached.
[1707] Output: The authentication result (and user information) is sent to the terminal.
[1708] Step 4:
[1709] Input: Authentication result
[1710] Specific operation: The device receives the authentication result, and if it contains user information, it notifies the user via voice, "Hello, [username]. Where are you going today?"
[1711] Data processing: Convert text to speech using speech synthesis technology.
[1712] Output: An audio notification is generated for the user.
[1713] Step 5:
[1714] Input: User voice input ("What are my plans for today?")
[1715] Specific action: The user asks aloud, "What's on the schedule for today?"
[1716] Data processing: The speech recognition system converts the audio into text data.
[1717] Output: Text data ("What are your plans for today?") is generated.
[1718] Step 6:
[1719] Input: Text data ("What are your plans for today?")
[1720] Specific action: This text data is sent to the generating AI.
[1721] Data processing: The server uses generative AI to retrieve user appointments from cloud calendars and scheduling information.
[1722] Output: User schedule information is generated.
[1723] Step 7:
[1724] Input: User's schedule information
[1725] Specific operation: Based on the acquired information, the server generates a suggestion message such as, "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[1726] Data calculation: Text message generation
[1727] Output: A suggestion message is generated.
[1728] Step 8:
[1729] Input: Suggestion message
[1730] Specific action: The terminal reads out the suggestion message.
[1731] Data processing: Text is converted to speech using speech synthesis technology.
[1732] Output: An audio notification is generated for the user.
[1733] Step 9:
[1734] Input: User voice input ("Yes, I'm going to the office.")
[1735] Specific action: The user responds, "Yes, I will go to the office."
[1736] Data processing: The speech recognition system converts the audio into text data.
[1737] Output: Text data ("Yes, I will go to the office") is generated.
[1738] Step 10:
[1739] Input: Text data ("Yes, I will go to the office")
[1740] Specific action: The terminal receives this text data and sets the office as the destination.
[1741] Data calculation: Setting destination information
[1742] Output: Destination setting complete.
[1743] Step 11:
[1744] Input: Destination information
[1745] Specific action: The terminal activates the navigation system and begins calculating the optimal route.
[1746] Data calculation: Calculation of the optimal route
[1747] Output: A navigation route is generated.
[1748] Step 12:
[1749] Input: Vehicle sensor data (speed, fuel level, location information)
[1750] Specific operation: The terminal collects data from the vehicle's sensors in real time and sends it to the server.
[1751] Data processing: The server analyzes the collected IoT data and compares it with traffic and weather information.
[1752] Output: Analysis results are generated.
[1753] Step 13:
[1754] Input: Analysis results
[1755] Specific operation: The server recalculates the optimal route based on the analysis results and sends it to the terminal.
[1756] Data calculation: Recalculating navigation routes
[1757] Output: The recalculated navigation route is sent to the terminal.
[1758] Step 14:
[1759] Input: Recalculated navigation route
[1760] Specific operation: The terminal provides the user with the latest recalculated route information through speech synthesis.
[1761] Data processing: Text is converted to speech using speech synthesis technology.
[1762] Output: An audio notification is generated for the user.
[1763] Step 15:
[1764] Input: User voice input ("Find a nearby parking lot")
[1765] Specific action: The user reaches near their destination and gives a voice command saying, "Find a nearby parking lot."
[1766] Data processing: The speech recognition system converts the audio into text data.
[1767] Output: Text data ("Find a nearby parking lot") is generated.
[1768] Step 16:
[1769] Input: Text data ("Find nearby parking")
[1770] Specific operation: The server receives text data and collects information about the nearest parking lot.
[1771] Data processing: The server checks the availability of parking spaces and selects a suitable one.
[1772] Output: A message recommending parking is generated.
[1773] Step 17:
[1774] Input: Recommended parking message
[1775] Specific action: The device reads out a recommendation message and notifies the user, "The nearest parking lot is XX. Would you like to make a reservation?"
[1776] Data processing: Text is converted to speech using speech synthesis technology.
[1777] Output: An audio notification is generated for the user.
[1778] Step 18:
[1779] Input: User voice input ("Yes")
[1780] Specific action: When the user responds with "yes," the server accesses the parking reservation system and completes the reservation.
[1781] Data processing: Generating confirmation information for parking reservations
[1782] Output: Reservation confirmation information is generated and sent to the terminal.
[1783] Step 19:
[1784] Input: Reservation confirmation information
[1785] Specific operation: The device receives reservation confirmation information and notifies the user via voice.
[1786] Data processing: Text is converted to speech using speech synthesis technology.
[1787] Output: An audio notification is generated for the user.
[1788] Through these steps, users can travel to their destination comfortably and efficiently while receiving personalized support. The system as a whole utilizes speech recognition, generative AI, and real-time data collection and analysis technologies.
[1789] (Application Example 1)
[1790] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1791] Current autonomous vehicle systems face numerous challenges in improving the user experience, particularly in areas such as user authentication, destination setting, and real-time route optimization. For example, issues include difficulties with smooth facial recognition and cumbersome destination setting. Furthermore, a lack of proper integration of real-time optimal route calculations during driving and parking reservation information near the destination significantly reduces user convenience. Addressing these challenges requires advanced speech recognition, generative AI, and the integration of IoT data.
[1792] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1793] In this invention, the server includes means for capturing the user's face and performing facial recognition, means for transmitting real-time vehicle sensor data to the server and analyzing it, and means for collecting parking information near the destination, checking availability, and proposing and completing reservations. This makes it possible to smoothly perform everything from user authentication to destination setting, navigation, real-time route optimization, and parking reservation.
[1794] "User authentication" refers to the authentication process used when a user logs into a system, utilizing personal characteristics such as their face or voice.
[1795] "Speech recognition" is a technology that allows a machine to understand what a user says and convert it into text data.
[1796] "Generative AI" is an artificial intelligence technology that generates new data or answers based on pre-trained data.
[1797] "IoT data" refers to real-time information collected from various sensors and devices.
[1798] "Calculating the optimal route" means determining the most efficient route to the destination, taking into account current traffic conditions and road congestion.
[1799] "Speech synthesis" is a technology that converts text data into speech and provides information to the user as audio.
[1800] A "smartphone" is an electronic device that combines the functions of a mobile phone and a computer, enabling internet connectivity and application usage.
[1801] "Vehicle sensor data" refers to data such as speed, fuel level, and location information measured by sensors installed in the vehicle.
[1802] "Parking information" refers to information such as the location, availability, and fees of the parking lot.
[1803] A "reservation" is a procedure aimed at securing goods or services in advance.
[1804] "Authentication" refers to the process of verifying that someone is a person.
[1805] A "proposal" is an expression of an idea or plan regarding a particular subject.
[1806] "Analysis" refers to the process of analyzing collected data in detail.
[1807] This invention is a system that provides personalized support for autonomous vehicles, including user facial recognition, voice recognition, acquisition of schedule information using generative AI, real-time optimal route calculation, and parking reservation.
[1808] Hardware and software configuration
[1809] Hardware: Smartphones, cameras, in-vehicle sensors (speed sensors, fuel sensors, GPS, etc.)
[1810] Software: OpenCV, face_recognition, speech_recognition, cloud server, generative AI model
[1811] This system uses the user's smartphone camera to perform facial recognition before the user enters the vehicle. The facial data captured by the camera is encrypted and sent to a cloud server for facial recognition. If facial recognition is successful, the server retrieves the corresponding user's information and sends it back to the device.
[1812] Users set destinations and give other instructions using voice commands. These voice commands are received through the smartphone's microphone and converted into text data by speech_recognition software. The text data is sent to a generative AI model, which suggests appropriate destinations based on the user's schedule information. The generative AI model retrieves the user's schedule from sources such as cloud calendars and generates the optimal route and action plan.
[1813] Once a destination is set, data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors and sent to a cloud server. Based on the collected data, the server calculates the optimal route in real time and sends the result back to the terminal. The terminal uses speech synthesis to provide the user with the latest route information in real time. For example, it may provide specific instructions such as, "We have calculated the optimal route to avoid traffic. Please turn right."
[1814] When the user arrives near their destination, they instruct the terminal by voice saying, "Find a nearby parking lot." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a suggestion message such as, "The nearest parking lot is XX. Would you like to make a reservation?" If the user responds with "Yes," the server accesses the parking lot reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and the user is notified by voice.
[1815] Examples of prompt statements
[1816] "Please suggest destinations based on the user's schedule."
[1817] "Please calculate the optimal route to avoid traffic congestion."
[1818] "Please provide information on the availability and reservation status of nearby parking lots."
[1819] This system allows users to enjoy efficient and safe driving while receiving personalized support.
[1820] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1821] Step 1:
[1822] The user approaches the vehicle using their smartphone. The device (smartphone) activates its camera and captures the user's face. This image data is processed using a facial recognition framework (e.g., OpenCV and face_recognition). Facial recognition extracts facial feature points, which are then encrypted. The encrypted data is sent to a cloud server.
[1823] Input: Smartphone camera image
[1824] Data processing: Extraction and encryption of facial feature points.
[1825] Output: Encrypted facial data
[1826] Step 2:
[1827] The server authenticates the user by comparing the received encrypted facial data with a database in the cloud. If authentication is successful, the user's individual information is sent back from the server to the device. This completes the user authentication process.
[1828] Input: Encrypted facial data
[1829] Data processing: Database matching
[1830] Output: Authentication results and user information
[1831] Step 3:
[1832] The user gives voice commands to the device. The device receives the user's voice via its microphone and converts it into text data using the speech_recognition library. This text data is then sent to a cloud server.
[1833] Input: Audio data
[1834] Data processing: Speech-to-text conversion
[1835] Output: Text data of voice commands
[1836] Step 4:
[1837] The server analyzes the received text data based on a generative AI model. The generative AI model retrieves the user's schedule information from a cloud calendar and other sources, and suggests destinations. These suggestions are sent to the terminal and notified to the user via speech synthesis.
[1838] Input: Text data of voice commands
[1839] Data processing: Schedule acquisition and destination suggestion using generative AI models.
[1840] Output: Destination suggestion message
[1841] Step 5:
[1842] The user accepts the destination suggested by voice. For example, they might respond by saying, "Yes, I'm going to the office." This response is then converted back into text data by speech_recognition and sent to the server.
[1843] Input: Voice response
[1844] Data processing: Speech-to-text conversion
[1845] Output: Text data of the response
[1846] Step 6:
[1847] The server analyzes the received text data and sets the destination. Furthermore, real-time data from vehicle sensors (speed, fuel level, location) is collected and sent to the server. Based on this data, the server calculates the optimal route and sends the result to the terminal.
[1848] Input: Response text data and sensor data
[1849] Data processing: Calculation of the optimal route
[1850] Output: Optimal route information
[1851] Step 7:
[1852] The terminal uses speech synthesis to guide the user to the optimal route. For example, it might provide instructions such as, "The optimal route has been set. Please turn right at the next intersection."
[1853] Input: Optimal route information
[1854] Data processing: Speech synthesis
[1855] Output: Voice guidance
[1856] Step 8:
[1857] When the user approaches their destination, they will voice-instruct the device to "find a nearby parking lot." This instruction is also converted into text data via speech_recognition and sent to the server.
[1858] Input: Audio data for parking instructions
[1859] Data processing: Speech-to-text conversion
[1860] Output: Text data of parking instructions
[1861] Step 9:
[1862] The server analyzes the text data of the parking instructions and collects information about nearby parking lots. It checks availability and sends a suggestion message to the terminal indicating a suitable parking lot. A message such as "The nearest parking lot is XX. Would you like to make a reservation?" is generated.
[1863] Input: Text data for parking instructions
[1864] Data processing: Obtaining parking information and checking availability.
[1865] Output: Parking suggestion message
[1866] Step 10:
[1867] When the user responds with "yes," the terminal converts the response into text using its speech recognition system and sends it to the server. The server accesses the parking reservation system and completes the reservation. Reservation confirmation information is sent to the terminal and notified to the user via voice.
[1868] Input: Audio data of parking reservation response
[1869] Data processing: Text conversion and reservation processing of responses
[1870] Output: Reservation confirmation information
[1871] Through the processing steps described above, this system provides consistent support from user facial recognition to parking reservation, offering a safe and efficient driving environment.
[1872] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1873] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[1874] User authentication
[1875] The terminal captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the terminal. The terminal then notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[1876] Setting a destination
[1877] The user asks the system by voice, "What's on my schedule today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to a generating AI, which retrieves the user's schedule from a cloud calendar and other scheduling information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1878] Navigation and real-time support
[1879] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[1880] How the emotion engine works
[1881] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[1882] Transaction processing
[1883] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[1884] summary
[1885] Through the above processes, users can enjoy an efficient and comfortable journey to their destination while receiving personalized support. This system can further enhance the user experience by recognizing and appropriately responding to emotional states, in addition to voice recognition, generative AI, and real-time data collection and analysis.
[1886] The following describes the processing flow.
[1887] Specific processing steps for carrying out the invention
[1888] User authentication
[1889] Step 1:
[1890] Terminal: Simultaneously with the vehicle's engine starting, the on-board camera captures the user's face.
[1891] Step 2:
[1892] Terminal: Encrypts the captured facial image and sends it to the cloud server.
[1893] Step 3:
[1894] Server: The server compares facial image data received on the cloud with the user database and performs authentication.
[1895] Step 4:
[1896] Server: If authentication is successful, the server sends the authentication result, including the user's associated data, back to the terminal.
[1897] Step 5:
[1898] Terminal: Receives the authentication result and notifies the user by voice, "Hello, Mr. / Ms. Yamada. Where are you going today?"
[1899] Setting a destination
[1900] Step 1:
[1901] User: "What's on the schedule for today?" asks the system by voice.
[1902] Step 2:
[1903] Terminal: The speech recognition system analyzes the user's voice and converts it into text data.
[1904] Step 3:
[1905] Terminal: Sends the converted text data to the generating AI.
[1906] Step 4:
[1907] Server: The generating AI references the user's schedule from cloud calendars and scheduling management systems.
[1908] Step 5:
[1909] Server: Based on the schedule, it generates a suggestion message: "There is a meeting at the office at 9:00 today. Shall we head to the office?"
[1910] Step 6:
[1911] Terminal: The generated message is read aloud to the user using speech synthesis.
[1912] Step 7:
[1913] User: "Yes, I'm going to the office," they reply, confirming the destination.
[1914] Navigation and real-time support
[1915] Step 1:
[1916] Terminal: Once the destination is determined, the navigation system activates and calculates the optimal route.
[1917] Step 2:
[1918] Terminal: Collects IoT data such as speed, fuel level, and location information from vehicle sensors.
[1919] Step 3:
[1920] Terminal: Sends collected data to the server in real time.
[1921] Step 4:
[1922] Server: Analyzes received data, compares it with traffic conditions and weather information, and recalculates the optimal route.
[1923] Step 5:
[1924] Server: Sends updated routes and additional information to the terminal.
[1925] Step 6:
[1926] Terminal: Notifies the user of updated information using speech synthesis.
[1927] Step 7:
[1928] User: Drive according to instructions based on traffic congestion and accident information.
[1929] How the emotion engine works
[1930] Step 1:
[1931] Terminal: Uses the vehicle's camera and microphone to capture the user's facial expressions and voice tone in real time.
[1932] Step 2:
[1933] Terminal: Sends captured data to the emotion engine.
[1934] Step 3:
[1935] Server: The emotion engine analyzes the data and determines the user's emotional state.
[1936] Step 4:
[1937] Server: Based on the user's emotional state, the generating AI proposes appropriate countermeasures.
[1938] Step 5:
[1939] Device: The generating AI reads aloud messages and suggestions tailored to the user's emotions using speech synthesis. For example, if the user is feeling stressed, it might suggest, "Would you like to play some music to relax?"
[1940] Transaction processing
[1941] Step 1:
[1942] User: "Find a nearby parking lot," gives a voice command.
[1943] Step 2:
[1944] Terminal: Analyzes instructions using a voice recognition system and converts them into text data.
[1945] Step 3:
[1946] Terminal: Sends the converted text data to the server.
[1947] Step 4:
[1948] Server: Searches for the nearest parking lot and checks its availability.
[1949] Step 5:
[1950] Server: Generates information about the selected parking lot and a suggestion message: "The nearest parking lot is XX. Would you like to make a reservation?"
[1951] Step 6:
[1952] Terminal: The suggested message is read aloud to the user using speech synthesis.
[1953] Step 7:
[1954] User: Responds with "Yes" and instructs to reserve a parking space.
[1955] Step 8:
[1956] Server: Executes parking reservations through the system.
[1957] Step 9:
[1958] Server: Sends a reservation completion notification to the device.
[1959] Step 10:
[1960] Terminal: Notifies the user of the reservation completion via voice and displays the reservation information.
[1961] Through the processing flow described above, users can enjoy a comfortable driving environment that takes their emotional state into consideration, while also receiving individually personalized support.
[1962] (Example 2)
[1963] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1964] Conventional driver assistance systems typically only offer navigation and voice recognition functions, lacking features to address the individual needs and emotional states of users. As a result, while safe and efficient driver assistance may be achievable, providing emotional satisfaction and a personalized experience for users remains a challenge.
[1965] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1966] In this invention, the server includes means for authenticating the user's identity, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generative AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, means using an emotion engine to analyze the user's emotional state, and means for the generative AI to make suggestions based on the emotional state. This makes it possible to respond to the user's individual needs and improve emotional satisfaction.
[1967] "Means of authenticating a user's identity" refers to a function that uses sensors such as cameras to acquire the user's biometric information and verifies the user's identity based on that information.
[1968] "A means of receiving user commands through voice recognition" refers to a technology that uses acoustic devices such as microphones to capture user voice instructions, analyzes them, and converts them into text data or commands.
[1969] "A method for acquiring user schedule information and suggesting destinations using generative AI" refers to a system that uses artificial intelligence to access user schedule data and suggests the optimal destination based on that data.
[1970] "A method for collecting IoT data and calculating the optimal route in real time" refers to a technology that analyzes traffic conditions and road information based on data collected from various sensors in a vehicle, and calculates the optimal route in real time.
[1971] "A means of providing information to users through speech synthesis" refers to a technology that converts text data into speech output and notifies users of necessary information via voice.
[1972] "Methods using an emotion engine to analyze the user's emotional state" refers to technologies that analyze the user's voice and facial expressions to estimate their emotional state and take appropriate action.
[1973] "A means by which AI generates suggestions based on emotional state" refers to a function in which the generating AI takes into account the user's emotional state and provides the user with appropriate responses and suggestions.
[1974] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This not only realizes a safe and efficient driving environment but also addresses the user's emotional needs and provides a more comfortable and personalized experience.
[1975] First, as a means of user authentication, the terminal uses a camera mounted on the vehicle. When the engine starts, facial recognition begins automatically, and the acquired facial data is encrypted and sent to the server. The server compares it with user information stored in the cloud and performs authentication. If authentication is successful, the server sends the authentication result, including the corresponding user's individual information, back to the terminal, and the terminal notifies the user by voice, "Hello, Mr. Yamada. Where are you going today?"
[1976] Next, regarding setting the destination, the user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9am today. Shall we head to the office?", which the terminal reads aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[1977] Once a destination is set, the terminal activates its navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, upon receiving traffic congestion information, it might instruct the user, "We have calculated the optimal route to avoid congestion. Please turn right."
[1978] Regarding the operation of the emotion engine, the agent system collects the user's voice and facial expressions using the camera and microphone and sends them to the emotion engine. The emotion engine analyzes this data to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to make suggestions. Specifically, if the user is expressing stress, the device will make suggestions via voice, such as, "Shall I play some music to help you relax?" or "Shall I find a nearby resting place?"
[1979] Regarding transaction processing, when a user arrives near their destination and needs to find parking, they instruct the terminal to "find nearby parking." This instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The terminal reads this message aloud. If the user responds with "yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the terminal and notified by voice.
[1980] As a concrete example of its operation, when a user gets into a car and asks, "What are my plans for today?", the system suggests, "I have a meeting at the office at 9:00," and navigation begins. Also, when near the destination, if the user says, "Find a nearby parking lot," the terminal converts the voice into text data, and the server searches for and reserves the nearest parking lot.
[1981] As an example of a prompt statement,
[1982] 1. "Please tell me about your appointments and plans for today."
[1983] 2. "Where are you planning to go today?"
[1984] There is.
[1985] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1986] Step 1:
[1987] The device uses a camera mounted on the vehicle to capture the user's face. When the user starts the engine, facial recognition begins automatically. The input data is the facial image captured by the camera, which is processed by a facial recognition algorithm to generate encrypted facial data. The output is the encrypted facial data.
[1988] Step 2:
[1989] The device sends encrypted facial data to the server. The input is encrypted facial data, and the output is the data sent to the server. Specifically, it uses an encrypted communication protocol to send the data.
[1990] Step 3:
[1991] The server performs authentication by comparing the received encrypted data with user information stored in the cloud. The input consists of encrypted facial data and facial data stored in the cloud, and a match is verified through a database search operation. The output is the authentication result.
[1992] Step 4:
[1993] If authentication is successful, the server returns an authentication result to the terminal, which includes the individual user information of the corresponding user. The input consists of the authentication result and user information, which are encrypted and sent to the terminal. The output is the authentication information received by the terminal.
[1994] Step 5:
[1995] Upon successful authentication, the device uses speech synthesis to notify the user, "Hello, Mr. / Ms. Yamada. Where are you going today?" The input is authentication information, which is converted from text to speech via a speech synthesis engine. The output is a voice notification.
[1996] Step 6:
[1997] The user asks the system, "What's on the schedule today?" using voice. The input is the user's voice command, which the terminal collects. The output is voice data.
[1998] Step 7:
[1999] The terminal analyzes voice commands using a speech recognition system and converts them into text data. The input is voice data, and the speech recognition engine converts the voice to text. The output is text data.
[2000] Step 8:
[2001] The device sends text data to the generating AI, and the server retrieves the user's schedule from a cloud calendar or other scheduling information. The input is text data, which is fed into the generating AI model. The output is the user's schedule information.
[2002] Step 9:
[2003] The server generates a suggestion message based on the acquired information: "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is schedule information, and a generation AI is used to generate the suggestion message. The output is the suggestion message.
[2004] Step 10:
[2005] The device notifies the user of this suggestion message via voice. The input is the suggestion message, which is converted from text to speech using a speech synthesis engine. The output is a voice notification.
[2006] Step 11:
[2007] The user responds by voice, "Yes, I'm going to the office," which sets the destination. The input is the user's voice response, which the terminal collects. The output is voice data.
[2008] Step 12:
[2009] The terminal uses a speech recognition system to convert the user's voice responses into text data and confirm the destination setting. The input is voice data, which is converted to text using the speech recognition engine. The output is text data.
[2010] Step 13:
[2011] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. The input is destination information, and the system uses a route calculation algorithm to generate the best route. The output is the optimal route information.
[2012] Step 14:
[2013] The terminal collects IoT data such as speed, fuel level, and location information from vehicle sensors in real time. The input is data from vehicle sensors, which is collected and analyzed. The output is real-time data.
[2014] Step 15:
[2015] The collected IoT data is sent to the server. The input is real-time data, and the output is data transmission to the server.
[2016] Step 16:
[2017] The server analyzes IoT data by cross-referencing it with traffic and weather information, and recalculates the optimal route. Inputs are real-time data and external information, and the optimal route is generated through data analysis. The output is the recalculated route information.
[2018] Step 17:
[2019] This function sends the latest route information to the terminal. The input is the recalculated route information, and the output is the data transmission to the terminal.
[2020] Step 18:
[2021] The device provides the user with the latest route information through speech synthesis. The input is the latest route information, which is converted into speech using the speech synthesis engine. The output is a voice notification.
[2022] (Application Example 2)
[2023] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2024] While conventional autonomous driving systems supported user authentication and destination setting, they did not provide driving support based on the user's emotional state. Furthermore, although they offered real-time optimal route calculations and voice-based information, they did not address the user's emotional needs. As a result, while they provided a safe and efficient driving environment, they were insufficient to fully deliver a comfortable driving experience for the user.
[2025] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for authenticating the user's personal information, means for receiving user commands by voice recognition, means for acquiring user schedule information and suggesting a destination using a generation AI, means for collecting IoT data and calculating the optimal route in real time, means for providing information to the user by speech synthesis, and means for recognizing the user's emotional state using an emotion engine and providing support based on this. This makes it possible to respond to the user's emotional needs and to provide an even more comfortable and personalized driving experience.
[2026] "User personal authentication" is a method of recognizing the user's face using a camera installed in the vehicle and matching it with data in the cloud.
[2027] "Speech recognition" is a method of converting a user's voice into text data, allowing the system to accept user commands.
[2028] "Generative AI" is a method that, based on the user's voice instructions, retrieves schedule information from cloud calendars and other sources and suggests destinations.
[2029] "IoT data" refers to data such as speed, fuel level, and location information collected from vehicle sensors.
[2030] "A method for calculating the optimal route in real time" refers to a method that uses IoT data to calculate the optimal route by comparing it with traffic and weather information.
[2031] "Speech synthesis" is a method of providing users with calculation results or suggested messages in voice.
[2032] An "emotion engine" is a method of recognizing a user's emotional state by analyzing their voice and facial expressions.
[2033] "Means of providing support" refers to methods of proposing solutions that meet the user's emotional needs based on the analysis results of the emotion engine.
[2034] This invention provides a system that, in addition to user authentication, voice recognition, generative AI, IoT data collection, and real-time optimal route calculation, recognizes the user's emotional state using an emotion engine and provides support based on that. This system not only realizes a safe and efficient driving environment but also addresses the user's emotional needs, providing a more comfortable and personalized driving experience.
[2035] User authentication
[2036] The server captures the user's face using a camera mounted on the vehicle. When the engine starts, facial recognition automatically begins, encrypting the acquired facial data and sending it to the cloud server. The server compares the encrypted data with user information stored in the cloud and performs authentication. If authentication is successful, the authentication result, including the corresponding user's individual information, is sent back to the vehicle terminal. The terminal notifies the user by voice, "Hello, user. Where are you going today?"
[2037] Setting a destination
[2038] The user asks the system by voice, "What's my schedule for today?" This voice command is analyzed by the vehicle terminal's voice recognition system and converted into text data. The text data is sent to the generating AI, and the server retrieves the user's schedule from the cloud calendar and other schedule information. Based on the retrieved information, the system generates a suggestion message, "You have a meeting at the office at 9:00 today. Shall we head to the office?" and the terminal reads it aloud. The user responds, "Yes, I'll go to the office," and the destination is set.
[2039] Navigation and real-time support
[2040] Once a destination is set, the terminal activates the navigation system and begins calculating the optimal route. Simultaneously, IoT data such as speed, fuel level, and location information is collected in real time from the vehicle's sensors. This data is sent to a server and cross-referenced with traffic and weather information. The server analyzes this data to recalculate the optimal route and sends it back to the terminal. The terminal continuously provides the user with the latest route information through speech synthesis. For example, if traffic congestion information is received, it will provide instructions such as, "We have calculated the optimal route to avoid congestion. Please turn right."
[2041] How the emotion engine works
[2042] This agent system analyzes the user's voice and facial expressions and uses an emotion engine to recognize the user's emotional state. For example, if the user is feeling stressed, the emotion engine will detect this state. Based on the analysis results of the emotion engine, the server uses generative AI to suggest countermeasures. Specifically, if the user is complaining of stress, the system will make suggestions such as, "Shall we play some music to help you relax?" or "Shall we find a nearby resting place?"
[2043] Parking reservation support
[2044] When a user arrives near their destination and needs to find parking, they instruct the device to "Find nearby parking." The instruction is converted into text data by a voice recognition system and sent to the server. The server collects information on the nearest parking lots and checks their availability. It selects a suitable parking lot and generates a message saying, "The nearest parking lot is XX. Would you like to make a reservation?" The device reads this aloud, and if the user responds with "Yes," the server accesses the parking reservation system and completes the reservation. After that, reservation confirmation information is sent to the device and notified by voice.
[2045] Examples of specific cases and prompt statements
[2046] For example, if a user asks, "What's on my schedule today?", the system retrieves information from a cloud calendar and suggests, "You have a meeting at the office at 9:00." Also, if the user is feeling stressed, the system might suggest, "Would you like some relaxing music?"
[2047] Example of a prompt:
[2048] What are your plans for today? Please tell me your schedule.
[2049]
[2050] Please suggest ways to address situations where users are experiencing stress.
[2051] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2052] Step 1:
[2053] The camera is activated and the user's face is captured. For authentication, the facial image acquired from the camera is encrypted and sent to the server. The server compares the encrypted data with user information stored in the cloud and performs authentication. The input is facial image data, and the output is the authentication result. Specifically, facial recognition is performed using OpenCV, and if authentication is successful, the corresponding user information is sent back to the terminal.
[2054] Step 2:
[2055] The device notifies the user via speech synthesis, "Hello, user. Where are you going today?" The user's voice response is captured by the microphone and converted into text data using a speech recognition system. This converted text data is the input, and the speech recognition result is the output. Specifically, the Google Cloud Speech-to-Text API is used to convert speech to text.
[2056] Step 3:
[2057] When a user asks "What's on my schedule today?", a generative AI retrieves the user's schedule information. The generative AI obtains the user's schedule from a cloud calendar and other schedule information and generates destination suggestions. The input to this process is the user's voice command "What's on my schedule today?", and the output is destination suggestions. Specifically, it uses the OpenAI GPT model to obtain schedule information based on the prompt sentence and generates suggestion messages.
[2058] Step 4:
[2059] The device synthesizes suggested messages into speech and notifies the user. For example, it might read aloud the message, "There is a meeting at the office at 9:00 today. Shall we head to the office?" The input is a suggested message from the generation AI, and the output is a synthesized speech notification.
[2060] Step 5:
[2061] Once the user confirms the destination setting, the navigation system activates and begins calculating the optimal route. The device collects IoT data such as speed, fuel level, and location information in real time from sensors and sends it to the server. The server compares this data with traffic and weather information and recalculates the optimal route. The input is IoT data and real-time traffic and weather information, and the output is instructions for the optimal route.
[2062] Step 6:
[2063] The device uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotional state. The input is the user's voice and facial expression data, and the output is the result of the emotional state analy...
Claims
1. A means of authenticating the user's identity, A means of receiving user commands through voice recognition, A method for obtaining user schedule information using generative AI and suggesting destinations, A means of collecting IoT data and calculating the optimal route in real time, A means of providing information to users through speech synthesis, A system that includes this.
2. The system according to claim 1, comprising means for authenticating a user's face using a camera and matching the authentication result with the cloud.
3. The system according to claim 1, further comprising means for a speech recognition system to convert speech into text data and a generating AI to retrieve appointments from a cloud calendar.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A