system
The system addresses the limitations of conventional navigation by using voice/text inputs, real-time data, and generative models to offer personalized and adaptive route guidance, ensuring efficient and stress-free travel.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Conventional car navigation systems struggle to provide personalized route suggestions that consider users' preferences and real-time traffic and weather conditions, lacking flexibility in accepting voice or text instructions and real-time route corrections.
A system that receives voice or text instructions, records user location and history, acquires real-time traffic and weather information, and uses a generative model to calculate and present optimal routes, allowing for flexible and personalized navigation with real-time updates.
Enables users to reach their destinations comfortably by providing individually optimized routes that adapt to changing conditions and user preferences, reducing travel stress and improving navigation efficiency.
Smart Images

Figure 2026070197000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional car navigation systems, there was a problem that it was difficult to propose a route that fully considered the user's personal preferences and situations. In particular, it was difficult to provide the optimal route for the user because it was impossible to utilize the user's past behavior history and comprehensively use real-time traffic and weather information to make a personalized route proposal. Furthermore, flexible acceptance of instructions in voice or text and real-time route correction were required.
Means for Solving the Problems
[0005] This invention provides means for receiving voice or text instructions from the user, thereby allowing for flexible setting of destinations and conditions. Furthermore, it includes means for recording and acquiring the user's current location information and past activity history, enabling more personalized navigation. It also has means for acquiring real-time traffic and weather information, and uses this data for a generative model to calculate the optimal route to the destination. This optimized route is presented to the user and also has a function to re-evaluate and update the route based on further voice instructions. This makes it possible to provide a flexible and optimal car navigation experience that is tailored to the user's actual situation and preferences.
[0006] A "user" refers to an individual or legal entity that uses the system to receive route guidance.
[0007] "Instructions" refer to information that users use to communicate destinations and route conditions to the system, and can be entered via voice or text.
[0008] "Current location information" refers to location data acquired by the user's device to indicate its current location.
[0009] "Behavioral history" refers to recorded information such as what routes a user has chosen in the past and what traffic conditions they have experienced.
[0010] "Traffic information" refers to information about road congestion, accidents, road construction, etc., which is obtained in real time.
[0011] "Weather information" refers to information about weather conditions such as rain, snow, and wind speed, which can potentially affect navigation.
[0012] A "generative model" is an algorithm or system that calculates the optimal path for a user based on acquired data.
[0013] The "optimal route" refers to the most efficient route to the destination, selected by the system based on the user's preferences, conditions, and real-time circumstances.
[0014] "Presentation" refers to the act of conveying calculated route information to the user visually or audibly.
[0015] "Re-evaluation" refers to the process of reviewing decisions regarding existing routes based on additional instructions from users and then optimizing them accordingly. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] As an embodiment of the present invention, a scenario is shown in which a user operates the car navigation system using a portable information terminal.
[0038] The user gets into the car and activates a portable information terminal. The terminal allows the user to specify the destination and driving conditions via voice or text input. For example, the user might voice-input "I want to avoid traffic jams." This input is analyzed by the terminal and sent to the server as text data.
[0039] The device obtains the user's current location information using GPS and references their past activity history. This information, along with real-time traffic and weather information, is sent to the server. The server integrates and analyzes this data and calculates the optimal route using a generative model.
[0040] The server calculates the optimal route, which is then sent to the terminal, where it is presented to the user visually and audibly. The information presented includes estimated travel time to the destination, distance, and potential stops along the way. Once the user approves the proposed route, the terminal begins navigation. During navigation, the terminal monitors traffic conditions in real time, re-evaluating and updating the route as needed, and providing the user with updated information.
[0041] As a concrete example, consider a situation where a user wants to go to a specific cafe, but the route shown on their car's navigation system is congested. The system receives a voice command saying, "I would like an alternative route that avoids the current congestion," and the server recalculates a new route that avoids the traffic. Based on this result, this route is sent back to the terminal and presented to the user. The user can then select this route and arrive at their destination cafe smoothly, avoiding the congestion.
[0042] Thus, the system of the present invention flexibly reflects user instructions and provides individually optimized routes, thereby supporting users in comfortably reaching their destinations.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user activates their mobile information terminal and inputs their destination and desired conditions via voice or text. For example, they might say, "I want to go to a nearby cafe to avoid traffic."
[0046] Step 2:
[0047] The device converts voice input into text data and uses natural language processing to analyze the user's intent. Specifically, it extracts destination categories and conditions (such as avoiding traffic jams).
[0048] Step 3:
[0049] The device uses GPS to obtain the user's current location and sends the location information, the user's past activity history, and the analyzed instructions to the server.
[0050] Step 4:
[0051] Based on the data received by the server, potential destinations near the current location are collected from a database and external APIs. This includes information such as nearby cafes.
[0052] Step 5:
[0053] The server acquires traffic and weather information in real time and uses a generative model to calculate the optimal route that reflects the user's desired conditions.
[0054] Step 6:
[0055] The server sends optimal route information to the terminal and creates multiple route options, including details such as travel time and distance.
[0056] Step 7:
[0057] The device presents route options to the user and prompts them to start navigation. Navigation begins based on the route selected by the user.
[0058] Step 8:
[0059] If the user gives a voice command such as "Find a shorter route" during navigation, the device will send a new command to the server.
[0060] Step 9:
[0061] The server receives additional instructions and recalculates the route, incorporating current real-time information. It then sends the new optimal route to the terminal.
[0062] Step 10:
[0063] The device presents the user with a recalculated route and updates the navigation based on the new route if selected.
[0064] (Example 1)
[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0066] In today's transportation environment, users often face difficulties in reaching their destinations smoothly. Conventional car navigation systems struggle to suggest optimal routes that take into account real-time traffic information and individual user preferences. Furthermore, they cannot reflect the user's past travel history or preferences when setting routes, which can result in inefficient route selection. There is a need for navigation systems that can solve these problems and enable users to reach their destinations comfortably.
[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0068] In this invention, the server includes voice analysis means and text analysis means for receiving instructions from the user, location estimation means for acquiring current location information and past behavioral history, and information acquisition means for acquiring traffic information and weather information in real time. This makes it possible to calculate and present the optimal route that takes into account the user's individual preferences and current traffic conditions.
[0069] "Voice analysis means" refers to a function or device for converting a user's voice instructions into digital data and analyzing the intended meaning.
[0070] "Text analysis means" refers to a function or device for analyzing text data entered by a user and understanding the content of the instructions.
[0071] "Location estimation means" refers to a function or device that uses GPS technology or similar to determine the user's current location and acquire their past activity history.
[0072] "Information acquisition means" refers to a function or device for collecting traffic information and weather information in real time from external information sources.
[0073] A "generative model based on deep learning" is a model that uses multi-layer neural network technology to analyze data and generate the optimal path.
[0074] A "route calculation means" is a function or device for calculating the optimal route to a destination based on acquired information.
[0075] "Presentation means" refers to a function or device for presenting calculated route information to the user visually or audibly.
[0076] A "re-evaluation means" is a function or device that receives additional instructions from the user, re-evaluates existing routes, and updates the routes as necessary.
[0077] The user utilizes a portable information terminal to implement the navigation system of the present invention. The user, upon entering the vehicle, activates the portable terminal and inputs the destination and driving conditions via voice or text. The terminal analyzes this user input using voice analysis software (e.g., a speech recognition API) and processes it as text. The analyzed text data is then linked with a location estimation means and transmitted to a server along with GPS information and past activity history.
[0078] The server integrates the received information and calculates the optimal route using a generative AI model based on deep learning. Real-time traffic and weather information is obtained from external information services (for example, map API services) and is also included in the analysis. Once the optimal route is calculated, the server sends this route information back to the terminal.
[0079] The route information sent is displayed on the terminal and guided to the user using speech synthesis software (e.g., a speech synthesis API). Navigation begins once the user confirms and approves the route. While the user is moving, the terminal and server communicate continuously, and if new traffic information is acquired, the route is re-evaluated based on that information, and updated information is provided to the user.
[0080] As a concrete example, a user might instruct the system by voice, "I want to avoid traffic." This input is analyzed by the server, and a new, optimized route is calculated by a generative AI model. This route is then presented to the user's device.
[0081] An example of a prompt for a generative AI model is, "If a user requests to avoid traffic congestion, please tell me how to calculate the optimal route considering the current traffic conditions." This enables flexible guidance tailored to the user's needs.
[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0083] Step 1:
[0084] The user activates their mobile device and inputs their destination and driving conditions via voice or text. The user's input is converted into text data by a voice analysis system. The converted text data is temporarily stored on the device. Here, input is voice or text, and output is text data for analysis. Voice analysis software (e.g., speech recognition API) is used to convert voice to text.
[0085] Step 2:
[0086] The device uses GPS technology to obtain the user's current location. Furthermore, it accesses and obtains the user's past activity history. The current location data and activity history data serve as input for transmission to the server. The output is integrated location data used by the server. Specifically, a location estimation system collects GPS data in real time and organizes it.
[0087] Step 3:
[0088] The server integrates text data sent from the terminal, current location, past activity history, and real-time traffic and weather information obtained via external APIs. The integrated data is input into a generative AI model. The output here is a dataset for predictive analysis by the generative AI model. The server retrieves necessary information from various external information provision services (e.g., map API services), structures the data, and prepares it for input into the model.
[0089] Step 4:
[0090] The generative AI model calculates the optimal path. This process involves simulating the optimal path based on an integrated dataset. The input is the integrated dataset, and the output is information about the optimal path. This step is performed on a server using deep learning techniques.
[0091] Step 5:
[0092] The server sends the calculated optimal route to the terminal. The terminal presents the received route information to the user using a map display application or speech synthesis software. The input is the optimal route information, and the output is navigation information presented to the user visually and audibly. The terminal provides information to the user through both visual map display and voice guidance.
[0093] Step 6:
[0094] Once the user approves the route, the device begins navigation. During navigation, the device continues to communicate with the server, re-evaluating the route based on changes in traffic and weather conditions, and updating it as needed. Input is real-time traffic information, and output is updated route information. The device receives real-time updated information and provides it to the user, correcting it as appropriate.
[0095] (Application Example 1)
[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] There is a need for a way for drivers of autonomous vehicles to intuitively obtain optimal route information to reach their destination. Furthermore, a challenge is to realize a system that provides drivers with visual and audio guidance while responding to real-time changes in traffic and weather conditions.
[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0099] In this invention, the server includes means for receiving instructions from the user via voice or text, means for recording and acquiring the user's current location information and past activity history, and means for acquiring traffic and weather information in real time. This allows the driver to visually confirm route information using an additional device and always reach their destination via the optimal route.
[0100] "Means for receiving user instructions via voice or text" refers to an interface for receiving user requests for destinations and routes via voice or text input and processing them within the system.
[0101] "Means for recording and acquiring a user's current location information and past behavioral history" refers to technologies that use GPS or similar methods to acquire a user's real-time location and record and refer to historical data such as places visited in the past.
[0102] "Means for obtaining real-time traffic and weather information" refers to technologies that continuously acquire the latest traffic conditions and weather data via a network.
[0103] "Means for calculating the optimal route to the destination based on the acquired information using a generative model" refers to a method of calculating the optimal travel route based on collected data using machine learning or AI technology.
[0104] "Means for presenting the calculated optimal route to the user" refers to technical means for providing route information to the user through displays, audio, etc.
[0105] "Means for re-evaluating and updating routes based on additional instructions from the user" refers to a method for recalculating and updating routes in real time in response to new route information and changes in conditions provided by the user.
[0106] "Means of providing route information as visual information through the driver's auxiliary device to assist route guidance" refers to methods of intuitively presenting route information to the driver visually using smart glasses or head-mounted displays.
[0107] As an embodiment of this invention, a navigation system for an autonomous vehicle utilizing a smart device is specifically described. The user can wear smart glasses equipped in the autonomous vehicle and obtain comfortable route guidance while visually checking traffic information in real time.
[0108] The server first receives voice or text input from the user and interprets the destination information and driving conditions. For this, Google® Speech-to-Text can be used as the speech recognition API. Next, the server obtains the user's current location using a GPS module and accesses past activity history stored on a cloud server. Furthermore, a network connection is required to collect traffic and weather information from the internet in real time.
[0109] Using a generative AI model (e.g., Google AI), the server calculates the optimal route from this information. The calculated route is then presented to the user as visual information through smart glasses. For example, if a user requests a route that is not congested, the server will present a new route that reflects real-time congestion. An example of a prompt in this case would be: "The user wants to go to a specific destination via a route that is not congested. Please provide the optimal route considering the current traffic conditions."
[0110] This system allows drivers to arrive at their destination via the optimal route at all times, while receiving visual and audible feedback. By flexibly responding to real-time changing traffic and weather conditions, it supports a comfortable travel experience for users.
[0111] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0112] Step 1:
[0113] The terminal starts up, and the user inputs the destination and driving conditions by voice or text. The input is received via the terminal's microphone or touchscreen. The terminal converts this into text data using a speech recognition API.
[0114] Step 2:
[0115] The device uses a GPS module to obtain the user's current location information. In addition, it retrieves past activity history from a cloud server and uses it as reference information for movement patterns and destinations. The input is current location information and past data, and the output is the user's location and associated history.
[0116] Step 3:
[0117] The server collects real-time traffic and weather information via the internet. This allows for monitoring of current traffic conditions and weather. Input is information from external APIs, and output is the latest traffic and weather data.
[0118] Step 4:
[0119] The server uses a generative AI model to calculate the optimal route based on user information and external data obtained so far. It utilizes prompts to personalize the route according to user requests. Specifically, it generates a route based on instructions such as, "The user wants to go to a specific destination via a 'congested' route." The output is optimized route information.
[0120] Step 5:
[0121] The calculated optimal route is transmitted to the terminal and presented to the user visually and audibly. The user can confirm their direction of travel through maps and route information displayed on smart glasses. The input is route information from the server, and the output is visual and audible guidance.
[0122] Step 6:
[0123] If the user provides additional instructions, the server re-evaluates the route and updates it as needed. This enables route guidance that adapts to dynamic changes in circumstances. The re-evaluated route is sent back to the terminal. The input is the user's new instructions, and the output is the updated route information.
[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0125] An embodiment of the present invention is shown below: a car navigation system incorporating an emotion engine. This system operates when a user receives route guidance using a portable information terminal.
[0126] The user activates a portable information terminal and provides instructions to the system regarding destination and route conditions via voice or text. The system incorporates an emotion engine that analyzes the user's emotions through voice input, allowing it to understand the user's emotional state based on the input voice data.
[0127] The emotion engine adjusts the generative model to prioritize scenic routes when the user is relaxed, for example, and to calculate the fastest route when the user is in a hurry or stressed. The server provides comprehensive route suggestions, taking into account the user's current location, past behavior history, and real-time traffic and weather information.
[0128] The calculated optimal route is transmitted to the terminal, which then presents this information to the user visually and audibly. The presentation method can also be adjusted according to the user's emotional state. For example, if the user is stressed, the information may be conveyed in a calm voice.
[0129] For example, if a user is in a hurry because they are running late for their scheduled departure time and gives the voice command, "I absolutely must avoid traffic," the system will recognize the urgency through its emotion engine, and the server will calculate a route that prioritizes speed. This calculated route information will be quickly presented to the user through the terminal, using a calm voice tone to reduce stress.
[0130] Thus, the system of the present invention, which incorporates an emotion engine, further personalizes the car navigation experience according to the user's psychological state, supporting more flexible and comfortable travel.
[0131] The following describes the processing flow.
[0132] Step 1:
[0133] The user activates their portable information terminal and inputs a voice command saying, "I want to go to my destination while avoiding the highway." At this time, the terminal records the voice data.
[0134] Step 2:
[0135] The device converts voice input into text data and simultaneously uses an emotion engine to analyze the user's emotional state from the tone and speed of their voice. The analysis result might be determined to indicate, for example, "the user is irritated."
[0136] Step 3:
[0137] The device obtains its current location information via GPS and collects data on the user's past behavior. This information, along with the results of sentiment analysis, is sent to the server.
[0138] Step 4:
[0139] Based on the server's received location information, the user's desired route conditions, and emotional state, real-time traffic and weather information is obtained.
[0140] Step 5:
[0141] The server uses a generative model to calculate the optimal route to the destination. This calculation takes into account the user's emotional state, prioritizing routes that, for example, allow for quick and stress-reducing travel.
[0142] Step 6:
[0143] The server calculates the optimal route and sends it to the terminal, then prepares detailed route information.
[0144] Step 7:
[0145] The device presents the route information it has received to the user. Depending on the user's emotional state, it provides guidance in a calm and reassuring voice, for example.
[0146] Step 8:
[0147] The user confirms the suggested route and begins navigation. If the user gives a new instruction during navigation, such as "I want to arrive sooner," the device sends this information back to the server.
[0148] Step 9:
[0149] The server receives the new instructions and, as in the previous step, re-evaluates the user's emotions and requests, and recalculates a more optimized path.
[0150] Step 10:
[0151] The device will suggest a new route, and if selected, it will immediately update the navigation to help the user have the most satisfying journey.
[0152] (Example 2)
[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0154] Conventional car navigation systems typically suggest the optimal route without considering the user's emotional state. This lack of flexible guidance, adapted to the user's psychological condition, can increase stress. Furthermore, they lack methods to adjust guidance based on emotional state, highlighting the need for improved user experience.
[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0156] In this invention, the server includes means for receiving instructions from the user in voice or text, means for analyzing the user's emotional state, and means for recording and acquiring the user's current location information and past behavioral history. This enables the personalization of the optimal route based on the user's emotional state, and route guidance that is appropriate to their psychological state.
[0157] "Means for receiving user instructions via voice or text" refers to a function that provides an interface for users to specify destination and route conditions to the system via voice or text.
[0158] "Means for analyzing a user's emotional state" refers to technologies or devices used to identify a user's emotional state from voice data or user interactions. This makes it possible to understand the user's psychological state.
[0159] "Means for recording and acquiring a user's current location information and past behavioral history" refers to a device or method for acquiring a user's location information in real time, storing it, and analyzing and utilizing past movement patterns.
[0160] "Means for obtaining real-time traffic and weather information" refers to technologies or services that collect the latest traffic and weather conditions from external sources and provide them to a system.
[0161] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a method or apparatus that uses AI technology or the like to calculate the most suitable travel route for the user based on the acquired information.
[0162] "A means of presenting a calculated optimal route to the user and adjusting the presented content to the user's emotional state" refers to a technology that presents calculated route information to the user using visual and auditory methods, and applies communication methods that correspond to the user's emotions at that time.
[0163] "Means for re-evaluating and updating routes based on additional instructions from users" refers to a mechanism that reviews existing routes and updates them to the optimal ones in response to new requests or changes in conditions from users.
[0164] The system based on this invention realizes an advanced car navigation system equipped with emotion analysis capabilities. The system uses the user's mobile device as its primary interface and accepts instructions from the user via voice or text.
[0165] The terminal uses hardware such as a voice input device and a touchscreen to acquire information about the user's destination and route conditions. The acquired voice data is sent to a server equipped with an emotion analysis engine. This server uses speech recognition software (e.g., a common speech recognition API) to analyze the user's emotional state from their voice. In this process, by analyzing features such as tone of voice, word choice, and speaking speed, the server identifies the user's emotional state, such as whether they are relaxed, in a hurry, or stressed.
[0166] Once the emotional state is identified, the server uses a generative AI model to calculate the optimal route based on it. The generative AI model takes into account real-time traffic information (e.g., traffic information APIs), weather information, the user's current location, and past behavioral history. As a result, an efficient and personalized route that matches the user's emotional state is proposed.
[0167] The calculated route information is transmitted to the terminal and presented visually as a route on a map and audibly as a guide. This presentation is adjusted according to the user's emotional state; for example, a user experiencing stress will be provided with guidance in a calm voice.
[0168] For example, if a user gives a voice command requesting the "fastest route while avoiding traffic," the terminal sends this command to a server, which uses an emotion analysis engine to recognize the user's sense of urgency. Then, using a generative AI model, it calculates the fastest route while avoiding traffic and hands it over to the terminal. The terminal then presents this to the user in a calm voice, alleviating the user's anxiety.
[0169] An example of a prompt is, "Design a car navigation system that uses an emotion engine to suggest the best traffic route when the user is in a hurry." This prompt is used to train a generative AI model and design its responses.
[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0171] Step 1:
[0172] The terminal accepts voice or text input from the user. Specifically, if the user specifies a destination as a "major landmark," the terminal acquires that voice or text data. The input in this step is the user's instruction data, and the output is digital data obtained by the server by converting this data into a parseable format.
[0173] Step 2:
[0174] The terminal sends the acquired audio data to the server. The server uses speech recognition software to convert the audio data into text data and analyze the user's voice characteristics to identify their emotional state. The input for this step is audio data, and the output is text data containing metadata indicating the user's emotional state.
[0175] Step 3:
[0176] The server uses a generative AI model based on the user's current emotional state and acquired text data to calculate the optimal route to the destination. This calculation incorporates real-time traffic and weather information. Inputs include the user's emotional state, location, and traffic information, while the output is data on the optimal route based on emotions.
[0177] Step 4:
[0178] The server sends the calculated optimal route data to the terminal. The terminal receives this data, visually displays the route on a map, and prepares audio guidance. The input for this step is the optimal route data, and the output is the visual and audio guidance information to be presented to the user.
[0179] Step 5:
[0180] The device adjusts and presents route guidance content according to the user's emotional state. Specifically, it uses a calmer voice tone and provides navigation instructions to users who are feeling stressed. The input for this step is the user's emotional state and route guidance information, and the output is route guidance optimized for the user's psychological state.
[0181] Step 6:
[0182] While receiving route guidance, users can provide additional instructions to the terminal if necessary. The server receives the new instructions, performs sentiment analysis and route calculation again, and proposes an updated route. The input for this step is the additional instructions, and the output is the updated optimal route information.
[0183] (Application Example 2)
[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0185] In autonomous vehicles, there is a need for personalized navigation and adjustments to the in-vehicle environment that take into account the emotional state of passengers. However, conventional systems have difficulty providing appropriate route guidance and in-vehicle settings that reflect the psychological state of passengers, and may lack passenger comfort. The present invention aims to solve these problems and provide a flexible travel experience that responds to the emotions of passengers.
[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0187] In this invention, the server includes means for receiving instructions from the user via voice or text, means for analyzing and understanding the user's emotions, and means for acquiring travel-related and weather information in real time. This enables highly personalized settings based on the user's emotions, providing a flexible and comfortable travel experience.
[0188] "Means of receiving user instructions via voice or text" refers to an interface that allows users to send instructions to the system via voice or text.
[0189] "Means for analyzing and understanding user emotions" refers to technologies that identify and understand emotional states from user voice and other input data.
[0190] "Means for recording and obtaining a user's current location information and past behavioral history" refers to technologies that store a user's current location and past behavioral records, and allow access to them as needed.
[0191] "Means for obtaining travel-related and weather information in real time" refers to technologies that instantly collect the latest data, such as road conditions and weather, and provide it to the system.
[0192] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a technique that uses machine learning or other algorithms to calculate the optimal travel route from collected data.
[0193] "Means of presenting routes in a manner appropriate to the user's emotions based on analyzed emotions" refers to a function that provides route information to the user visually or audibly in an appropriate manner according to the user's emotional state.
[0194] "Means for re-evaluating and updating routes based on additional instructions or emotional states from the user" refers to a function that modifies the initial route as needed and re-presents a more appropriate route.
[0195] This invention relates to a navigation system for autonomous vehicles incorporating an emotion engine. The system accepts user voice or text input, uses voice analysis technology to understand emotions, and acquires real-time location, traffic, and weather information. Based on this data, a server uses a generative AI model to calculate the optimal route to the destination. In doing so, the system takes the user's emotional state into consideration and personalizes the optimal route and in-vehicle environment.
[0196] The system analyzes the user's voice instructions using natural language processing software and a speech recognition engine. The hardware includes a GPS module for location acquisition and a high-performance processor for real-time data processing. The generated optimal route is presented in a voice and visual format that takes the user's emotions into consideration.
[0197] For example, if a user wants to relax when they get in, the autonomous vehicle will guide them along a scenic route and automatically adjust the in-car music and lighting to a relaxing setting. On the other hand, if they want to get home quickly, the vehicle will select the fastest route and provide voice-guided traffic information.
[0198] An example of a prompt message might be, "Please tell me about an AI model that uses an emotion engine to select routes and adjust the in-car environment according to the passenger's emotions."
[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0200] Step 1:
[0201] The user inputs instructions into the device via voice or text. The device converts the input voice data into text data using speech recognition software and passes it to an emotion analysis engine to extract the user's requests and emotions from their utterances.
[0202] Step 2:
[0203] The device uses an emotion analysis engine to analyze the user's emotions from voice data and retrieves the results. The analyzed emotion data is output as states such as relaxed, hurried, and stressed. These analysis results influence the subsequent optimal path calculation.
[0204] Step 3:
[0205] The server receives the user's current location, emotional state, and past behavioral history sent from the terminal. The server uses a GPS module and a database to verify this information and obtain real-time traffic and weather information. Based on this information, the generative AI model calculates the optimal route to the destination.
[0206] Step 4:
[0207] The server inputs emotional states and real-time information into a generating AI model to create the optimal route. The model selects routes with better scenery than usual, the fastest routes, etc., and outputs them as route data. This creates navigation optimized for the user's requests and emotions.
[0208] Step 5:
[0209] The terminal receives the optimal route from the server and presents the information to the user visually and audibly. The terminal guides the user with a tone and presentation that matches the user's emotional state. For example, it uses a calm voice when the user is relaxed and conveys information efficiently when the user is in a hurry.
[0210] Step 6:
[0211] If the user gives additional instructions or if there is a change in their emotional state, the device resends information to the server. The server re-evaluates the route based on the new information and provides the updated route information to the device. This maintains flexible navigation that adapts to the situation.
[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0215] [Second Embodiment]
[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0228] As an embodiment of the present invention, a scenario is shown in which a user operates the car navigation system using a portable information terminal.
[0229] The user gets into the car and activates a portable information terminal. The terminal allows the user to specify the destination and driving conditions via voice or text input. For example, the user might voice-input "I want to avoid traffic jams." This input is analyzed by the terminal and sent to the server as text data.
[0230] The device obtains the user's current location information using GPS and references their past activity history. This information, along with real-time traffic and weather information, is sent to the server. The server integrates and analyzes this data and calculates the optimal route using a generative model.
[0231] The server calculates the optimal route, which is then sent to the terminal, where it is presented to the user visually and audibly. The information presented includes estimated travel time to the destination, distance, and potential stops along the way. Once the user approves the proposed route, the terminal begins navigation. During navigation, the terminal monitors traffic conditions in real time, re-evaluating and updating the route as needed, and providing the user with updated information.
[0232] As a concrete example, consider a situation where a user wants to go to a specific cafe, but the route shown on their car's navigation system is congested. The system receives a voice command saying, "I would like an alternative route that avoids the current congestion," and the server recalculates a new route that avoids the traffic. Based on this result, this route is sent back to the terminal and presented to the user. The user can then select this route and arrive at their destination cafe smoothly, avoiding the congestion.
[0233] Thus, the system of the present invention flexibly reflects user instructions and provides individually optimized routes, thereby supporting users in comfortably reaching their destinations.
[0234] The following describes the processing flow.
[0235] Step 1:
[0236] The user activates their mobile information terminal and inputs their destination and desired conditions via voice or text. For example, they might say, "I want to go to a nearby cafe to avoid traffic."
[0237] Step 2:
[0238] The device converts voice input into text data and uses natural language processing to analyze the user's intent. Specifically, it extracts destination categories and conditions (such as avoiding traffic jams).
[0239] Step 3:
[0240] The device uses GPS to obtain the user's current location and sends the location information, the user's past activity history, and the analyzed instructions to the server.
[0241] Step 4:
[0242] Based on the data received by the server, potential destinations near the current location are collected from a database and external APIs. This includes information such as nearby cafes.
[0243] Step 5:
[0244] The server acquires traffic and weather information in real time and uses a generative model to calculate the optimal route that reflects the user's desired conditions.
[0245] Step 6:
[0246] The server sends optimal route information to the terminal and creates multiple route options, including details such as travel time and distance.
[0247] Step 7:
[0248] The device presents route options to the user and prompts them to start navigation. Navigation begins based on the route selected by the user.
[0249] Step 8:
[0250] If the user gives a voice command such as "Find a shorter route" during navigation, the device will send a new command to the server.
[0251] Step 9:
[0252] The server receives additional instructions and recalculates the route, incorporating current real-time information. It then sends the new optimal route to the terminal.
[0253] Step 10:
[0254] The device presents the user with a recalculated route and updates the navigation based on the new route if selected.
[0255] (Example 1)
[0256] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0257] In today's transportation environment, users often face difficulties in reaching their destinations smoothly. Conventional car navigation systems struggle to suggest optimal routes that take into account real-time traffic information and individual user preferences. Furthermore, they cannot reflect the user's past travel history or preferences when setting routes, which can result in inefficient route selection. There is a need for navigation systems that can solve these problems and enable users to reach their destinations comfortably.
[0258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0259] In this invention, the server includes voice analysis means and text analysis means for receiving instructions from the user, location estimation means for acquiring current location information and past behavioral history, and information acquisition means for acquiring traffic information and weather information in real time. This makes it possible to calculate and present the optimal route that takes into account the user's individual preferences and current traffic conditions.
[0260] "Voice analysis means" refers to a function or device for converting a user's voice instructions into digital data and analyzing the intended meaning.
[0261] "Text analysis means" refers to a function or device for analyzing text data entered by a user and understanding the content of the instructions.
[0262] "Location estimation means" refers to a function or device that uses GPS technology or similar to determine the user's current location and acquire their past activity history.
[0263] "Information acquisition means" refers to a function or device for collecting traffic information and weather information in real time from external information sources.
[0264] A "generative model based on deep learning" is a model that uses multi-layer neural network technology to analyze data and generate the optimal path.
[0265] A "route calculation means" is a function or device for calculating the optimal route to a destination based on acquired information.
[0266] "Presentation means" refers to a function or device for presenting calculated route information to the user visually or audibly.
[0267] A "re-evaluation means" is a function or device that receives additional instructions from the user, re-evaluates existing routes, and updates the routes as necessary.
[0268] The user utilizes a portable information terminal to implement the navigation system of the present invention. The user, upon entering the vehicle, activates the portable terminal and inputs the destination and driving conditions via voice or text. The terminal analyzes this user input using voice analysis software (e.g., a speech recognition API) and processes it as text. The analyzed text data is then linked with a location estimation means and transmitted to a server along with GPS information and past activity history.
[0269] The server integrates the received information and calculates the optimal route using a generative AI model based on deep learning. Real-time traffic and weather information is obtained from external information services (for example, map API services) and is also included in the analysis. Once the optimal route is calculated, the server sends this route information back to the terminal.
[0270] The route information sent is displayed on the terminal and guided to the user using speech synthesis software (e.g., a speech synthesis API). Navigation begins once the user confirms and approves the route. While the user is moving, the terminal and server communicate continuously, and if new traffic information is acquired, the route is re-evaluated based on that information, and updated information is provided to the user.
[0271] As a concrete example, a user might instruct the system by voice, "I want to avoid traffic." This input is analyzed by the server, and a new, optimized route is calculated by a generative AI model. This route is then presented to the user's device.
[0272] An example of a prompt for a generative AI model is, "If a user requests to avoid traffic congestion, please tell me how to calculate the optimal route considering the current traffic conditions." This enables flexible guidance tailored to the user's needs.
[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0274] Step 1:
[0275] The user activates their mobile device and inputs their destination and driving conditions via voice or text. The user's input is converted into text data by a voice analysis system. The converted text data is temporarily stored on the device. Here, input is voice or text, and output is text data for analysis. Voice analysis software (e.g., speech recognition API) is used to convert voice to text.
[0276] Step 2:
[0277] The terminal uses GPS technology to obtain the user's current location. Furthermore, it refers to the user's past behavior history and also obtains this. The current location data and the behavior history data serve as inputs for transmission to the server. The output is the integrated location data used by the server. Specifically, the location estimation means collects GPS data in real time and organizes it.
[0278] Step 3:
[0279] The server integrates the text data, current location, past behavior history transmitted from the terminal, and real-time traffic information and weather information obtained via an external API. The integrated data is input into the generative AI model. The output here is a dataset for predictive analysis by the generative AI model. The server fetches the necessary information from various external information providing services (e.g., map API services), structures the data, and prepares it for input into the model.
[0280] Step 4:
[0281] The generative AI model calculates the optimal route. In this process, a simulation is performed based on the integrated dataset to calculate the optimal route. The input here is the integrated dataset, and the output is the information on the optimal route. In this step, the calculation is performed using deep learning technology on the server.
[0282] Step 5:
[0283] The server transmits the calculated optimal route to the terminal. The terminal presents the received route information to the user using a map display application or voice synthesis software. The input is the optimal route information, and the output is the navigation information presented to the user visually and audibly. The terminal provides information to the user both visually through map display and audibly through voice guidance.
[0284] Step 6:
[0285] When the user approves the route, the terminal starts navigation. During navigation, the terminal continues to communicate with the server, re-evaluates the route according to changes in traffic conditions and weather information, and updates it as necessary. The input is real-time traffic information, and the output is updated route information. The terminal receives the information updated in real time and appropriately modifies and provides the information to the user.
[0286] (Application Example 1)
[0287] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0288] There is a need for a means by which the driver of an autonomous vehicle can intuitively obtain optimal route information for reaching the destination. Also, it is an issue to realize a system in which the driver can receive visual and voice guidance while coping with traffic conditions and weather conditions that change in real time.
[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0290] In this invention, the server includes means for receiving an instruction from the user by voice or text, means for recording and acquiring the user's current location information and past behavior history, and means for acquiring traffic information and weather information in real time. Thereby, the driver can visually confirm the route information using an additional device and can always reach the destination by the optimal route.
[0291] The "means for receiving an instruction from the user by voice or text" is an interface for receiving the user's wishes for the destination and route by voice or text input and processing it in the system.
[0292] "Means for recording and acquiring a user's current location information and past behavioral history" refers to technologies that use GPS or similar methods to acquire a user's real-time location and record and refer to historical data such as places visited in the past.
[0293] "Means for obtaining real-time traffic and weather information" refers to technologies that continuously acquire the latest traffic conditions and weather data via a network.
[0294] "Means for calculating the optimal route to the destination based on the acquired information using a generative model" refers to a method of calculating the optimal travel route based on collected data using machine learning or AI technology.
[0295] "Means for presenting the calculated optimal route to the user" refers to technical means for providing route information to the user through displays, audio, etc.
[0296] "Means for re-evaluating and updating routes based on additional instructions from the user" refers to a method for recalculating and updating routes in real time in response to new route information and changes in conditions provided by the user.
[0297] "Means of providing route information as visual information through the driver's auxiliary device to assist route guidance" refers to methods of intuitively presenting route information to the driver visually using smart glasses or head-mounted displays.
[0298] As an embodiment of this invention, a navigation system for an autonomous vehicle utilizing a smart device is specifically described. The user can wear smart glasses equipped in the autonomous vehicle and obtain comfortable route guidance while visually checking traffic information in real time.
[0299] The server first receives voice or text input from the user and interprets the destination information and driving conditions. Google Speech-to-Text can be used as the speech recognition API for this purpose. Next, the server obtains the user's current location using a GPS module and accesses past activity history stored on a cloud server. Furthermore, a network connection is required to collect traffic and weather information from the internet in real time.
[0300] Using a generative AI model (e.g., Google AI), the server calculates the optimal route from this information. The calculated route is then presented to the user as visual information through smart glasses. For example, if a user requests a route that is not congested, the server will present a new route that reflects real-time congestion. An example of a prompt in this case would be: "The user wants to go to a specific destination via a route that is not congested. Please provide the optimal route considering the current traffic conditions."
[0301] This system allows drivers to arrive at their destination via the optimal route at all times, while receiving visual and audible feedback. By flexibly responding to real-time changing traffic and weather conditions, it supports a comfortable travel experience for users.
[0302] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0303] Step 1:
[0304] The terminal starts up, and the user inputs the destination and driving conditions by voice or text. The input is received via the terminal's microphone or touchscreen. The terminal converts this into text data using a speech recognition API.
[0305] Step 2:
[0306] The terminal uses a GPS module to obtain the user's current location information. In addition to this, it obtains the past behavior history from the cloud server and utilizes it as reference information for movement patterns and destinations. The input is the current location information and past data, and the output is the user location and related history.
[0307] Step 3:
[0308] The server collects real-time traffic information and weather information via the Internet. Thereby, the current traffic situation and weather can be grasped. The input is the information from the external API, and the output is the latest traffic and weather data.
[0309] Step 4:
[0310] The server uses a generative AI model to calculate the optimal route based on the user information and external data obtained so far. Utilize the prompt text to perform personalization according to the user's request. Specifically, a route is generated according to an instruction such as "the user wants to go to a specific destination by a 'congestion-free' route". The output is the optimized route information.
[0311] Step 5:
[0312] The calculated optimal route is sent to the terminal and presented to the user visually and audibly. The user can confirm the direction of travel through the map and route information displayed on the smart glasses. The input is the route information from the server, and the output is visual and audible guidance.
[0313] Step 6:
[0314] When the user gives an additional instruction, the server re-evaluates the route again and updates it as necessary. Thereby, route guidance corresponding to dynamic situation changes becomes possible. The re-evaluated route is sent back to the terminal again. The input is the user's new instruction, and the output is the updated route information.
[0315] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0316] An embodiment of the present invention is shown below: a car navigation system incorporating an emotion engine. This system operates when a user receives route guidance using a portable information terminal.
[0317] The user activates a portable information terminal and provides instructions to the system regarding destination and route conditions via voice or text. The system incorporates an emotion engine that analyzes the user's emotions through voice input, allowing it to understand the user's emotional state based on the input voice data.
[0318] The emotion engine adjusts the generative model to prioritize scenic routes when the user is relaxed, for example, and to calculate the fastest route when the user is in a hurry or stressed. The server provides comprehensive route suggestions, taking into account the user's current location, past behavior history, and real-time traffic and weather information.
[0319] The calculated optimal route is transmitted to the terminal, which then presents this information to the user visually and audibly. The presentation method can also be adjusted according to the user's emotional state. For example, if the user is stressed, the information may be conveyed in a calm voice.
[0320] For example, if a user is in a hurry because they are running late for their scheduled departure time and gives the voice command, "I absolutely must avoid traffic," the system will recognize the urgency through its emotion engine, and the server will calculate a route that prioritizes speed. This calculated route information will be quickly presented to the user through the terminal, using a calm voice tone to reduce stress.
[0321] Thus, the system of the present invention, which incorporates an emotion engine, further personalizes the car navigation experience according to the user's psychological state, supporting more flexible and comfortable travel.
[0322] The following describes the processing flow.
[0323] Step 1:
[0324] The user activates their portable information terminal and inputs a voice command saying, "I want to go to my destination while avoiding the highway." At this time, the terminal records the voice data.
[0325] Step 2:
[0326] The device converts voice input into text data and simultaneously uses an emotion engine to analyze the user's emotional state from the tone and speed of their voice. The analysis result might be determined to indicate, for example, "the user is irritated."
[0327] Step 3:
[0328] The device obtains its current location information via GPS and collects data on the user's past behavior. This information, along with the results of sentiment analysis, is sent to the server.
[0329] Step 4:
[0330] Based on the server's received location information, the user's desired route conditions, and emotional state, real-time traffic and weather information is obtained.
[0331] Step 5:
[0332] The server uses a generative model to calculate the optimal route to the destination. This calculation takes into account the user's emotional state, prioritizing routes that, for example, allow for quick and stress-reducing travel.
[0333] Step 6:
[0334] The server calculates the optimal route and sends it to the terminal, then prepares detailed route information.
[0335] Step 7:
[0336] The device presents the route information it has received to the user. Depending on the user's emotional state, it provides guidance in a calm and reassuring voice, for example.
[0337] Step 8:
[0338] The user confirms the suggested route and begins navigation. If the user gives a new instruction during navigation, such as "I want to arrive sooner," the device sends this information back to the server.
[0339] Step 9:
[0340] The server receives the new instructions and, as in the previous step, re-evaluates the user's emotions and requests, and recalculates a more optimized path.
[0341] Step 10:
[0342] The device will suggest a new route, and if selected, it will immediately update the navigation to help the user have the most satisfying journey.
[0343] (Example 2)
[0344] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0345] Conventional car navigation systems typically suggest the optimal route without considering the user's emotional state. This lack of flexible guidance, adapted to the user's psychological condition, can increase stress. Furthermore, they lack methods to adjust guidance based on emotional state, highlighting the need for improved user experience.
[0346] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0347] In this invention, the server includes means for receiving instructions from the user in voice or text, means for analyzing the user's emotional state, and means for recording and acquiring the user's current location information and past behavioral history. This enables the personalization of the optimal route based on the user's emotional state, and route guidance that is appropriate to their psychological state.
[0348] "Means for receiving user instructions via voice or text" refers to a function that provides an interface for users to specify destination and route conditions to the system via voice or text.
[0349] "Means for analyzing a user's emotional state" refers to technologies or devices used to identify a user's emotional state from voice data or user interactions. This makes it possible to understand the user's psychological state.
[0350] "Means for recording and acquiring a user's current location information and past behavioral history" refers to a device or method for acquiring a user's location information in real time, storing it, and analyzing and utilizing past movement patterns.
[0351] "Means for obtaining real-time traffic and weather information" refers to technologies or services that collect the latest traffic and weather conditions from external sources and provide them to a system.
[0352] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a method or apparatus that uses AI technology or the like to calculate the most suitable travel route for the user based on the acquired information.
[0353] "A means of presenting a calculated optimal route to the user and adjusting the presented content to the user's emotional state" refers to a technology that presents calculated route information to the user using visual and auditory methods, and applies communication methods that correspond to the user's emotions at that time.
[0354] "Means for re-evaluating and updating routes based on additional instructions from users" refers to a mechanism that reviews existing routes and updates them to the optimal ones in response to new requests or changes in conditions from users.
[0355] The system based on this invention realizes an advanced car navigation system equipped with emotion analysis capabilities. The system uses the user's mobile device as its primary interface and accepts instructions from the user via voice or text.
[0356] The terminal uses hardware such as a voice input device and a touchscreen to acquire information about the user's destination and route conditions. The acquired voice data is sent to a server equipped with an emotion analysis engine. This server uses speech recognition software (e.g., a common speech recognition API) to analyze the user's emotional state from their voice. In this process, by analyzing features such as tone of voice, word choice, and speaking speed, the server identifies the user's emotional state, such as whether they are relaxed, in a hurry, or stressed.
[0357] Once the emotional state is identified, the server uses a generative AI model to calculate the optimal route based on it. The generative AI model takes into account real-time traffic information (e.g., traffic information APIs), weather information, the user's current location, and past behavioral history. As a result, an efficient and personalized route that matches the user's emotional state is proposed.
[0358] The calculated route information is transmitted to the terminal and presented visually as a route on a map and audibly as a guide. This presentation is adjusted according to the user's emotional state; for example, a user experiencing stress will be provided with guidance in a calm voice.
[0359] For example, if a user gives a voice command requesting the "fastest route while avoiding traffic," the terminal sends this command to a server, which uses an emotion analysis engine to recognize the user's sense of urgency. Then, using a generative AI model, it calculates the fastest route while avoiding traffic and hands it over to the terminal. The terminal then presents this to the user in a calm voice, alleviating the user's anxiety.
[0360] An example of a prompt is, "Design a car navigation system that uses an emotion engine to suggest the best traffic route when the user is in a hurry." This prompt is used to train a generative AI model and design its responses.
[0361] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0362] Step 1:
[0363] The terminal accepts voice or text input from the user. Specifically, if the user specifies a destination as a "major landmark," the terminal acquires that voice or text data. The input in this step is the user's instruction data, and the output is digital data obtained by the server by converting this data into a parseable format.
[0364] Step 2:
[0365] The terminal sends the acquired audio data to the server. The server uses speech recognition software to convert the audio data into text data and analyze the user's voice characteristics to identify their emotional state. The input for this step is audio data, and the output is text data containing metadata indicating the user's emotional state.
[0366] Step 3:
[0367] The server uses a generative AI model based on the user's current emotional state and acquired text data to calculate the optimal route to the destination. This calculation incorporates real-time traffic and weather information. Inputs include the user's emotional state, location, and traffic information, while the output is data on the optimal route based on emotions.
[0368] Step 4:
[0369] The server sends the calculated optimal route data to the terminal. The terminal receives this data, visually displays the route on a map, and prepares audio guidance. The input for this step is the optimal route data, and the output is the visual and audio guidance information to be presented to the user.
[0370] Step 5:
[0371] The device adjusts and presents route guidance content according to the user's emotional state. Specifically, it uses a calmer voice tone and provides navigation instructions to users who are feeling stressed. The input for this step is the user's emotional state and route guidance information, and the output is route guidance optimized for the user's psychological state.
[0372] Step 6:
[0373] While receiving route guidance, users can provide additional instructions to the terminal if necessary. The server receives the new instructions, performs sentiment analysis and route calculation again, and proposes an updated route. The input for this step is the additional instructions, and the output is the updated optimal route information.
[0374] (Application Example 2)
[0375] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0376] In autonomous vehicles, there is a need for personalized navigation and adjustments to the in-vehicle environment that take into account the emotional state of passengers. However, conventional systems have difficulty providing appropriate route guidance and in-vehicle settings that reflect the psychological state of passengers, and may lack passenger comfort. The present invention aims to solve these problems and provide a flexible travel experience that responds to the emotions of passengers.
[0377] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0378] In this invention, the server includes means for receiving instructions from the user via voice or text, means for analyzing and understanding the user's emotions, and means for acquiring travel-related and weather information in real time. This enables highly personalized settings based on the user's emotions, providing a flexible and comfortable travel experience.
[0379] "Means of receiving user instructions via voice or text" refers to an interface that allows users to send instructions to the system via voice or text.
[0380] "Means for analyzing and understanding user emotions" refers to technologies that identify and understand emotional states from user voice and other input data.
[0381] "Means for recording and obtaining a user's current location information and past behavioral history" refers to technologies that store a user's current location and past behavioral records, and allow access to them as needed.
[0382] "Means for obtaining travel-related and weather information in real time" refers to technologies that instantly collect the latest data, such as road conditions and weather, and provide it to the system.
[0383] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a technique that uses machine learning or other algorithms to calculate the optimal travel route from collected data.
[0384] "Means of presenting routes in a manner appropriate to the user's emotions based on analyzed emotions" refers to a function that provides route information to the user visually or audibly in an appropriate manner according to the user's emotional state.
[0385] "Means for re-evaluating and updating routes based on additional instructions or emotional states from the user" refers to a function that modifies the initial route as needed and re-presents a more appropriate route.
[0386] This invention relates to a navigation system for autonomous vehicles incorporating an emotion engine. The system accepts user voice or text input, uses voice analysis technology to understand emotions, and acquires real-time location, traffic, and weather information. Based on this data, a server uses a generative AI model to calculate the optimal route to the destination. In doing so, the system takes the user's emotional state into consideration and personalizes the optimal route and in-vehicle environment.
[0387] The system analyzes the user's voice instructions using natural language processing software and a speech recognition engine. The hardware includes a GPS module for location acquisition and a high-performance processor for real-time data processing. The generated optimal route is presented in a voice and visual format that takes the user's emotions into consideration.
[0388] For example, if a user wants to relax when they get in, the autonomous vehicle will guide them along a scenic route and automatically adjust the in-car music and lighting to a relaxing setting. On the other hand, if they want to get home quickly, the vehicle will select the fastest route and provide voice-guided traffic information.
[0389] An example of a prompt message might be, "Please tell me about an AI model that uses an emotion engine to select routes and adjust the in-car environment according to the passenger's emotions."
[0390] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0391] Step 1:
[0392] The user inputs instructions into the device via voice or text. The device converts the input voice data into text data using speech recognition software and passes it to an emotion analysis engine to extract the user's requests and emotions from their utterances.
[0393] Step 2:
[0394] The device uses an emotion analysis engine to analyze the user's emotions from voice data and retrieves the results. The analyzed emotion data is output as states such as relaxed, hurried, and stressed. These analysis results influence the subsequent optimal path calculation.
[0395] Step 3:
[0396] The server receives the user's current location, emotional state, and past behavioral history sent from the terminal. The server uses a GPS module and a database to verify this information and obtain real-time traffic and weather information. Based on this information, the generative AI model calculates the optimal route to the destination.
[0397] Step 4:
[0398] The server inputs emotional states and real-time information into a generating AI model to create the optimal route. The model selects routes with better scenery than usual, the fastest routes, etc., and outputs them as route data. This creates navigation optimized for the user's requests and emotions.
[0399] Step 5:
[0400] The terminal receives the optimal route from the server and presents the information to the user visually and audibly. The terminal guides the user with a tone and presentation that matches the user's emotional state. For example, it uses a calm voice when the user is relaxed and conveys information efficiently when the user is in a hurry.
[0401] Step 6:
[0402] If the user gives additional instructions or if there is a change in their emotional state, the device resends information to the server. The server re-evaluates the route based on the new information and provides the updated route information to the device. This maintains flexible navigation that adapts to the situation.
[0403] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0404] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0405] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0406] [Third Embodiment]
[0407] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0408] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0409] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0410] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0411] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0412] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0413] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0414] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0415] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0416] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0417] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0418] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0419] As an embodiment of the present invention, a scenario is shown in which a user operates the car navigation system using a portable information terminal.
[0420] The user gets into the car and activates a portable information terminal. The terminal allows the user to specify the destination and driving conditions via voice or text input. For example, the user might voice-input "I want to avoid traffic jams." This input is analyzed by the terminal and sent to the server as text data.
[0421] The device obtains the user's current location information using GPS and references their past activity history. This information, along with real-time traffic and weather information, is sent to the server. The server integrates and analyzes this data and calculates the optimal route using a generative model.
[0422] The server calculates the optimal route, which is then sent to the terminal, where it is presented to the user visually and audibly. The information presented includes estimated travel time to the destination, distance, and potential stops along the way. Once the user approves the proposed route, the terminal begins navigation. During navigation, the terminal monitors traffic conditions in real time, re-evaluating and updating the route as needed, and providing the user with updated information.
[0423] As a concrete example, consider a situation where a user wants to go to a specific cafe, but the route shown on their car's navigation system is congested. The system receives a voice command saying, "I would like an alternative route that avoids the current congestion," and the server recalculates a new route that avoids the traffic. Based on this result, this route is sent back to the terminal and presented to the user. The user can then select this route and arrive at their destination cafe smoothly, avoiding the congestion.
[0424] Thus, the system of the present invention flexibly reflects user instructions and provides individually optimized routes, thereby supporting users in comfortably reaching their destinations.
[0425] The following describes the processing flow.
[0426] Step 1:
[0427] The user activates their mobile information terminal and inputs their destination and desired conditions via voice or text. For example, they might say, "I want to go to a nearby cafe to avoid traffic."
[0428] Step 2:
[0429] The device converts voice input into text data and uses natural language processing to analyze the user's intent. Specifically, it extracts destination categories and conditions (such as avoiding traffic jams).
[0430] Step 3:
[0431] The device uses GPS to obtain the user's current location and sends the location information, the user's past activity history, and the analyzed instructions to the server.
[0432] Step 4:
[0433] Based on the data received by the server, potential destinations near the current location are collected from a database and external APIs. This includes information such as nearby cafes.
[0434] Step 5:
[0435] The server acquires traffic and weather information in real time and uses a generative model to calculate the optimal route that reflects the user's desired conditions.
[0436] Step 6:
[0437] The server sends optimal route information to the terminal and creates multiple route options, including details such as travel time and distance.
[0438] Step 7:
[0439] The device presents route options to the user and prompts them to start navigation. Navigation begins based on the route selected by the user.
[0440] Step 8:
[0441] If the user gives a voice command such as "Find a shorter route" during navigation, the device will send a new command to the server.
[0442] Step 9:
[0443] The server receives additional instructions and recalculates the route, incorporating current real-time information. It then sends the new optimal route to the terminal.
[0444] Step 10:
[0445] The device presents the user with a recalculated route and updates the navigation based on the new route if selected.
[0446] (Example 1)
[0447] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0448] In today's transportation environment, users often face difficulties in reaching their destinations smoothly. Conventional car navigation systems struggle to suggest optimal routes that take into account real-time traffic information and individual user preferences. Furthermore, they cannot reflect the user's past travel history or preferences when setting routes, which can result in inefficient route selection. There is a need for navigation systems that can solve these problems and enable users to reach their destinations comfortably.
[0449] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0450] In this invention, the server includes voice analysis means and text analysis means for receiving instructions from the user, location estimation means for acquiring current location information and past behavioral history, and information acquisition means for acquiring traffic information and weather information in real time. This makes it possible to calculate and present the optimal route that takes into account the user's individual preferences and current traffic conditions.
[0451] "Voice analysis means" refers to a function or device for converting a user's voice instructions into digital data and analyzing the intended meaning.
[0452] "Text analysis means" refers to a function or device for analyzing text data entered by a user and understanding the content of the instructions.
[0453] "Location estimation means" refers to a function or device that uses GPS technology or similar to determine the user's current location and acquire their past activity history.
[0454] "Information acquisition means" refers to a function or device for collecting traffic information and weather information in real time from external information sources.
[0455] A "generative model based on deep learning" is a model that uses multi-layer neural network technology to analyze data and generate the optimal path.
[0456] A "route calculation means" is a function or device for calculating the optimal route to a destination based on acquired information.
[0457] "Presentation means" refers to a function or device for presenting calculated route information to the user visually or audibly.
[0458] A "re-evaluation means" is a function or device that receives additional instructions from the user, re-evaluates existing routes, and updates the routes as necessary.
[0459] The user utilizes a portable information terminal to implement the navigation system of the present invention. The user, upon entering the vehicle, activates the portable terminal and inputs the destination and driving conditions via voice or text. The terminal analyzes this user input using voice analysis software (e.g., a speech recognition API) and processes it as text. The analyzed text data is then linked with a location estimation means and transmitted to a server along with GPS information and past activity history.
[0460] The server integrates the received information and calculates the optimal route using a generative AI model based on deep learning. Real-time traffic and weather information is obtained from external information services (for example, map API services) and is also included in the analysis. Once the optimal route is calculated, the server sends this route information back to the terminal.
[0461] The route information sent is displayed on the terminal and guided to the user using speech synthesis software (e.g., a speech synthesis API). Navigation begins once the user confirms and approves the route. While the user is moving, the terminal and server communicate continuously, and if new traffic information is acquired, the route is re-evaluated based on that information, and updated information is provided to the user.
[0462] As a concrete example, a user might instruct the system by voice, "I want to avoid traffic." This input is analyzed by the server, and a new, optimized route is calculated by a generative AI model. This route is then presented to the user's device.
[0463] An example of a prompt for a generative AI model is, "If a user requests to avoid traffic congestion, please tell me how to calculate the optimal route considering the current traffic conditions." This enables flexible guidance tailored to the user's needs.
[0464] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0465] Step 1:
[0466] The user activates their mobile device and inputs their destination and driving conditions via voice or text. The user's input is converted into text data by a voice analysis system. The converted text data is temporarily stored on the device. Here, input is voice or text, and output is text data for analysis. Voice analysis software (e.g., speech recognition API) is used to convert voice to text.
[0467] Step 2:
[0468] The device uses GPS technology to obtain the user's current location. Furthermore, it accesses and obtains the user's past activity history. The current location data and activity history data serve as input for transmission to the server. The output is integrated location data used by the server. Specifically, a location estimation system collects GPS data in real time and organizes it.
[0469] Step 3:
[0470] The server integrates text data sent from the terminal, current location, past activity history, and real-time traffic and weather information obtained via external APIs. The integrated data is input into a generative AI model. The output here is a dataset for predictive analysis by the generative AI model. The server retrieves necessary information from various external information provision services (e.g., map API services), structures the data, and prepares it for input into the model.
[0471] Step 4:
[0472] The generative AI model calculates the optimal path. This process involves simulating the optimal path based on an integrated dataset. The input is the integrated dataset, and the output is information about the optimal path. This step is performed on a server using deep learning techniques.
[0473] Step 5:
[0474] The server sends the calculated optimal route to the terminal. The terminal presents the received route information to the user using a map display application or speech synthesis software. The input is the optimal route information, and the output is navigation information presented to the user visually and audibly. The terminal provides information to the user through both visual map display and voice guidance.
[0475] Step 6:
[0476] Once the user approves the route, the device begins navigation. During navigation, the device continues to communicate with the server, re-evaluating the route based on changes in traffic and weather conditions, and updating it as needed. Input is real-time traffic information, and output is updated route information. The device receives real-time updated information and provides it to the user, correcting it as appropriate.
[0477] (Application Example 1)
[0478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0479] There is a need for a way for drivers of autonomous vehicles to intuitively obtain optimal route information to reach their destination. Furthermore, a challenge is to realize a system that provides drivers with visual and audio guidance while responding to real-time changes in traffic and weather conditions.
[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0481] In this invention, the server includes means for receiving instructions from the user via voice or text, means for recording and acquiring the user's current location information and past activity history, and means for acquiring traffic and weather information in real time. This allows the driver to visually confirm route information using an additional device and always reach their destination via the optimal route.
[0482] "Means for receiving user instructions via voice or text" refers to an interface for receiving user requests for destinations and routes via voice or text input and processing them within the system.
[0483] "Means for recording and acquiring a user's current location information and past behavioral history" refers to technologies that use GPS or similar methods to acquire a user's real-time location and record and refer to historical data such as places visited in the past.
[0484] "Means for obtaining real-time traffic and weather information" refers to technologies that continuously acquire the latest traffic conditions and weather data via a network.
[0485] "Means for calculating the optimal route to the destination based on the acquired information using a generative model" refers to a method of calculating the optimal travel route based on collected data using machine learning or AI technology.
[0486] "Means for presenting the calculated optimal route to the user" refers to technical means for providing route information to the user through displays, audio, etc.
[0487] "Means for re-evaluating and updating routes based on additional instructions from the user" refers to a method for recalculating and updating routes in real time in response to new route information and changes in conditions provided by the user.
[0488] "Means of providing route information as visual information through the driver's auxiliary device to assist route guidance" refers to methods of intuitively presenting route information to the driver visually using smart glasses or head-mounted displays.
[0489] As an embodiment of this invention, a navigation system for an autonomous vehicle utilizing a smart device is specifically described. The user can wear smart glasses equipped in the autonomous vehicle and obtain comfortable route guidance while visually checking traffic information in real time.
[0490] The server first receives voice or text input from the user and interprets the destination information and driving conditions. Google Speech-to-Text can be used as the speech recognition API for this purpose. Next, the server obtains the user's current location using a GPS module and accesses past activity history stored on a cloud server. Furthermore, a network connection is required to collect traffic and weather information from the internet in real time.
[0491] Using a generative AI model (e.g., Google AI), the server calculates the optimal route from this information. The calculated route is then presented to the user as visual information through smart glasses. For example, if a user requests a route that is not congested, the server will present a new route that reflects real-time congestion. An example of a prompt in this case would be: "The user wants to go to a specific destination via a route that is not congested. Please provide the optimal route considering the current traffic conditions."
[0492] This system allows drivers to arrive at their destination via the optimal route at all times, while receiving visual and audible feedback. By flexibly responding to real-time changing traffic and weather conditions, it supports a comfortable travel experience for users.
[0493] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0494] Step 1:
[0495] The terminal starts up, and the user inputs the destination and driving conditions by voice or text. The input is received via the terminal's microphone or touchscreen. The terminal converts this into text data using a speech recognition API.
[0496] Step 2:
[0497] The device uses a GPS module to obtain the user's current location information. In addition, it retrieves past activity history from a cloud server and uses it as reference information for movement patterns and destinations. The input is current location information and past data, and the output is the user's location and associated history.
[0498] Step 3:
[0499] The server collects real-time traffic and weather information via the internet. This allows for monitoring of current traffic conditions and weather. Input is information from external APIs, and output is the latest traffic and weather data.
[0500] Step 4:
[0501] The server uses a generative AI model to calculate the optimal route based on user information and external data obtained so far. It utilizes prompts to personalize the route according to user requests. Specifically, it generates a route based on instructions such as, "The user wants to go to a specific destination via a 'congested' route." The output is optimized route information.
[0502] Step 5:
[0503] The calculated optimal route is transmitted to the terminal and presented to the user visually and audibly. The user can confirm their direction of travel through maps and route information displayed on smart glasses. The input is route information from the server, and the output is visual and audible guidance.
[0504] Step 6:
[0505] If the user provides additional instructions, the server re-evaluates the route and updates it as needed. This enables route guidance that adapts to dynamic changes in circumstances. The re-evaluated route is sent back to the terminal. The input is the user's new instructions, and the output is the updated route information.
[0506] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0507] An embodiment of the present invention is shown below: a car navigation system incorporating an emotion engine. This system operates when a user receives route guidance using a portable information terminal.
[0508] The user activates a portable information terminal and provides instructions to the system regarding destination and route conditions via voice or text. The system incorporates an emotion engine that analyzes the user's emotions through voice input, allowing it to understand the user's emotional state based on the input voice data.
[0509] The emotion engine adjusts the generative model to prioritize scenic routes when the user is relaxed, for example, and to calculate the fastest route when the user is in a hurry or stressed. The server provides comprehensive route suggestions, taking into account the user's current location, past behavior history, and real-time traffic and weather information.
[0510] The calculated optimal route is transmitted to the terminal, which then presents this information to the user visually and audibly. The presentation method can also be adjusted according to the user's emotional state. For example, if the user is stressed, the information may be conveyed in a calm voice.
[0511] For example, if a user is in a hurry because they are running late for their scheduled departure time and gives the voice command, "I absolutely must avoid traffic," the system will recognize the urgency through its emotion engine, and the server will calculate a route that prioritizes speed. This calculated route information will be quickly presented to the user through the terminal, using a calm voice tone to reduce stress.
[0512] Thus, the system of the present invention, which incorporates an emotion engine, further personalizes the car navigation experience according to the user's psychological state, supporting more flexible and comfortable travel.
[0513] The following describes the processing flow.
[0514] Step 1:
[0515] The user activates their portable information terminal and inputs a voice command saying, "I want to go to my destination while avoiding the highway." At this time, the terminal records the voice data.
[0516] Step 2:
[0517] The device converts voice input into text data and simultaneously uses an emotion engine to analyze the user's emotional state from the tone and speed of their voice. The analysis result might be determined to indicate, for example, "the user is irritated."
[0518] Step 3:
[0519] The device obtains its current location information via GPS and collects data on the user's past behavior. This information, along with the results of sentiment analysis, is sent to the server.
[0520] Step 4:
[0521] Based on the server's received location information, the user's desired route conditions, and emotional state, real-time traffic and weather information is obtained.
[0522] Step 5:
[0523] The server uses a generative model to calculate the optimal route to the destination. This calculation takes into account the user's emotional state, prioritizing routes that, for example, allow for quick and stress-reducing travel.
[0524] Step 6:
[0525] The server calculates the optimal route and sends it to the terminal, then prepares detailed route information.
[0526] Step 7:
[0527] The device presents the route information it has received to the user. Depending on the user's emotional state, it provides guidance in a calm and reassuring voice, for example.
[0528] Step 8:
[0529] The user confirms the suggested route and begins navigation. If the user gives a new instruction during navigation, such as "I want to arrive sooner," the device sends this information back to the server.
[0530] Step 9:
[0531] The server receives the new instructions and, as in the previous step, re-evaluates the user's emotions and requests, and recalculates a more optimized path.
[0532] Step 10:
[0533] The device will suggest a new route, and if selected, it will immediately update the navigation to help the user have the most satisfying journey.
[0534] (Example 2)
[0535] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] Conventional car navigation systems typically suggest the optimal route without considering the user's emotional state. This lack of flexible guidance, adapted to the user's psychological condition, can increase stress. Furthermore, they lack methods to adjust guidance based on emotional state, highlighting the need for improved user experience.
[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0538] In this invention, the server includes means for receiving instructions from the user in voice or text, means for analyzing the user's emotional state, and means for recording and acquiring the user's current location information and past behavioral history. This enables the personalization of the optimal route based on the user's emotional state, and route guidance that is appropriate to their psychological state.
[0539] "Means for receiving user instructions via voice or text" refers to a function that provides an interface for users to specify destination and route conditions to the system via voice or text.
[0540] "Means for analyzing a user's emotional state" refers to technologies or devices used to identify a user's emotional state from voice data or user interactions. This makes it possible to understand the user's psychological state.
[0541] "Means for recording and acquiring a user's current location information and past behavioral history" refers to a device or method for acquiring a user's location information in real time, storing it, and analyzing and utilizing past movement patterns.
[0542] "Means for obtaining real-time traffic and weather information" refers to technologies or services that collect the latest traffic and weather conditions from external sources and provide them to a system.
[0543] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a method or apparatus that uses AI technology or the like to calculate the most suitable travel route for the user based on the acquired information.
[0544] "A means of presenting a calculated optimal route to the user and adjusting the presented content to the user's emotional state" refers to a technology that presents calculated route information to the user using visual and auditory methods, and applies communication methods that correspond to the user's emotions at that time.
[0545] "Means for re-evaluating and updating routes based on additional instructions from users" refers to a mechanism that reviews existing routes and updates them to the optimal ones in response to new requests or changes in conditions from users.
[0546] The system based on this invention realizes an advanced car navigation system equipped with emotion analysis capabilities. The system uses the user's mobile device as its primary interface and accepts instructions from the user via voice or text.
[0547] The terminal uses hardware such as a voice input device and a touchscreen to acquire information about the user's destination and route conditions. The acquired voice data is sent to a server equipped with an emotion analysis engine. This server uses speech recognition software (e.g., a common speech recognition API) to analyze the user's emotional state from their voice. In this process, by analyzing features such as tone of voice, word choice, and speaking speed, the server identifies the user's emotional state, such as whether they are relaxed, in a hurry, or stressed.
[0548] Once the emotional state is identified, the server uses a generative AI model to calculate the optimal route based on it. The generative AI model takes into account real-time traffic information (e.g., traffic information APIs), weather information, the user's current location, and past behavioral history. As a result, an efficient and personalized route that matches the user's emotional state is proposed.
[0549] The calculated route information is transmitted to the terminal and presented visually as a route on a map and audibly as a guide. This presentation is adjusted according to the user's emotional state; for example, a user experiencing stress will be provided with guidance in a calm voice.
[0550] For example, if a user gives a voice command requesting the "fastest route while avoiding traffic," the terminal sends this command to a server, which uses an emotion analysis engine to recognize the user's sense of urgency. Then, using a generative AI model, it calculates the fastest route while avoiding traffic and hands it over to the terminal. The terminal then presents this to the user in a calm voice, alleviating the user's anxiety.
[0551] An example of a prompt is, "Design a car navigation system that uses an emotion engine to suggest the best traffic route when the user is in a hurry." This prompt is used to train a generative AI model and design its responses.
[0552] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0553] Step 1:
[0554] The terminal accepts voice or text input from the user. Specifically, if the user specifies a destination as a "major landmark," the terminal acquires that voice or text data. The input in this step is the user's instruction data, and the output is digital data obtained by the server by converting this data into a parseable format.
[0555] Step 2:
[0556] The terminal sends the acquired audio data to the server. The server uses speech recognition software to convert the audio data into text data and analyze the user's voice characteristics to identify their emotional state. The input for this step is audio data, and the output is text data containing metadata indicating the user's emotional state.
[0557] Step 3:
[0558] The server uses a generative AI model based on the user's current emotional state and acquired text data to calculate the optimal route to the destination. This calculation incorporates real-time traffic and weather information. Inputs include the user's emotional state, location, and traffic information, while the output is data on the optimal route based on emotions.
[0559] Step 4:
[0560] The server sends the calculated optimal route data to the terminal. The terminal receives this data, visually displays the route on a map, and prepares audio guidance. The input for this step is the optimal route data, and the output is the visual and audio guidance information to be presented to the user.
[0561] Step 5:
[0562] The device adjusts and presents route guidance content according to the user's emotional state. Specifically, it uses a calmer voice tone and provides navigation instructions to users who are feeling stressed. The input for this step is the user's emotional state and route guidance information, and the output is route guidance optimized for the user's psychological state.
[0563] Step 6:
[0564] While receiving route guidance, users can provide additional instructions to the terminal if necessary. The server receives the new instructions, performs sentiment analysis and route calculation again, and proposes an updated route. The input for this step is the additional instructions, and the output is the updated optimal route information.
[0565] (Application Example 2)
[0566] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0567] In autonomous vehicles, there is a need for personalized navigation and adjustments to the in-vehicle environment that take into account the emotional state of passengers. However, conventional systems have difficulty providing appropriate route guidance and in-vehicle settings that reflect the psychological state of passengers, and may lack passenger comfort. The present invention aims to solve these problems and provide a flexible travel experience that responds to the emotions of passengers.
[0568] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0569] In this invention, the server includes means for receiving instructions from the user via voice or text, means for analyzing and understanding the user's emotions, and means for acquiring travel-related and weather information in real time. This enables highly personalized settings based on the user's emotions, providing a flexible and comfortable travel experience.
[0570] "Means of receiving user instructions via voice or text" refers to an interface that allows users to send instructions to the system via voice or text.
[0571] "Means for analyzing and understanding user emotions" refers to technologies that identify and understand emotional states from user voice and other input data.
[0572] "Means for recording and obtaining a user's current location information and past behavioral history" refers to technologies that store a user's current location and past behavioral records, and allow access to them as needed.
[0573] "Means for obtaining travel-related and weather information in real time" refers to technologies that instantly collect the latest data, such as road conditions and weather, and provide it to the system.
[0574] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a technique that uses machine learning or other algorithms to calculate the optimal travel route from collected data.
[0575] "Means of presenting routes in a manner appropriate to the user's emotions based on analyzed emotions" refers to a function that provides route information to the user visually or audibly in an appropriate manner according to the user's emotional state.
[0576] "Means for re-evaluating and updating routes based on additional instructions or emotional states from the user" refers to a function that modifies the initial route as needed and re-presents a more appropriate route.
[0577] This invention relates to a navigation system for autonomous vehicles incorporating an emotion engine. The system accepts user voice or text input, uses voice analysis technology to understand emotions, and acquires real-time location, traffic, and weather information. Based on this data, a server uses a generative AI model to calculate the optimal route to the destination. In doing so, the system takes the user's emotional state into consideration and personalizes the optimal route and in-vehicle environment.
[0578] The system analyzes the user's voice instructions using natural language processing software and a speech recognition engine. The hardware includes a GPS module for location acquisition and a high-performance processor for real-time data processing. The generated optimal route is presented in a voice and visual format that takes the user's emotions into consideration.
[0579] For example, if a user wants to relax when they get in, the autonomous vehicle will guide them along a scenic route and automatically adjust the in-car music and lighting to a relaxing setting. On the other hand, if they want to get home quickly, the vehicle will select the fastest route and provide voice-guided traffic information.
[0580] An example of a prompt message might be, "Please tell me about an AI model that uses an emotion engine to select routes and adjust the in-car environment according to the passenger's emotions."
[0581] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0582] Step 1:
[0583] The user inputs instructions into the device via voice or text. The device converts the input voice data into text data using speech recognition software and passes it to an emotion analysis engine to extract the user's requests and emotions from their utterances.
[0584] Step 2:
[0585] The device uses an emotion analysis engine to analyze the user's emotions from voice data and retrieves the results. The analyzed emotion data is output as states such as relaxed, hurried, and stressed. These analysis results influence the subsequent optimal path calculation.
[0586] Step 3:
[0587] The server receives the user's current location, emotional state, and past behavioral history sent from the terminal. The server uses a GPS module and a database to verify this information and obtain real-time traffic and weather information. Based on this information, the generative AI model calculates the optimal route to the destination.
[0588] Step 4:
[0589] The server inputs emotional states and real-time information into a generating AI model to create the optimal route. The model selects routes with better scenery than usual, the fastest routes, etc., and outputs them as route data. This creates navigation optimized for the user's requests and emotions.
[0590] Step 5:
[0591] The terminal receives the optimal route from the server and presents the information to the user visually and audibly. The terminal guides the user with a tone and presentation that matches the user's emotional state. For example, it uses a calm voice when the user is relaxed and conveys information efficiently when the user is in a hurry.
[0592] Step 6:
[0593] If the user gives additional instructions or if there is a change in their emotional state, the device resends information to the server. The server re-evaluates the route based on the new information and provides the updated route information to the device. This maintains flexible navigation that adapts to the situation.
[0594] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0595] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0596] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0597] [Fourth Embodiment]
[0598] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0599] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0600] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0601] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0602] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0603] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0604] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0605] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0606] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0607] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0608] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0609] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0610] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0611] As an embodiment of the present invention, a scenario is shown in which a user operates the car navigation system using a portable information terminal.
[0612] The user gets into the car and activates a portable information terminal. The terminal allows the user to specify the destination and driving conditions via voice or text input. For example, the user might voice-input "I want to avoid traffic jams." This input is analyzed by the terminal and sent to the server as text data.
[0613] The device obtains the user's current location information using GPS and references their past activity history. This information, along with real-time traffic and weather information, is sent to the server. The server integrates and analyzes this data and calculates the optimal route using a generative model.
[0614] The server calculates the optimal route, which is then sent to the terminal, where it is presented to the user visually and audibly. The information presented includes estimated travel time to the destination, distance, and potential stops along the way. Once the user approves the proposed route, the terminal begins navigation. During navigation, the terminal monitors traffic conditions in real time, re-evaluating and updating the route as needed, and providing the user with updated information.
[0615] As a concrete example, consider a situation where a user wants to go to a specific cafe, but the route shown on their car's navigation system is congested. The system receives a voice command saying, "I would like an alternative route that avoids the current congestion," and the server recalculates a new route that avoids the traffic. Based on this result, this route is sent back to the terminal and presented to the user. The user can then select this route and arrive at their destination cafe smoothly, avoiding the congestion.
[0616] Thus, the system of the present invention flexibly reflects user instructions and provides individually optimized routes, thereby supporting users in comfortably reaching their destinations.
[0617] The following describes the processing flow.
[0618] Step 1:
[0619] The user activates their mobile information terminal and inputs their destination and desired conditions via voice or text. For example, they might say, "I want to go to a nearby cafe to avoid traffic."
[0620] Step 2:
[0621] The device converts voice input into text data and uses natural language processing to analyze the user's intent. Specifically, it extracts destination categories and conditions (such as avoiding traffic jams).
[0622] Step 3:
[0623] The device uses GPS to obtain the user's current location and sends the location information, the user's past activity history, and the analyzed instructions to the server.
[0624] Step 4:
[0625] Based on the data received by the server, potential destinations near the current location are collected from a database and external APIs. This includes information such as nearby cafes.
[0626] Step 5:
[0627] The server acquires traffic and weather information in real time and uses a generative model to calculate the optimal route that reflects the user's desired conditions.
[0628] Step 6:
[0629] The server sends optimal route information to the terminal and creates multiple route options, including details such as travel time and distance.
[0630] Step 7:
[0631] The device presents route options to the user and prompts them to start navigation. Navigation begins based on the route selected by the user.
[0632] Step 8:
[0633] If the user gives a voice command such as "Find a shorter route" during navigation, the device will send a new command to the server.
[0634] Step 9:
[0635] The server receives additional instructions and recalculates the route, incorporating current real-time information. It then sends the new optimal route to the terminal.
[0636] Step 10:
[0637] The device presents the user with a recalculated route and updates the navigation based on the new route if selected.
[0638] (Example 1)
[0639] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0640] In today's transportation environment, users often face difficulties in reaching their destinations smoothly. Conventional car navigation systems struggle to suggest optimal routes that take into account real-time traffic information and individual user preferences. Furthermore, they cannot reflect the user's past travel history or preferences when setting routes, which can result in inefficient route selection. There is a need for navigation systems that can solve these problems and enable users to reach their destinations comfortably.
[0641] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0642] In this invention, the server includes voice analysis means and text analysis means for receiving instructions from the user, location estimation means for acquiring current location information and past behavioral history, and information acquisition means for acquiring traffic information and weather information in real time. This makes it possible to calculate and present the optimal route that takes into account the user's individual preferences and current traffic conditions.
[0643] "Voice analysis means" refers to a function or device for converting a user's voice instructions into digital data and analyzing the intended meaning.
[0644] "Text analysis means" refers to a function or device for analyzing text data entered by a user and understanding the content of the instructions.
[0645] "Location estimation means" refers to a function or device that uses GPS technology or similar to determine the user's current location and acquire their past activity history.
[0646] "Information acquisition means" refers to a function or device for collecting traffic information and weather information in real time from external information sources.
[0647] A "generative model based on deep learning" is a model that uses multi-layer neural network technology to analyze data and generate the optimal path.
[0648] A "route calculation means" is a function or device for calculating the optimal route to a destination based on acquired information.
[0649] "Presentation means" refers to a function or device for presenting calculated route information to the user visually or audibly.
[0650] A "re-evaluation means" is a function or device that receives additional instructions from the user, re-evaluates existing routes, and updates the routes as necessary.
[0651] The user utilizes a portable information terminal to implement the navigation system of the present invention. The user, upon entering the vehicle, activates the portable terminal and inputs the destination and driving conditions via voice or text. The terminal analyzes this user input using voice analysis software (e.g., a speech recognition API) and processes it as text. The analyzed text data is then linked with a location estimation means and transmitted to a server along with GPS information and past activity history.
[0652] The server integrates the received information and calculates the optimal route using a generative AI model based on deep learning. Real-time traffic and weather information is obtained from external information services (for example, map API services) and is also included in the analysis. Once the optimal route is calculated, the server sends this route information back to the terminal.
[0653] The route information sent is displayed on the terminal and guided to the user using speech synthesis software (e.g., a speech synthesis API). Navigation begins once the user confirms and approves the route. While the user is moving, the terminal and server communicate continuously, and if new traffic information is acquired, the route is re-evaluated based on that information, and updated information is provided to the user.
[0654] As a concrete example, a user might instruct the system by voice, "I want to avoid traffic." This input is analyzed by the server, and a new, optimized route is calculated by a generative AI model. This route is then presented to the user's device.
[0655] An example of a prompt for a generative AI model is, "If a user requests to avoid traffic congestion, please tell me how to calculate the optimal route considering the current traffic conditions." This enables flexible guidance tailored to the user's needs.
[0656] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0657] Step 1:
[0658] The user activates their mobile device and inputs their destination and driving conditions via voice or text. The user's input is converted into text data by a voice analysis system. The converted text data is temporarily stored on the device. Here, input is voice or text, and output is text data for analysis. Voice analysis software (e.g., speech recognition API) is used to convert voice to text.
[0659] Step 2:
[0660] The device uses GPS technology to obtain the user's current location. Furthermore, it accesses and obtains the user's past activity history. The current location data and activity history data serve as input for transmission to the server. The output is integrated location data used by the server. Specifically, a location estimation system collects GPS data in real time and organizes it.
[0661] Step 3:
[0662] The server integrates text data sent from the terminal, current location, past activity history, and real-time traffic and weather information obtained via external APIs. The integrated data is input into a generative AI model. The output here is a dataset for predictive analysis by the generative AI model. The server retrieves necessary information from various external information provision services (e.g., map API services), structures the data, and prepares it for input into the model.
[0663] Step 4:
[0664] The generative AI model calculates the optimal path. This process involves simulating the optimal path based on an integrated dataset. The input is the integrated dataset, and the output is information about the optimal path. This step is performed on a server using deep learning techniques.
[0665] Step 5:
[0666] The server sends the calculated optimal route to the terminal. The terminal presents the received route information to the user using a map display application or speech synthesis software. The input is the optimal route information, and the output is navigation information presented to the user visually and audibly. The terminal provides information to the user through both visual map display and voice guidance.
[0667] Step 6:
[0668] Once the user approves the route, the device begins navigation. During navigation, the device continues to communicate with the server, re-evaluating the route based on changes in traffic and weather conditions, and updating it as needed. Input is real-time traffic information, and output is updated route information. The device receives real-time updated information and provides it to the user, correcting it as appropriate.
[0669] (Application Example 1)
[0670] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] There is a need for a way for drivers of autonomous vehicles to intuitively obtain optimal route information to reach their destination. Furthermore, a challenge is to realize a system that provides drivers with visual and audio guidance while responding to real-time changes in traffic and weather conditions.
[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0673] In this invention, the server includes means for receiving instructions from the user via voice or text, means for recording and acquiring the user's current location information and past activity history, and means for acquiring traffic and weather information in real time. This allows the driver to visually confirm route information using an additional device and always reach their destination via the optimal route.
[0674] "Means for receiving user instructions via voice or text" refers to an interface for receiving user requests for destinations and routes via voice or text input and processing them within the system.
[0675] "Means for recording and acquiring a user's current location information and past behavioral history" refers to technologies that use GPS or similar methods to acquire a user's real-time location and record and refer to historical data such as places visited in the past.
[0676] "Means for obtaining real-time traffic and weather information" refers to technologies that continuously acquire the latest traffic conditions and weather data via a network.
[0677] "Means for calculating the optimal route to the destination based on the acquired information using a generative model" refers to a method of calculating the optimal travel route based on collected data using machine learning or AI technology.
[0678] "Means for presenting the calculated optimal route to the user" refers to technical means for providing route information to the user through displays, audio, etc.
[0679] "Means for re-evaluating and updating routes based on additional instructions from the user" refers to a method for recalculating and updating routes in real time in response to new route information and changes in conditions provided by the user.
[0680] "Means of providing route information as visual information through the driver's auxiliary device to assist route guidance" refers to methods of intuitively presenting route information to the driver visually using smart glasses or head-mounted displays.
[0681] As an embodiment of this invention, a navigation system for an autonomous vehicle utilizing a smart device is specifically described. The user can wear smart glasses equipped in the autonomous vehicle and obtain comfortable route guidance while visually checking traffic information in real time.
[0682] The server first receives voice or text input from the user and interprets the destination information and driving conditions. Google Speech-to-Text can be used as the speech recognition API for this purpose. Next, the server obtains the user's current location using a GPS module and accesses past activity history stored on a cloud server. Furthermore, a network connection is required to collect traffic and weather information from the internet in real time.
[0683] Using a generative AI model (e.g., Google AI), the server calculates the optimal route from this information. The calculated route is then presented to the user as visual information through smart glasses. For example, if a user requests a route that is not congested, the server will present a new route that reflects real-time congestion. An example of a prompt in this case would be: "The user wants to go to a specific destination via a route that is not congested. Please provide the optimal route considering the current traffic conditions."
[0684] This system allows drivers to arrive at their destination via the optimal route at all times, while receiving visual and audible feedback. By flexibly responding to real-time changing traffic and weather conditions, it supports a comfortable travel experience for users.
[0685] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0686] Step 1:
[0687] The terminal starts up, and the user inputs the destination and driving conditions by voice or text. The input is received via the terminal's microphone or touchscreen. The terminal converts this into text data using a speech recognition API.
[0688] Step 2:
[0689] The device uses a GPS module to obtain the user's current location information. In addition, it retrieves past activity history from a cloud server and uses it as reference information for movement patterns and destinations. The input is current location information and past data, and the output is the user's location and associated history.
[0690] Step 3:
[0691] The server collects real-time traffic and weather information via the internet. This allows for monitoring of current traffic conditions and weather. Input is information from external APIs, and output is the latest traffic and weather data.
[0692] Step 4:
[0693] The server uses a generative AI model to calculate the optimal route based on user information and external data obtained so far. It utilizes prompts to personalize the route according to user requests. Specifically, it generates a route based on instructions such as, "The user wants to go to a specific destination via a 'congested' route." The output is optimized route information.
[0694] Step 5:
[0695] The calculated optimal route is transmitted to the terminal and presented to the user visually and audibly. The user can confirm their direction of travel through maps and route information displayed on smart glasses. The input is route information from the server, and the output is visual and audible guidance.
[0696] Step 6:
[0697] If the user provides additional instructions, the server re-evaluates the route and updates it as needed. This enables route guidance that adapts to dynamic changes in circumstances. The re-evaluated route is sent back to the terminal. The input is the user's new instructions, and the output is the updated route information.
[0698] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0699] An embodiment of the present invention is shown below: a car navigation system incorporating an emotion engine. This system operates when a user receives route guidance using a portable information terminal.
[0700] The user activates a portable information terminal and provides instructions to the system regarding destination and route conditions via voice or text. The system incorporates an emotion engine that analyzes the user's emotions through voice input, allowing it to understand the user's emotional state based on the input voice data.
[0701] The emotion engine adjusts the generative model to prioritize scenic routes when the user is relaxed, for example, and to calculate the fastest route when the user is in a hurry or stressed. The server provides comprehensive route suggestions, taking into account the user's current location, past behavior history, and real-time traffic and weather information.
[0702] The calculated optimal route is transmitted to the terminal, which then presents this information to the user visually and audibly. The presentation method can also be adjusted according to the user's emotional state. For example, if the user is stressed, the information may be conveyed in a calm voice.
[0703] For example, if a user is in a hurry because they are running late for their scheduled departure time and gives the voice command, "I absolutely must avoid traffic," the system will recognize the urgency through its emotion engine, and the server will calculate a route that prioritizes speed. This calculated route information will be quickly presented to the user through the terminal, using a calm voice tone to reduce stress.
[0704] Thus, the system of the present invention, which incorporates an emotion engine, further personalizes the car navigation experience according to the user's psychological state, supporting more flexible and comfortable travel.
[0705] The following describes the processing flow.
[0706] Step 1:
[0707] The user activates their portable information terminal and inputs a voice command saying, "I want to go to my destination while avoiding the highway." At this time, the terminal records the voice data.
[0708] Step 2:
[0709] The device converts voice input into text data and simultaneously uses an emotion engine to analyze the user's emotional state from the tone and speed of their voice. The analysis result might be determined to indicate, for example, "the user is irritated."
[0710] Step 3:
[0711] The device obtains its current location information via GPS and collects data on the user's past behavior. This information, along with the results of sentiment analysis, is sent to the server.
[0712] Step 4:
[0713] Based on the server's received location information, the user's desired route conditions, and emotional state, real-time traffic and weather information is obtained.
[0714] Step 5:
[0715] The server uses a generative model to calculate the optimal route to the destination. This calculation takes into account the user's emotional state, prioritizing routes that, for example, allow for quick and stress-reducing travel.
[0716] Step 6:
[0717] The server calculates the optimal route and sends it to the terminal, then prepares detailed route information.
[0718] Step 7:
[0719] The device presents the route information it has received to the user. Depending on the user's emotional state, it provides guidance in a calm and reassuring voice, for example.
[0720] Step 8:
[0721] The user confirms the suggested route and begins navigation. If the user gives a new instruction during navigation, such as "I want to arrive sooner," the device sends this information back to the server.
[0722] Step 9:
[0723] The server receives the new instructions and, as in the previous step, re-evaluates the user's emotions and requests, and recalculates a more optimized path.
[0724] Step 10:
[0725] The device will suggest a new route, and if selected, it will immediately update the navigation to help the user have the most satisfying journey.
[0726] (Example 2)
[0727] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0728] Conventional car navigation systems typically suggest the optimal route without considering the user's emotional state. This lack of flexible guidance, adapted to the user's psychological condition, can increase stress. Furthermore, they lack methods to adjust guidance based on emotional state, highlighting the need for improved user experience.
[0729] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0730] In this invention, the server includes means for receiving instructions from the user in voice or text, means for analyzing the user's emotional state, and means for recording and acquiring the user's current location information and past behavioral history. This enables the personalization of the optimal route based on the user's emotional state, and route guidance that is appropriate to their psychological state.
[0731] "Means for receiving user instructions via voice or text" refers to a function that provides an interface for users to specify destination and route conditions to the system via voice or text.
[0732] "Means for analyzing a user's emotional state" refers to technologies or devices used to identify a user's emotional state from voice data or user interactions. This makes it possible to understand the user's psychological state.
[0733] "Means for recording and acquiring a user's current location information and past behavioral history" refers to a device or method for acquiring a user's location information in real time, storing it, and analyzing and utilizing past movement patterns.
[0734] "Means for obtaining real-time traffic and weather information" refers to technologies or services that collect the latest traffic and weather conditions from external sources and provide them to a system.
[0735] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a method or apparatus that uses AI technology or the like to calculate the most suitable travel route for the user based on the acquired information.
[0736] "A means of presenting a calculated optimal route to the user and adjusting the presented content to the user's emotional state" refers to a technology that presents calculated route information to the user using visual and auditory methods, and applies communication methods that correspond to the user's emotions at that time.
[0737] "Means for re-evaluating and updating routes based on additional instructions from users" refers to a mechanism that reviews existing routes and updates them to the optimal ones in response to new requests or changes in conditions from users.
[0738] The system based on this invention realizes an advanced car navigation system equipped with emotion analysis capabilities. The system uses the user's mobile device as its primary interface and accepts instructions from the user via voice or text.
[0739] The terminal uses hardware such as a voice input device and a touchscreen to acquire information about the user's destination and route conditions. The acquired voice data is sent to a server equipped with an emotion analysis engine. This server uses speech recognition software (e.g., a common speech recognition API) to analyze the user's emotional state from their voice. In this process, by analyzing features such as tone of voice, word choice, and speaking speed, the server identifies the user's emotional state, such as whether they are relaxed, in a hurry, or stressed.
[0740] Once the emotional state is identified, the server uses a generative AI model to calculate the optimal route based on it. The generative AI model takes into account real-time traffic information (e.g., traffic information APIs), weather information, the user's current location, and past behavioral history. As a result, an efficient and personalized route that matches the user's emotional state is proposed.
[0741] The calculated route information is transmitted to the terminal and presented visually as a route on a map and audibly as a guide. This presentation is adjusted according to the user's emotional state; for example, a user experiencing stress will be provided with guidance in a calm voice.
[0742] For example, if a user gives a voice command requesting the "fastest route while avoiding traffic," the terminal sends this command to a server, which uses an emotion analysis engine to recognize the user's sense of urgency. Then, using a generative AI model, it calculates the fastest route while avoiding traffic and hands it over to the terminal. The terminal then presents this to the user in a calm voice, alleviating the user's anxiety.
[0743] An example of a prompt is, "Design a car navigation system that uses an emotion engine to suggest the best traffic route when the user is in a hurry." This prompt is used to train a generative AI model and design its responses.
[0744] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0745] Step 1:
[0746] The terminal accepts voice or text input from the user. Specifically, if the user specifies a destination as a "major landmark," the terminal acquires that voice or text data. The input in this step is the user's instruction data, and the output is digital data obtained by the server by converting this data into a parseable format.
[0747] Step 2:
[0748] The terminal sends the acquired audio data to the server. The server uses speech recognition software to convert the audio data into text data and analyze the user's voice characteristics to identify their emotional state. The input for this step is audio data, and the output is text data containing metadata indicating the user's emotional state.
[0749] Step 3:
[0750] The server uses a generative AI model based on the user's current emotional state and acquired text data to calculate the optimal route to the destination. This calculation incorporates real-time traffic and weather information. Inputs include the user's emotional state, location, and traffic information, while the output is data on the optimal route based on emotions.
[0751] Step 4:
[0752] The server sends the calculated optimal route data to the terminal. The terminal receives this data, visually displays the route on a map, and prepares audio guidance. The input for this step is the optimal route data, and the output is the visual and audio guidance information to be presented to the user.
[0753] Step 5:
[0754] The device adjusts and presents route guidance content according to the user's emotional state. Specifically, it uses a calmer voice tone and provides navigation instructions to users who are feeling stressed. The input for this step is the user's emotional state and route guidance information, and the output is route guidance optimized for the user's psychological state.
[0755] Step 6:
[0756] While receiving route guidance, users can provide additional instructions to the terminal if necessary. The server receives the new instructions, performs sentiment analysis and route calculation again, and proposes an updated route. The input for this step is the additional instructions, and the output is the updated optimal route information.
[0757] (Application Example 2)
[0758] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0759] In autonomous vehicles, there is a need for personalized navigation and adjustments to the in-vehicle environment that take into account the emotional state of passengers. However, conventional systems have difficulty providing appropriate route guidance and in-vehicle settings that reflect the psychological state of passengers, and may lack passenger comfort. The present invention aims to solve these problems and provide a flexible travel experience that responds to the emotions of passengers.
[0760] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0761] In this invention, the server includes means for receiving instructions from the user via voice or text, means for analyzing and understanding the user's emotions, and means for acquiring travel-related and weather information in real time. This enables highly personalized settings based on the user's emotions, providing a flexible and comfortable travel experience.
[0762] "Means of receiving user instructions via voice or text" refers to an interface that allows users to send instructions to the system via voice or text.
[0763] "Means for analyzing and understanding user emotions" refers to technologies that identify and understand emotional states from user voice and other input data.
[0764] "Means for recording and obtaining a user's current location information and past behavioral history" refers to technologies that store a user's current location and past behavioral records, and allow access to them as needed.
[0765] "Means for obtaining travel-related and weather information in real time" refers to technologies that instantly collect the latest data, such as road conditions and weather, and provide it to the system.
[0766] "Means for calculating the optimal route to a destination based on the acquired information using a generative model" refers to a technique that uses machine learning or other algorithms to calculate the optimal travel route from collected data.
[0767] "Means of presenting routes in a manner appropriate to the user's emotions based on analyzed emotions" refers to a function that provides route information to the user visually or audibly in an appropriate manner according to the user's emotional state.
[0768] "Means for re-evaluating and updating routes based on additional instructions or emotional states from the user" refers to a function that modifies the initial route as needed and re-presents a more appropriate route.
[0769] This invention relates to a navigation system for autonomous vehicles incorporating an emotion engine. The system accepts user voice or text input, uses voice analysis technology to understand emotions, and acquires real-time location, traffic, and weather information. Based on this data, a server uses a generative AI model to calculate the optimal route to the destination. In doing so, the system takes the user's emotional state into consideration and personalizes the optimal route and in-vehicle environment.
[0770] The system analyzes the user's voice instructions using natural language processing software and a speech recognition engine. The hardware includes a GPS module for location acquisition and a high-performance processor for real-time data processing. The generated optimal route is presented in a voice and visual format that takes the user's emotions into consideration.
[0771] For example, if a user wants to relax when they get in, the autonomous vehicle will guide them along a scenic route and automatically adjust the in-car music and lighting to a relaxing setting. On the other hand, if they want to get home quickly, the vehicle will select the fastest route and provide voice-guided traffic information.
[0772] An example of a prompt message might be, "Please tell me about an AI model that uses an emotion engine to select routes and adjust the in-car environment according to the passenger's emotions."
[0773] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0774] Step 1:
[0775] The user inputs instructions into the device via voice or text. The device converts the input voice data into text data using speech recognition software and passes it to an emotion analysis engine to extract the user's requests and emotions from their utterances.
[0776] Step 2:
[0777] The device uses an emotion analysis engine to analyze the user's emotions from voice data and retrieves the results. The analyzed emotion data is output as states such as relaxed, hurried, and stressed. These analysis results influence the subsequent optimal path calculation.
[0778] Step 3:
[0779] The server receives the user's current location, emotional state, and past behavioral history sent from the terminal. The server uses a GPS module and a database to verify this information and obtain real-time traffic and weather information. Based on this information, the generative AI model calculates the optimal route to the destination.
[0780] Step 4:
[0781] The server inputs emotional states and real-time information into a generating AI model to create the optimal route. The model selects routes with better scenery than usual, the fastest routes, etc., and outputs them as route data. This creates navigation optimized for the user's requests and emotions.
[0782] Step 5:
[0783] The terminal receives the optimal route from the server and presents the information to the user visually and audibly. The terminal guides the user with a tone and presentation that matches the user's emotional state. For example, it uses a calm voice when the user is relaxed and conveys information efficiently when the user is in a hurry.
[0784] Step 6:
[0785] If the user gives additional instructions or if there is a change in their emotional state, the device resends information to the server. The server re-evaluates the route based on the new information and provides the updated route information to the device. This maintains flexible navigation that adapts to the situation.
[0786] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0787] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0788] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0789] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0790] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0791] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0792] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0793] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0794] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0795] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0796] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0797] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0798] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0799] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0800] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0801] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0802] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0803] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0804] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0805] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0806] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0807] The following is further disclosed regarding the embodiments described above.
[0808] (Claim 1)
[0809] A means of receiving instructions from the user via voice or text,
[0810] Means for recording and acquiring the user's current location information and past activity history,
[0811] A means of obtaining real-time traffic and weather information,
[0812] A means for calculating the optimal route to the destination based on the acquired information using a generative model,
[0813] A means of presenting the calculated optimal path to the user,
[0814] A means of re-evaluating and updating the route based on additional instructions from the user,
[0815] A system that includes this.
[0816] (Claim 2)
[0817] The system according to claim 1, wherein the generative model personalizes the route based on the user's past behavioral history and preferences.
[0818] (Claim 3)
[0819] The system according to claim 1, wherein the acceptance of user instructions and the presentation of routes are performed in a voice dialogue format.
[0820] "Example 1"
[0821] (Claim 1)
[0822] A voice analysis means and a text analysis means for receiving instructions from the user,
[0823] A location estimation means for obtaining current location information and past behavioral history,
[0824] A means of acquiring information for obtaining real-time traffic and weather information,
[0825] A path calculation means for calculating the optimal path to the destination based on the acquired information using a generative model based on deep learning,
[0826] A presentation means that presents the calculated optimal path visually and audibly,
[0827] A re-evaluation mechanism that accepts additional instructions from the user in order to re-evaluate and update the route,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, wherein the generative model based on deep learning optimizes the route based on the user's past behavior history and preferences.
[0831] (Claim 3)
[0832] The system according to claim 1, wherein the acceptance of user instructions and the presentation of routes are performed in a two-way voice dialogue format.
[0833] "Application Example 1"
[0834] (Claim 1)
[0835] A means of receiving instructions from the user via voice or text,
[0836] Means for recording and acquiring the user's current location information and past activity history,
[0837] A means of obtaining real-time traffic and weather information,
[0838] A means for calculating the optimal route to the destination based on the acquired information using a generative model,
[0839] A means of presenting the calculated optimal path to the user,
[0840] A means of re-evaluating and updating the route based on additional instructions from the user,
[0841] A means of providing route information as visual information through an additional device for the driver, and assisting in route guidance,
[0842] A system that includes this.
[0843] (Claim 2)
[0844] The system according to claim 1, wherein the generative model personalizes the route based on the user's past behavioral history and preferences.
[0845] (Claim 3)
[0846] The system according to claim 1, wherein the acceptance of user instructions and the presentation of routes are performed in a voice dialogue format.
[0847] "Example 2 of combining an emotion engine"
[0848] (Claim 1)
[0849] A means of receiving instructions from the user via voice or text,
[0850] A means of analyzing the user's emotional state,
[0851] Means for recording and acquiring the user's current location information and past activity history,
[0852] A means of obtaining real-time traffic and weather information,
[0853] A means for calculating the optimal route to the destination based on the acquired information using a generative model based on the user's emotional state,
[0854] A means of presenting the calculated optimal path to the user and adjusting the presented content according to the user's emotional state,
[0855] A means of re-evaluating and updating the route based on additional instructions from the user,
[0856] A system that includes this.
[0857] (Claim 2)
[0858] The system according to claim 1, wherein the generative model personalizes the route based on the user's past behavioral history and preferences, as well as their current emotional state.
[0859] (Claim 3)
[0860] The system according to claim 1, wherein the acceptance of user instructions and the presentation of routes are performed in a voice dialogue format that corresponds to the user's emotional state.
[0861] "Application example 2 when combining with an emotional engine"
[0862] (Claim 1)
[0863] A means of receiving instructions from the user via voice or text,
[0864] A means of analyzing and understanding user emotions,
[0865] Means for recording and acquiring the user's current location information and past activity history,
[0866] A means of obtaining information related to movement and weather information in real time,
[0867] A means for calculating the optimal route to the destination based on the acquired information using a generative model based on the analyzed emotions,
[0868] A means of presenting paths in a way that responds to the user's emotions,
[0869] Means for re-evaluating and updating the path based on additional instructions or emotional states from the user,
[0870] A system that includes this.
[0871] (Claim 2)
[0872] The system according to claim 1, wherein the generation model personalizes the route based on the user's past behavioral history and preferences, and further adjusts the in-car environment according to the user's emotional state.
[0873] (Claim 3)
[0874] The system according to claim 1, wherein the acceptance of user instructions and the presentation of routes are performed in a voice dialogue format based on the user's emotional state. [Explanation of Symbols]
[0875] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving instructions from the user via voice or text, Means for recording and acquiring the user's current location information and past activity history, A means of obtaining real-time traffic and weather information, A means for calculating the optimal route to the destination based on the acquired information using a generative model, A means of presenting the calculated optimal path to the user, A means of re-evaluating and updating the route based on additional instructions from the user, A system that includes this.
2. The system according to claim 1, wherein the generative model personalizes the route based on the user's past behavioral history and preferences.
3. The system according to claim 1, wherein the acceptance of user instructions and the presentation of routes are performed in a voice dialogue format.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A