system
The system addresses the limitations of current car navigation systems by using a generative model and vehicle sensors to provide interactive and responsive navigation, ensuring user safety and convenience through real-time anomaly detection and countermeasures.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Current car navigation systems lack interactive functionality, fail to provide real-time responses to user voice commands, and are inadequate in detecting vehicle abnormalities and presenting appropriate countermeasures, compromising user convenience and safety.
A system integrating a generative model for interpreting voice inputs, a car navigation terminal for dynamic information updates, vehicle sensors for anomaly detection, and customizable interfaces to enhance user interaction and safety.
Enables real-time, interactive navigation experiences and quick, specific responses to vehicle malfunctions, improving user convenience and safety.
Smart Images

Figure 2026047862000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Current car navigation systems have basic functions such as Internet connection and voice control. However, the voice control is based on simple instructions and there is a problem that it cannot provide the interactive usage experience required by users. Also, regarding the detection of vehicle abnormalities and the presentation of countermeasures, specific and prompt responses are required, but this is not effectively carried out in the current system. Therefore, in order to improve the convenience and safety of users, a car navigation system with advanced interactive functions using generative AI is needed.
Means for Solving the Problems
[0005] The present invention solves the above problems with a system that includes a model for interpreting voice input from a user using a generation model and generating an appropriate response, a car navigation terminal, means for dynamically updating the information displayed on the car navigation terminal based on the user's voice instructions, a sensor for collecting vehicle status data in real time and detecting anomalies, and means for transmitting the anomaly data collected from the sensor to the generation model and providing an appropriate countermeasure. Furthermore, by providing means for appropriately setting and updating the vehicle's navigation route in real time, and means for providing a customizable interface based on user settings, user convenience and safety can be further enhanced.
[0006] A "generative model" is an artificial intelligence technology that analyzes voice input from a user, understands their intent, and generates an appropriate response.
[0007] A "car navigation terminal" is an electronic device installed in a vehicle that displays maps and provides route guidance.
[0008] "A means of dynamically updating" refers to a technology that changes the information displayed on a car navigation terminal in real time based on the user's voice commands.
[0009] A "sensor" is a device that collects vehicle status data in real time and detects information when an abnormality occurs.
[0010] "Abnormal data" refers to data detected by sensors regarding problems or malfunctions that prevent the vehicle from operating normally.
[0011] "Means of providing appropriate countermeasures" refers to technologies that use generative models to present users with specific countermeasures based on abnormal data.
[0012] A "navigation route" refers to the optimal path from a specific starting point to a destination.
[0013] A "customizable interface" refers to settings for the operation screen and voice guidance of a car navigation terminal that can be changed according to the user's preferences. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention relates to a car navigation system that uses a generative model to achieve advanced interactive functions. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination.
[0036] System Configuration
[0037] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[0038] Car navigation terminal
[0039] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[0040] Generative model
[0041] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0042] Sensors and Anomaly Detection
[0043] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0044] User Interface
[0045] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[0046] Explanation of the program's processing
[0047] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0048] 2. Terminal: Activates the speech recognition engine to convert voice commands into text data. Then sends that text data to the server.
[0049] 3. Server: Passes text data to the generative model and generates an appropriate response.
[0050] 4. Generative Model: Analyzes user intent and generates appropriate responses. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0051] 5. Server: Sends the response data from the generated model to the car navigation terminal.
[0052] 6. Device: Provides responses to the user via voice and screen display.
[0053] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[0054] 8. Terminal: Calculates the selected route and starts the guidance.
[0055] 9. Sensors: Perform continuous monitoring of the vehicle's condition.
[0056] 10. Server: If an anomaly is detected, it sends that information to the generation model to generate an appropriate countermeasure.
[0057] 11. Generative Model: Provide specific solutions tailored to the type of anomaly. Example: "The oil change light is on. Please follow these steps: (1) (2) (3)"
[0058] 12. Server: Sends the response data from the generated model to the car navigation terminal.
[0059] 13. Device: Notifies the user via voice and display and guides them through the instructed course of action.
[0060] 14. User: Follow the provided procedures to address any vehicle malfunctions.
[0061] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users and enables quick and specific responses in the event of vehicle malfunctions.
[0062] The following describes the processing flow.
[0063] Step 1:
[0064] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0065] Step 2:
[0066] Terminal: Activates the speech recognition engine to convert user voice commands from speech to text, and then sends it to the server.
[0067] Step 3:
[0068] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[0069] Step 4:
[0070] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0071] Step 5:
[0072] Server: Sends the generated response data to the car navigation terminal.
[0073] Step 6:
[0074] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[0075] Step 7:
[0076] User: Choose from the recommended options. Example: "I want to go to spot (1)."
[0077] Step 8:
[0078] Terminal: Calculates the selected route and starts navigation. Updates route information in real time.
[0079] Step 9:
[0080] Sensors: Monitor and collect various vehicle status data in real time.
[0081] Step 10:
[0082] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0083] Step 11:
[0084] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[0085] Step 12:
[0086] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[0087] Step 13:
[0088] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[0089] Step 14:
[0090] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[0091] (Example 1)
[0092] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0093] Conventional car navigation systems had problems with responding quickly and appropriately to user voice input and handling vehicle malfunctions. In particular, it was difficult to provide real-time responses to voice commands and to effectively coordinate vehicle status monitoring and anomaly detection. Furthermore, they lacked an intuitive interface that users could easily operate. As a result, the user experience was significantly reduced, and there was a risk of compromising driving safety.
[0094] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0095] In this invention, the server includes means for converting user voice into text data and transmitting it, means for passing the received text data to a generation model to generate a response, and means for transmitting the generated response to a car navigation terminal and providing it to the user in both voice and screen display. This enables real-time responses to user voice commands, vehicle status monitoring, and provision of appropriate countermeasures.
[0096] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[0097] A "car navigation terminal" is a device installed in a vehicle that takes voice commands as input, outputs generated responses, and displays route guidance.
[0098] A "speech recognition engine" refers to software or hardware that converts a user's voice into text data.
[0099] A "server" is a computing system on which the generative model operates and processes various types of data in conjunction with car navigation terminals and sensors.
[0100] A "sensor" is a device used to monitor the vehicle's condition in real time and collect data.
[0101] "Text data" refers to the character data generated by a speech recognition engine after analyzing speech.
[0102] "User interface" refers to the means, such as screens and sounds, that users use to operate a system.
[0103] "Abnormal data" refers to information about abnormal conditions in a vehicle detected by sensors.
[0104] "Handling procedures" refers to the steps and instructions for taking appropriate action in response to a vehicle malfunction.
[0105] "Response" refers to the content of the reply generated by the generative model based on the user's voice input.
[0106] This invention relates to a car navigation system that achieves advanced interactive functions using a generative model. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination. The configuration and specific operation of this system are described in detail below.
[0107] System Configuration
[0108] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[0109] Car navigation terminal
[0110] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and send it to a server.
[0111] Generative model
[0112] The server passes the text data received from the user to a generative model (e.g., OpenAI GPT-4) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0113] Specific example
[0114] When a user enters a voice command into the car navigation system, such as "I want to go for a drive to clear my head," the navigation system uses its voice recognition engine to convert the voice command into text data. This text data is then sent to a server via the internet.
[0115] The server passes the received text data to the generative model, which generates a response such as, "It's nighttime now, so we recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~". The generated response is then sent back to the car navigation terminal via the server.
[0116] The car navigation terminal uses a speech synthesis engine (e.g., Amazon Polly) and a screen display to provide responses to the user. The user selects "I want to go to spot (1)" from the recommended options, and the car navigation terminal calculates the selected route using a GPS navigation system (e.g., Google Maps API) and begins providing guidance via voice and screen.
[0117] Sensors and Anomaly Detection
[0118] Sensors installed inside the vehicle constantly monitor the vehicle's status. Vehicle status data is collected in real time and transmitted to a server. The server analyzes this data and, for example, if the oil lamp illuminates, sends that information to a generating model. The generating model then generates specific instructions for an oil change, such as "The oil change lamp is illuminated. Please follow these steps: (1) ~~ (2) ~~ (3) ~~," and transmits this information back to the car navigation terminal via the server.
[0119] The car navigation system notifies the user through voice and display, guiding them through the instructed course of action. The user then follows the suggested steps, for example, heading to the nearest gas station for an oil change.
[0120] Through the configuration and specific operation described above, this system realizes the interactive navigation experience that users desire and enables quick and specific responses in the event of a vehicle malfunction.
[0121] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0122] Step 1:
[0123] The user enters a voice command.
[0124] The specific action involves the user issuing a voice command to the car navigation terminal, such as "I want to go for a drive to clear my head." At this stage, the input is the user's voice.
[0125] Step 2:
[0126] The terminal converts voice commands into text data.
[0127] A speech recognition engine (e.g., Google Cloud Speech-to-Text) works to analyze the user's voice commands and generate text data. The input is the user's voice data, and the output is the corresponding text data.
[0128] Step 3:
[0129] The terminal sends text data to the server.
[0130] Text data converted from speech is sent from the terminal to the server via the network. The input is the text data generated by the speech recognition engine, and the output is the completion of the transmission of the text data to the server.
[0131] Step 4:
[0132] The server passes text data to the generative model.
[0133] The server provides the received text data to a generative AI model (e.g., OpenAI GPT-4). The generative model analyzes the text data and understands the user's intent. The input is the text data received by the server, and the output is the analysis request to the generative model.
[0134] Step 5:
[0135] The generative model generates the response.
[0136] The generative model generates an appropriate response based on the analysis results. For example, it might generate a response like, "Since it's nighttime, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~". The input is text data passed from the server, and the output is the generated response.
[0137] Step 6:
[0138] The server sends the generative model response to the terminal.
[0139] The server sends the response message obtained from the generative model to the car navigation terminal. The input is the response message generated by the generative model, and the output is the completion of sending that response message to the terminal.
[0140] Step 7:
[0141] The device provides responses to the user through voice and on-screen displays.
[0142] The device uses a speech synthesis engine (e.g., Amazon Polly) and a display to provide the user with generated responses as audio and visual information. The input is the response text sent from the server, and the output is the audio and screen display response to the user.
[0143] Step 8:
[0144] The user selects from the recommended options.
[0145] The user makes selections using voice or a touch panel, for example, "I want to go to spot (1)." Input is provided by the terminal display and voice guidance, while output is the user's selection.
[0146] Step 9:
[0147] The device calculates the selected route and begins providing directions.
[0148] The device uses a GPS navigation system (e.g., Google Maps API) to calculate the selected route. The input is the user's selection, and the output is the selected route and the start of navigation.
[0149] Step 10:
[0150] Sensors monitor the vehicle's status.
[0151] Sensors installed inside the vehicle monitor engine temperature, tire pressure, and other parameters in real time. Inputs are various vehicle status data, and outputs are data collected by the sensors.
[0152] Step 11:
[0153] If the server detects an anomaly, it sends information to the generative model.
[0154] The server analyzes data from the sensors and, if an anomaly is detected, sends that information to the generative model. The input is state data from the sensors, and the output is the transmission of anomaly data to the generative model.
[0155] Step 12:
[0156] The generative model generates methods for dealing with anomalies.
[0157] The generative model generates specific corrective actions based on the type of anomaly. The input is the anomaly data, and the output is a response message describing the corrective action. For example, "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[0158] Step 13:
[0159] The server sends the generative model response to the terminal.
[0160] The server sends the solution obtained from the generative model to the car navigation terminal. The input is the response message of the solution generated by the generative model, and the output is the completion of sending that response message to the terminal.
[0161] Step 14:
[0162] The device notifies the user and provides instructions.
[0163] The terminal uses voice and screen display to notify and guide the user of the solutions derived from the generative model. Input is the response text of the solution sent from the server, and output is voice and display guidance to the user.
[0164] Step 15:
[0165] The user will take action by following the provided instructions.
[0166] The user follows the instructions provided, for example, by going to the nearest gas station for an oil change. Input is the guidance from the terminal, and output is the action taken to address the vehicle problem.
[0167] (Application Example 1)
[0168] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0169] Conventional car navigation systems can only provide static responses to user voice commands, making real-time responses to vehicle abnormalities difficult. Furthermore, there has been a lack of systems capable of sophisticated user interaction in the operation of autonomous vehicles. Solving these problems is essential.
[0170] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0171] In this invention, the server includes means for interpreting voice input from a user using a generative model and generating an appropriate response; means for dynamically updating information displayed on a car navigation terminal based on the user's voice instructions; sensors for collecting vehicle status data in real time and detecting anomalies; means for transmitting anomaly data collected from the sensors to the generative model and providing appropriate countermeasures; means for calculating a route to a destination through navigation using the generative model and providing guidance to the autonomous vehicle; and means for monitoring the vehicle status in real time and guiding the user to appropriate countermeasures via the generative model when an anomaly occurs. This enables the provision of dynamic and appropriate responses to the user's voice instructions, and allows for navigation of the autonomous vehicle and rapid response to anomalies.
[0172] A "generative model" is a machine learning model that interprets voice input from a user and generates an appropriate response.
[0173] A "car navigation terminal" is a device installed in a vehicle that displays information based on the user's voice commands.
[0174] "Means" refer to methods or devices used to achieve a specific function or role.
[0175] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[0176] An "autonomous vehicle" is a vehicle that can drive autonomously without the intervention of a human driver.
[0177] "Guidance" refers to information and instructions that provide route information to a destination and guide a vehicle to its destination.
[0178] A "user interface" refers to equipment such as screens and audio output devices that allow users to operate a system and obtain information.
[0179] "Real-time" refers to actions or processes that occur immediately or with a very short delay.
[0180] This invention is a system that uses a generative model to realize advanced interactive functions. The system consists of a car navigation terminal, a server on which the generative model operates, vehicle sensors, and various interfaces.
[0181] System Configuration
[0182] Car navigation terminal
[0183] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[0184] Generative model
[0185] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0186] Sensors and Anomaly Detection
[0187] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0188] User Interface
[0189] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[0190] Operation overview
[0191] 1. Receiving and analyzing voice commands:
[0192] The car navigation system uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice commands into text data.
[0193] This text data will be sent to the server.
[0194] 2. Response generation using generative models:
[0195] The server passes the transformed text data to a generation model (e.g., the text generation model in the transformers library) to generate an appropriate response.
[0196] The generated response is sent back to the car navigation terminal via the server and provided to the user via voice and screen display.
[0197] 3. Route calculation and navigation:
[0198] The car navigation terminal calculates the optimal route to the destination and performs navigation based on the response of the generative model.
[0199] 4. Detection and response to vehicle abnormalities:
[0200] The sensors collect vehicle status data in real time and transmit it to the server.
[0201] When an anomaly is detected, the server passes that information to the generative model, which then generates an appropriate response.
[0202] The generative model provides specific solutions tailored to the type of anomaly and notifies the user via voice and display.
[0203] Specific example
[0204] For example, when a user enters a voice command on a smartphone, the following prompt message is used:
[0205] "Hey GPS, tell me some good restaurants around here."
[0206] In response to this, the system provides the following answer:
[0207] "We have two recommended restaurants nearby: (1) Restaurant A and (2) Cafe B. Which one would you like to go to?"
[0208] Once the user makes a selection, the system calculates the route to the specified destination and provides directions to the autonomous vehicle.
[0209] In this way, we can realize the interactive navigation experience that users desire and enable quick and specific responses in the event of a vehicle malfunction.
[0210] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0211] Step 1:
[0212] Receiving voice commands
[0213] Operation: The user enters voice commands into the car navigation terminal.
[0214] Input: User voice command (e.g., "Tell me some good restaurants nearby").
[0215] Output: Audio data.
[0216] Specific operation: The car navigation terminal collects the user's voice through the microphone.
[0217] Step 2:
[0218] Converting audio data to text
[0219] Operation: The device uses a speech recognition engine to convert speech data into text data.
[0220] Input: Audio data.
[0221] Output: Text data (e.g., "Tell me some good restaurants nearby").
[0222] Specific operation: Using a speech recognition library (e.g., speech_recognition), the system analyzes audio data and converts it into text data.
[0223] Step 3:
[0224] Sending text data to the server
[0225] Operation: The terminal sends the converted text data to the server.
[0226] Input: Text data.
[0227] Output: Notification that the text data has been successfully sent to the server.
[0228] Specific operation: Use the communication module in the terminal to send text data to the server.
[0229] Step 4:
[0230] Response generation using generative models
[0231] Operation: The server inputs the received text data into a generative model and generates an appropriate response.
[0232] Input: Text data.
[0233] Output: Generated response (e.g., "Recommended nearby restaurants are (1) Restaurant A and (2) Cafe B.").
[0234] Specific operation: The server inputs text data into a generative model (e.g., the text generation model in the transformers library), analyzes the user's intent, and generates an appropriate response.
[0235] Step 5:
[0236] Sending response data to the terminal
[0237] Operation: The server sends the generated response data to the terminal.
[0238] Input: Generated response data.
[0239] Output: Notification that the response data has been successfully sent to the terminal.
[0240] Specific operation: The generated response data is sent to the car navigation terminal using the communication module within the server.
[0241] Step 6:
[0242] Providing responses to users
[0243] Operation: The terminal provides the user with received response data via voice and screen display.
[0244] Input: Response data.
[0245] Output: Audio output and screen display (e.g., "Recommended nearby restaurants include (1) Restaurant A and (2) Cafe B.").
[0246] Specific operation: The device will reply using a speech synthesis engine (e.g., tts_engine) and simultaneously display it on the screen.
[0247] Step 7:
[0248] User Selection
[0249] Operation: The user makes a selection regarding the response via voice or touch.
[0250] Input: User selection (e.g., "I want to go to restaurant A (1)").
[0251] Output: Selected data.
[0252] Specific operation: The device collects user selections via voice recognition or screen touch.
[0253] Step 8:
[0254] Route calculation and navigation start
[0255] Operation: The device calculates the route to the selected destination and begins navigation.
[0256] Input: Selected data (e.g., "Restaurant A").
[0257] Output: Route data and navigation guide.
[0258] Specific operation: The device uses a navigation library (e.g., navigation) to calculate the optimal route to the destination and starts navigation.
[0259] Step 9:
[0260] Real-time monitoring of vehicle status
[0261] Operation: Sensors collect vehicle status data in real time and send it to the server.
[0262] Input: Vehicle status data.
[0263] Output: Notification that vehicle status data has been successfully sent to the server.
[0264] Specific operation: Sensors inside the vehicle collect status data in real time and send it to the server.
[0265] Step 10:
[0266] Vehicle abnormality detection
[0267] Operation: The server analyzes the received vehicle status data and detects any abnormalities.
[0268] Input: Vehicle status data.
[0269] Output: Anomaly detection notification (e.g., "Oil change required").
[0270] Specific operation: The analysis engine within the server analyzes the status data and detects anomalies.
[0271] Step 11:
[0272] Generating solutions using generative models
[0273] Operation: The server inputs the anomaly detection data into the generation model and generates an appropriate countermeasure method.
[0274] Input: Anomaly detection data.
[0275] Output: Countermeasure method (e.g., "Oil change is required. Please follow the next steps to handle it.").
[0276] Specific operation: The server inputs the anomaly detection data into the generation model and generates a specific countermeasure method for the user.
[0277] Step 12:
[0278] Transmission of the countermeasure method to the terminal
[0279] Operation: The server transmits the generated countermeasure method data to the terminal.
[0280] Input: Countermeasure method data.
[0281] Output: Notification of the completion of the transmission of the countermeasure method data to the terminal.
[0282] Specific operation: Using the communication module in the server, the generated countermeasure method data is transmitted to the car navigation terminal.
[0283] Step 13:
[0284] ]>Presentation of the countermeasure method to the user
[0285] Operation: The terminal provides the received countermeasure method data to the user through voice and screen display.
[0286] Input: Countermeasure method data.
[0287] Output: Voice output and screen display (e.g., "Oil change is required. Please follow the next steps to handle it.").
[0288] Specific operation: The device uses a speech synthesis engine to notify the user of the solution via voice, and simultaneously displays it on the screen.
[0289] Step 14:
[0290] User Action
[0291] Operation: The system will address the vehicle malfunction according to the troubleshooting steps provided to the user.
[0292] Input: Display of troubleshooting methods and voice guidance.
[0293] Output: Normal vehicle condition.
[0294] Specific actions: The user follows the provided instructions and performs the actual troubleshooting steps.
[0295] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0296] This invention combines an emotion engine with a car navigation system that interprets voice input from the user using a generative model and generates an appropriate response. The system consists of a car navigation terminal, a server, a generative model, vehicle sensors, and an emotion engine that recognizes the user's emotions. This improves the user's driving experience and enhances safety.
[0297] System configuration and operation
[0298] This system consists of a car navigation terminal, a server on which the generative model runs, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[0299] Car navigation terminal
[0300] The car navigation terminal is installed in the vehicle and is used to input voice instructions, output responses by the generation model and emotion engine, and display route guidance. When the user inputs a voice instruction, the terminal converts the voice into text data and sends it to the server.
[0301] Generation model
[0302] The server passes the text data received from the user to the generation model and generates an appropriate response. The generation model analyzes the user's intention and generates a specific and appropriate response. The generated response returns to the car navigation terminal through the server and is provided to the user.
[0303] Sensor and anomaly detection
[0304] The sensors in the vehicle collect the vehicle's state data in real time and send it to the server. The server analyzes these data and, if an anomaly is detected, sends that information to the generation model. The generation model generates an appropriate countermeasure according to the type of anomaly and guides the user through specific procedures.
[0305] Emotion engine
[0306] The emotion engine recognizes the user's emotion from the user's voice input and the in-vehicle situation. The user's emotion recognized by the emotion engine is reflected in the response of the generation model and the customization of navigation information. In particular, when signs of the user's stress or fatigue are detected, a function to recommend relaxation spots and rest places is included.
[0307] User interface
[0308] The user interface is composed of the screen and voice of the car navigation terminal and is designed to be easily operable by the user. The setting items include language, type of voice for voice guidance, theme color, etc., and can be customized according to the user's preference.
[0309] Explanation of the program's processing
[0310] Here, the system's operation is explained in natural language, describing the program's processing.
[0311] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0312] 2. Terminal: Activates the speech recognition engine to convert the user's voice commands from speech to text. This text data is then sent to the server.
[0313] 3. Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[0314] 4. Generative Model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0315] 5. Server: Sends the generated response data to the car navigation terminal.
[0316] 6. Terminal: Provides the user with the received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[0317] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[0318] 8. Terminal: Calculates the selected route and starts guidance. Updates route information in real time.
[0319] 9. Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[0320] 10. Server: Inputs data from the emotion engine into the generative model to generate a customized response based on emotion. Example: "Would you recommend a relaxation spot?"
[0321] 11. Sensors: Monitor and collect various vehicle status data in real time.
[0322] 12. Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0323] 13. Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[0324] 14. Server: Sends response data, including the solution created by the generative model, to the car navigation terminal.
[0325] 15. Device: Notify the user of how to resolve the issue via voice and on-screen display. Provide specific instructions.
[0326] 16. User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the user will be guided to the nearest gas station or repair shop.
[0327] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users, and further enables quick and specific responses in the event of vehicle malfunctions, thereby enhancing safety and comfort. By combining it with an emotional engine, flexible responses tailored to the user's psychological state become possible, further improving the driving experience.
[0328] The following describes the processing flow.
[0329] Step 1:
[0330] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0331] Step 2:
[0332] Terminal: Activates the speech recognition engine and converts the user's voice commands into text data. This text data is then sent to the server.
[0333] Step 3:
[0334] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[0335] Step 4:
[0336] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0337] Step 5:
[0338] Server: Sends the generated response data to the car navigation terminal.
[0339] Step 6:
[0340] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[0341] Step 7:
[0342] User: Select from the suggested options. Example: "I want to go to spot (1)."
[0343] Step 8:
[0344] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[0345] Step 9:
[0346] Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[0347] Step 10:
[0348] Server: Passes emotional data from the emotion engine to the generative model, which then generates an emotionally appropriate response. Example: "Would you recommend a relaxation spot?"
[0349] Step 11:
[0350] Generative model: Generates customized responses based on emotions. Example: "You seem stressed, so I recommend the following relaxation spots: (1)~~ (2)~~ (3)~~"
[0351] Step 12:
[0352] Server: Sends the generated response data to the car navigation terminal.
[0353] Step 13:
[0354] Terminal: Provides the user with received response data via voice output and screen display. Example: "We recommend the following relaxation spots. (1)~~ (2)~~ (3)~~ Where would you like us to take you?"
[0355] Step 14:
[0356] User: Select from the suggested relaxation spots. Example: "I want to go to relaxation spot (2)."
[0357] Step 15:
[0358] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[0359] Step 16:
[0360] Sensors: Monitor and collect various vehicle status data in real time.
[0361] Step 17:
[0362] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0363] Step 18:
[0364] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[0365] Step 19:
[0366] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[0367] Step 20:
[0368] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[0369] Step 21:
[0370] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[0371] This process allows the system to provide an interactive navigation experience while taking the user's emotional state into consideration, and enables quick and specific action in the event of a vehicle malfunction.
[0372] (Example 2)
[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0374] Conventional car navigation systems simply interpret user voice input to provide route guidance, making it difficult to respond in accordance with the user's emotions or the vehicle's malfunction. Furthermore, the lack of customized responses tailored to the user's psychological state meant that stress and fatigue during driving could not be reduced. Additionally, there was a problem with the inability to quickly provide specific solutions in the event of a vehicle malfunction.
[0375] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, means for dynamically updating information displayed on the navigation terminal based on the user's voice instructions, sensors for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensors to the generation model and providing an appropriate countermeasure, an emotion engine for recognizing the user's emotions from the user's voice input and the situation inside the vehicle and reflecting this in the response of the generation model, and means for customizing navigation information based on the user's emotions. This enables flexible responses according to the user's psychological state and quick and specific countermeasures in the event of a vehicle abnormality.
[0376] A "generative model" is a type of artificial intelligence that analyzes input data and generates appropriate responses or predictions based on that data.
[0377] A "navigation terminal" is a device installed inside a vehicle that provides map information and route guidance.
[0378] "Voice instructions" refer to instructions or commands spoken by the user, which serve as input for the navigation system to recognize and process.
[0379] A "sensor" is a device that collects vehicle status data in real time and detects specific conditions or abnormalities.
[0380] An "emotion engine" is software or an algorithm that analyzes the user's emotions from their voice input and the situation inside the vehicle, and reflects that emotional information in the response of a generative model.
[0381] "Handling instructions" refer to specific actions and procedures that the user should take when they detect a problem or abnormality in the vehicle.
[0382] "Customization" refers to modifying the interface and responses based on the user's settings and status to meet individual needs.
[0383] Modes for carrying out the invention
[0384] The car navigation system of the present invention enhances the user's driving experience and improves safety by combining a generative model, a voice recognition engine, an emotion engine, and various sensors. The present invention consists of a navigation terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[0385] Navigation terminal
[0386] The navigation terminal is installed in the vehicle, takes user voice commands as input, outputs responses using a generative model and emotion engine, and displays route guidance. When the user inputs a voice command such as "I want to go for a drive to clear my head," the terminal uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the voice into text data and sends it to the server.
[0387] Generative model
[0388] The server passes the text data received from the user to a generative model (e.g., GPT-3) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is sent to the navigation terminal via the server.
[0389] Sensors and Anomaly Detection
[0390] The vehicle is equipped with various sensors, including tire pressure sensors and an engine diagnostic system, to collect vehicle status data in real time. This data is sent to a server, which analyzes it and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0391] Emotional Engine
[0392] The emotion engine recognizes the user's emotions from their voice input and the environment inside the vehicle. For example, if the user is feeling stressed, the emotion engine detects this and sends the data to the generative model. The generative model then generates a customized response based on the emotion data, such as recommending a relaxation spot to the user.
[0393] User Interface
[0394] The user interface consists of a screen and audio for the navigation terminal and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, which can be customized to the user's preferences.
[0395] Specific example
[0396] When a user inputs a voice command such as "I want to go for a drive to clear my head," the voice recognition engine converts the voice into text and sends it to the server. The server passes this text data to a generative model, which generates a response such as "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~." The generated response is sent to the navigation terminal, where the user confirms the recommended spots on the screen and by voice and selects a destination. The navigation terminal then calculates the selected route and begins guidance.
[0397] Thus, by utilizing a generative AI model, a speech recognition engine, an emotion engine, and sensors, the present invention enables flexible responses that respond to the user's intentions and emotions, realizing a car navigation system that achieves both safety and comfort.
[0398] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0399] Program processing flow
[0400] Step 1:
[0401] User: The user enters a voice command. Specifically, they say, "I want to go for a drive to clear my head."
[0402] Input: User's voice
[0403] Output: Raw data audio file
[0404] Step 2:
[0405] Device: Uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert user voice commands from speech to text.
[0406] Input: Raw audio file
[0407] Output: Text data (Example: "I want to go for a drive to clear my head")
[0408] Step 3:
[0409] Terminal: Sends text data to the server.
[0410] Input: Text data
[0411] Output: Sending text data to the server
[0412] Step 4:
[0413] Server: Passes the received text data to a generative model (e.g., GPT-3) and starts the analysis.
[0414] Input: Text data
[0415] Output: User intent (e.g., "Recommendations for destinations suitable for a change of pace")
[0416] Step 5:
[0417] Generative model: Generates appropriate responses based on the user's intent. For example, it might create a response like, "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~".
[0418] Input: User intent
[0419] Output: Response text (Example: "Since it's nighttime now, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~")
[0420] Step 6:
[0421] Server: Sends the generated response text data to the navigation terminal.
[0422] Input: Response text data
[0423] Output: Sending response data to the navigation terminal
[0424] Step 7:
[0425] Terminal: Provides the user with the received response text via voice output (e.g., Amazon Polly) and screen display. Specifically, it would say, "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[0426] Input: Response text data
[0427] Output: Voice guidance and screen display
[0428] Step 8:
[0429] User: The user selects from the provided options and responds, "I want to go to spot (1)."
[0430] Input: User Selection
[0431] Output: Selection (Example: "(1) Spot")
[0432] Step 9:
[0433] Terminal: Calculates the selected route using the car navigation system's route calculation engine (e.g., Google Maps API) and begins guidance. Route information is updated in real time.
[0434] Input: Selection
[0435] Output: Route guidance and real-time updates
[0436] Step 10:
[0437] Emotion Engine: Analyzes user voice input and in-car environment to recognize user emotions. Specifically, it determines whether the user is experiencing stress.
[0438] Input: User voice, in-vehicle status data
[0439] Output: Recognized emotion (e.g., "I am feeling stressed")
[0440] Step 11:
[0441] Server: Inputs data from the emotion engine into the generative model to generate customized responses based on the user's emotions. For example, it might create a response such as, "Would you recommend a relaxation spot?"
[0442] Input: Sentiment data
[0443] Output: Response text data (e.g., "Do you recommend any relaxation spots?")
[0444] Step 12:
[0445] Sensors: Collect various status data (tire pressure, engine diagnostics, etc.) in real time.
[0446] Input: Vehicle condition
[0447] Output: Status data
[0448] Step 13:
[0449] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0450] Input: Status data from the sensor
[0451] Output: Abnormal data and notifications
[0452] Step 14:
[0453] Generative Model: Based on abnormal data, it generates appropriate countermeasures and creates instructions such as, "The oil change light is on. Please take the following steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[0454] Input: Abnormal data
[0455] Output: Text data of the solution.
[0456] Step 15:
[0457] Server: Sends response data containing the solution generated by the generative model to the navigation terminal.
[0458] Input: Text data of the solution
[0459] Output: Send to navigation terminal
[0460] Step 16:
[0461] Terminal: Notifies the user of how to resolve the issue via voice and on-screen display, and guides them through specific steps. For example, it might say, "The oil change light is on. Please follow these steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[0462] Input: Text data of the solution
[0463] Output: Voice guidance and screen display
[0464] Step 17:
[0465] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[0466] Input: Instructions on how to proceed
[0467] Output: Appropriate countermeasures
[0468] (Application Example 2)
[0469] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0470] In recent years, the demand for food delivery services has increased, requiring drivers to perform delivery duties efficiently and safely. However, excessive stress and fatigue during driving can reduce drivers' attention span and increase the risk of accidents. Furthermore, current car navigation systems cannot provide responses that take into account the driver's emotional state, and improvements are needed to enhance the driver's driving experience. Therefore, a system is needed that analyzes the driver's emotional state in real time and generates appropriate responses.
[0471] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, an in-vehicle information terminal, means for dynamically updating the information displayed on the in-vehicle information terminal based on the user's voice instructions, a sensor for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensor to the generation model and providing an appropriate countermeasure, emotion analysis means for analyzing the user's emotional state and generating a response corresponding to that emotion, and means for suggesting rest locations such as relaxation spots according to the user's emotional state recognized by the emotion analysis means. This makes it possible to detect the stress and fatigue of the driver and provide appropriate responses and rest suggestions, thereby enabling safe and efficient delivery operations.
[0472] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[0473] An "in-vehicle information terminal" is a device installed in a vehicle that allows for the input of voice commands, the output of responses from a generative model and sentiment analysis engine, and the display of route guidance.
[0474] A "dynamically updating method" refers to a method of changing the information displayed on the in-vehicle information terminal in real time based on the user's voice commands.
[0475] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[0476] "Means of providing appropriate countermeasures" refers to a method of transmitting anomaly data collected from sensors to a generation model and then presenting specific countermeasures based on that data.
[0477] "Emotion analysis means" refers to a method that recognizes the emotional state from the user's voice input or the situation inside the vehicle and provides that information to a generative model.
[0478] A "relaxation spot" is a place where users can rest and refresh themselves when they feel stressed or tired.
[0479] This invention is a system that supports food delivery drivers in performing their duties safely and efficiently, and consists of the following components.
[0480] System configuration
[0481] This system consists of an in-vehicle information terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion analysis engine.
[0482] In-vehicle information terminal
[0483] The in-vehicle information terminal is installed in the vehicle and handles voice command input, output of responses generated by a generative model and sentiment analysis engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to the server. For example, if a driver voice-inputs "I'm tired, I want to take a short break," the terminal converts that voice into text and sends it to the server.
[0484] Generative model
[0485] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and emotional state to generate a specific and appropriate response. The generated response is returned to the in-vehicle information terminal via the server and provided to the user. For example, in response to input such as "I'm tired, I want to take a break," a response like "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?" is generated.
[0486] Sensors and Anomaly Detection
[0487] Sensors within the vehicle collect vehicle status data in real time and send it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps. For example, if it detects that an oil change is needed, it will provide countermeasures such as, "The oil change lamp is lit. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[0488] Emotion analysis engine
[0489] The emotion analysis engine recognizes the user's emotional state from their voice input and the environment inside the vehicle. The user's emotional state recognized by the emotion analysis engine is reflected in the generative model's responses and the customization of navigation information. In particular, if the system detects that the user is feeling "tired" or "stressed," it includes a function that suggests relaxation spots and rest areas. For example, if the user inputs "tired," the system detects fatigue and suggests nearby relaxation spots.
[0490] User Interface
[0491] The user interface consists of the in-car information terminal's screen and voice, and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, all of which can be customized to the user's preferences. For example, simply saying "I want to change the voice guidance voice" will change it to the preferred voice type.
[0492] Specific example
[0493] 1. User input: "I'm tired, I want to take a break."
[0494] 2. Example of a prompt:
[0495] Emotion: Fatigue, Input: I'm tired and want to take a break. Can you suggest any relaxation spots?
[0496] 3. Response generated: "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[0497] This configuration allows for the detection of driver stress and fatigue, and enables safe and efficient delivery operations by providing appropriate responses and rest suggestions.
[0498] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0499] Step 1:
[0500] The user inputs a voice command. For example, they might say, "I'm tired, I want to take a break." This voice input becomes the starting point for the system's processing.
[0501] Step 2:
[0502] The device activates a speech recognition engine to convert the user's voice commands from speech to text. The input is the user's voice data, and the output is the text data of that voice. The device sends this text data to the server.
[0503] Step 3:
[0504] The server passes the received text data to the generative model. The input is text data sent from the terminal, and the generative model is used to analyze the user's intent and emotions. The output is data containing the analysis results.
[0505] Step 4:
[0506] The generative model generates appropriate responses based on the user's intentions and emotions. The input is the parsed data obtained in the previous step, and the output is the response sentence presented to the user. For example, it might generate a response such as, "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[0507] Step 5:
[0508] The server transmits the generated response data to the in-vehicle infotainment terminal. The input is the response data from the generative model, and the output is the communication data to the in-vehicle infotainment terminal.
[0509] Step 6:
[0510] The terminal provides the user with the received response data through voice output and screen display. The input is the response data from the server, and the output is voice guidance and screen display to the user. For example, it might provide voice guidance such as, "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[0511] Step 7:
[0512] The user selects an option from the suggested relaxation spots and gives a voice command. For example, they might respond, "I want to go to Cafe B." This voice data becomes the input for the next process.
[0513] Step 8:
[0514] The device restarts its speech recognition engine and converts the user's selection from speech to text. The input is the user's voice data, and the output is the text data of the selections. The device sends this text data to the server.
[0515] Step 9:
[0516] The server calculates the route to the selected relaxation spot based on the received text data. The input is the text data of the selected spot, and the output is the data of the optimal route.
[0517] Step 10:
[0518] The server transmits the calculated route to the in-vehicle information terminal. The input is the route calculation result data, and the output is the communication data to the in-vehicle information terminal.
[0519] Step 11:
[0520] The terminal starts navigation based on the received route data and updates the route information in real time. The input is route data from the server, and the output is navigation guidance and screen display.
[0521] Step 12:
[0522] The emotion analysis engine analyzes the user's voice input and the conditions inside the vehicle to recognize the user's emotions. The input is the user's voice data and sensor information, and the output is emotional state data.
[0523] Step 13:
[0524] The server inputs data obtained from the emotion analysis engine into a generative model to generate a customized response based on the emotion. The input is emotional state data, and the output is customized response data. For example, it can generate a response such as, "Would you recommend a relaxation spot?"
[0525] Step 14:
[0526] The sensors monitor and collect various vehicle status data in real time. The input is vehicle status information, and the output is data signals from the sensors.
[0527] Step 15:
[0528] The server continuously receives data from the sensor and, if it detects an anomaly, sends that information to the generative model. The input is the data signal from the sensor, and the output is the anomaly data sent to the generative model.
[0529] Step 16:
[0530] The generative model generates appropriate countermeasures based on abnormal data and creates specific procedures. The input is abnormal data, and the output is procedure data for the countermeasures. For example, it generates procedures such as "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[0531] Step 17:
[0532] The server sends response data, including the countermeasures created by the generative model, to the in-vehicle information terminal. The input is the procedural data of the countermeasures, and the output is the communication data to the in-vehicle information terminal.
[0533] Step 18:
[0534] The device notifies the user of troubleshooting methods via voice and on-screen display, guiding them through specific steps. Input is troubleshooting data from the server, and output is voice guidance and on-screen display to the user.
[0535] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0536] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0537] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0538] [Second Embodiment]
[0539] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0540] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0541] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0542] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0543] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0544] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0545] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0546] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0547] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0548] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0549] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0550] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0551] This invention relates to a car navigation system that uses a generative model to achieve advanced interactive functions. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination.
[0552] System Configuration
[0553] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[0554] Car navigation terminal
[0555] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[0556] Generative model
[0557] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0558] Sensors and Anomaly Detection
[0559] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0560] User Interface
[0561] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[0562] Explanation of the program's processing
[0563] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0564] 2. Terminal: Activates the speech recognition engine to convert voice commands into text data. Then sends that text data to the server.
[0565] 3. Server: Passes text data to the generative model and generates an appropriate response.
[0566] 4. Generative Model: Analyzes user intent and generates appropriate responses. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0567] 5. Server: Sends the response data from the generated model to the car navigation terminal.
[0568] 6. Device: Provides responses to the user via voice and screen display.
[0569] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[0570] 8. Terminal: Calculates the selected route and starts the guidance.
[0571] 9. Sensors: Perform continuous monitoring of the vehicle's condition.
[0572] 10. Server: If an anomaly is detected, it sends that information to the generation model to generate an appropriate countermeasure.
[0573] 11. Generative Model: Provide specific solutions tailored to the type of anomaly. Example: "The oil change light is on. Please follow these steps: (1) (2) (3)"
[0574] 12. Server: Sends the response data from the generated model to the car navigation terminal.
[0575] 13. Device: Notifies the user via voice and display and guides them through the instructed course of action.
[0576] 14. User: Follow the provided procedures to address any vehicle malfunctions.
[0577] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users and enables quick and specific responses in the event of vehicle malfunctions.
[0578] The following describes the processing flow.
[0579] Step 1:
[0580] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0581] Step 2:
[0582] Terminal: Activates the speech recognition engine to convert user voice commands from speech to text, and then sends it to the server.
[0583] Step 3:
[0584] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[0585] Step 4:
[0586] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0587] Step 5:
[0588] Server: Sends the generated response data to the car navigation terminal.
[0589] Step 6:
[0590] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[0591] Step 7:
[0592] User: Choose from the recommended options. Example: "I want to go to spot (1)."
[0593] Step 8:
[0594] Terminal: Calculates the selected route and starts navigation. Updates route information in real time.
[0595] Step 9:
[0596] Sensors: Monitor and collect various vehicle status data in real time.
[0597] Step 10:
[0598] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0599] Step 11:
[0600] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[0601] Step 12:
[0602] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[0603] Step 13:
[0604] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[0605] Step 14:
[0606] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[0607] (Example 1)
[0608] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0609] Conventional car navigation systems had problems with responding quickly and appropriately to user voice input and handling vehicle malfunctions. In particular, it was difficult to provide real-time responses to voice commands and to effectively coordinate vehicle status monitoring and anomaly detection. Furthermore, they lacked an intuitive interface that users could easily operate. As a result, the user experience was significantly reduced, and there was a risk of compromising driving safety.
[0610] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0611] In this invention, the server includes means for converting user voice into text data and transmitting it, means for passing the received text data to a generation model to generate a response, and means for transmitting the generated response to a car navigation terminal and providing it to the user in both voice and screen display. This enables real-time responses to user voice commands, vehicle status monitoring, and provision of appropriate countermeasures.
[0612] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[0613] A "car navigation terminal" is a device installed in a vehicle that takes voice commands as input, outputs generated responses, and displays route guidance.
[0614] A "speech recognition engine" refers to software or hardware that converts a user's voice into text data.
[0615] A "server" is a computing system on which the generative model operates and processes various types of data in conjunction with car navigation terminals and sensors.
[0616] A "sensor" is a device used to monitor the vehicle's condition in real time and collect data.
[0617] "Text data" refers to the character data generated by a speech recognition engine after analyzing speech.
[0618] "User interface" refers to the means, such as screens and sounds, that users use to operate a system.
[0619] "Abnormal data" refers to information about abnormal conditions in a vehicle detected by sensors.
[0620] "Handling procedures" refers to the steps and instructions for taking appropriate action in response to a vehicle malfunction.
[0621] "Response" refers to the content of the reply generated by the generative model based on the user's voice input.
[0622] This invention relates to a car navigation system that achieves advanced interactive functions using a generative model. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination. The configuration and specific operation of this system are described in detail below.
[0623] System Configuration
[0624] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[0625] Car navigation terminal
[0626] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and send it to a server.
[0627] Generative model
[0628] The server passes the text data received from the user to a generative model (e.g., OpenAI GPT-4) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0629] Specific example
[0630] When a user enters a voice command into the car navigation system, such as "I want to go for a drive to clear my head," the navigation system uses its voice recognition engine to convert the voice command into text data. This text data is then sent to a server via the internet.
[0631] The server passes the received text data to the generative model, which generates a response such as, "It's nighttime now, so we recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~". The generated response is then sent back to the car navigation terminal via the server.
[0632] The car navigation terminal uses a speech synthesis engine (e.g., Amazon Polly) and a screen display to provide responses to the user. The user selects "I want to go to spot (1)" from the recommended options, and the car navigation terminal calculates the selected route using a GPS navigation system (e.g., Google Maps API) and begins providing guidance via voice and screen.
[0633] Sensors and Anomaly Detection
[0634] Sensors installed inside the vehicle constantly monitor the vehicle's status. Vehicle status data is collected in real time and transmitted to a server. The server analyzes this data and, for example, if the oil lamp illuminates, sends that information to a generating model. The generating model then generates specific instructions for an oil change, such as "The oil change lamp is illuminated. Please follow these steps: (1) ~~ (2) ~~ (3) ~~," and transmits this information back to the car navigation terminal via the server.
[0635] The car navigation system notifies the user through voice and display, guiding them through the instructed course of action. The user then follows the suggested steps, for example, heading to the nearest gas station for an oil change.
[0636] Through the configuration and specific operation described above, this system realizes the interactive navigation experience that users desire and enables quick and specific responses in the event of a vehicle malfunction.
[0637] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0638] Step 1:
[0639] The user enters a voice command.
[0640] The specific action involves the user issuing a voice command to the car navigation terminal, such as "I want to go for a drive to clear my head." At this stage, the input is the user's voice.
[0641] Step 2:
[0642] The terminal converts voice commands into text data.
[0643] A speech recognition engine (e.g., Google Cloud Speech-to-Text) works to analyze the user's voice commands and generate text data. The input is the user's voice data, and the output is the corresponding text data.
[0644] Step 3:
[0645] The terminal sends text data to the server.
[0646] Text data converted from speech is sent from the terminal to the server via the network. The input is the text data generated by the speech recognition engine, and the output is the completion of the transmission of the text data to the server.
[0647] Step 4:
[0648] The server passes text data to the generative model.
[0649] The server provides the received text data to a generative AI model (e.g., OpenAI GPT-4). The generative model analyzes the text data and understands the user's intent. The input is the text data received by the server, and the output is the analysis request to the generative model.
[0650] Step 5:
[0651] The generative model generates the response.
[0652] The generative model generates an appropriate response based on the analysis results. For example, it might generate a response like, "Since it's nighttime, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~". The input is text data passed from the server, and the output is the generated response.
[0653] Step 6:
[0654] The server sends the generative model response to the terminal.
[0655] The server sends the response message obtained from the generative model to the car navigation terminal. The input is the response message generated by the generative model, and the output is the completion of sending that response message to the terminal.
[0656] Step 7:
[0657] The device provides responses to the user through voice and on-screen displays.
[0658] The device uses a speech synthesis engine (e.g., Amazon Polly) and a display to provide the user with generated responses as audio and visual information. The input is the response text sent from the server, and the output is the audio and screen display response to the user.
[0659] Step 8:
[0660] The user selects from the recommended options.
[0661] The user makes selections using voice or a touch panel, for example, "I want to go to spot (1)." Input is provided by the terminal display and voice guidance, while output is the user's selection.
[0662] Step 9:
[0663] The device calculates the selected route and begins providing directions.
[0664] The device uses a GPS navigation system (e.g., Google Maps API) to calculate the selected route. The input is the user's selection, and the output is the selected route and the start of navigation.
[0665] Step 10:
[0666] Sensors monitor the vehicle's status.
[0667] Sensors installed inside the vehicle monitor engine temperature, tire pressure, and other parameters in real time. Inputs are various vehicle status data, and outputs are data collected by the sensors.
[0668] Step 11:
[0669] If the server detects an anomaly, it sends information to the generative model.
[0670] The server analyzes data from the sensors and, if an anomaly is detected, sends that information to the generative model. The input is state data from the sensors, and the output is the transmission of anomaly data to the generative model.
[0671] Step 12:
[0672] The generative model generates methods for dealing with anomalies.
[0673] The generative model generates specific corrective actions based on the type of anomaly. The input is the anomaly data, and the output is a response message describing the corrective action. For example, "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[0674] Step 13:
[0675] The server sends the generative model response to the terminal.
[0676] The server sends the solution obtained from the generative model to the car navigation terminal. The input is the response message of the solution generated by the generative model, and the output is the completion of sending that response message to the terminal.
[0677] Step 14:
[0678] The device notifies the user and provides instructions.
[0679] The terminal uses voice and screen display to notify and guide the user of the solutions derived from the generative model. Input is the response text of the solution sent from the server, and output is voice and display guidance to the user.
[0680] Step 15:
[0681] The user will take action by following the provided instructions.
[0682] The user follows the instructions provided, for example, by going to the nearest gas station for an oil change. Input is the guidance from the terminal, and output is the action taken to address the vehicle problem.
[0683] (Application Example 1)
[0684] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0685] Conventional car navigation systems can only provide static responses to user voice commands, making real-time responses to vehicle abnormalities difficult. Furthermore, there has been a lack of systems capable of sophisticated user interaction in the operation of autonomous vehicles. Solving these problems is essential.
[0686] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0687] In this invention, the server includes means for interpreting voice input from a user using a generative model and generating an appropriate response; means for dynamically updating information displayed on a car navigation terminal based on the user's voice instructions; sensors for collecting vehicle status data in real time and detecting anomalies; means for transmitting anomaly data collected from the sensors to the generative model and providing appropriate countermeasures; means for calculating a route to a destination through navigation using the generative model and providing guidance to the autonomous vehicle; and means for monitoring the vehicle status in real time and guiding the user to appropriate countermeasures via the generative model when an anomaly occurs. This enables the provision of dynamic and appropriate responses to the user's voice instructions, and allows for navigation of the autonomous vehicle and rapid response to anomalies.
[0688] A "generative model" is a machine learning model that interprets voice input from a user and generates an appropriate response.
[0689] A "car navigation terminal" is a device installed in a vehicle that displays information based on the user's voice commands.
[0690] "Means" refer to methods or devices used to achieve a specific function or role.
[0691] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[0692] An "autonomous vehicle" is a vehicle that can drive autonomously without the intervention of a human driver.
[0693] "Guidance" refers to information and instructions that provide route information to a destination and guide a vehicle to its destination.
[0694] A "user interface" refers to equipment such as screens and audio output devices that allow users to operate a system and obtain information.
[0695] "Real-time" refers to actions or processes that occur immediately or with a very short delay.
[0696] This invention is a system that uses a generative model to realize advanced interactive functions. The system consists of a car navigation terminal, a server on which the generative model operates, vehicle sensors, and various interfaces.
[0697] System Configuration
[0698] Car navigation terminal
[0699] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[0700] Generative model
[0701] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0702] Sensors and Anomaly Detection
[0703] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0704] User Interface
[0705] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[0706] Operation overview
[0707] 1. Receiving and analyzing voice commands:
[0708] The car navigation system uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice commands into text data.
[0709] This text data will be sent to the server.
[0710] 2. Response generation using generative models:
[0711] The server passes the transformed text data to a generation model (e.g., the text generation model in the transformers library) to generate an appropriate response.
[0712] The generated response is sent back to the car navigation terminal via the server and provided to the user via voice and screen display.
[0713] 3. Route calculation and navigation:
[0714] The car navigation terminal calculates the optimal route to the destination and performs navigation based on the response of the generative model.
[0715] 4. Detection and response to vehicle abnormalities:
[0716] The sensors collect vehicle status data in real time and transmit it to the server.
[0717] When an anomaly is detected, the server passes that information to the generative model, which then generates an appropriate response.
[0718] The generative model provides specific solutions tailored to the type of anomaly and notifies the user via voice and display.
[0719] Specific example
[0720] For example, when a user enters a voice command on a smartphone, the following prompt message is used:
[0721] "Hey GPS, tell me some good restaurants around here."
[0722] In response to this, the system provides the following answer:
[0723] "We have two recommended restaurants nearby: (1) Restaurant A and (2) Cafe B. Which one would you like to go to?"
[0724] Once the user makes a selection, the system calculates the route to the specified destination and provides directions to the autonomous vehicle.
[0725] In this way, we can realize the interactive navigation experience that users desire and enable quick and specific responses in the event of a vehicle malfunction.
[0726] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0727] Step 1:
[0728] Receiving voice commands
[0729] Operation: The user enters voice commands into the car navigation terminal.
[0730] Input: User voice command (e.g., "Tell me some good restaurants nearby").
[0731] Output: Audio data.
[0732] Specific operation: The car navigation terminal collects the user's voice through the microphone.
[0733] Step 2:
[0734] Converting audio data to text
[0735] Operation: The device uses a speech recognition engine to convert speech data into text data.
[0736] Input: Audio data.
[0737] Output: Text data (e.g., "Tell me some good restaurants nearby").
[0738] Specific operation: Using a speech recognition library (e.g., speech_recognition), the system analyzes audio data and converts it into text data.
[0739] Step 3:
[0740] Sending text data to the server
[0741] Operation: The terminal sends the converted text data to the server.
[0742] Input: Text data.
[0743] Output: Notification that the text data has been successfully sent to the server.
[0744] Specific operation: Use the communication module in the terminal to send text data to the server.
[0745] Step 4:
[0746] Response generation using generative models
[0747] Operation: The server inputs the received text data into a generative model and generates an appropriate response.
[0748] Input: Text data.
[0749] Output: Generated response (e.g., "Recommended nearby restaurants are (1) Restaurant A and (2) Cafe B.").
[0750] Specific operation: The server inputs text data into a generative model (e.g., the text generation model in the transformers library), analyzes the user's intent, and generates an appropriate response.
[0751] Step 5:
[0752] Sending response data to the terminal
[0753] Operation: The server sends the generated response data to the terminal.
[0754] Input: Generated response data.
[0755] Output: Notification that the response data has been successfully sent to the terminal.
[0756] Specific operation: The generated response data is sent to the car navigation terminal using the communication module within the server.
[0757] Step 6:
[0758] Providing responses to users
[0759] Operation: The terminal provides the user with received response data via voice and screen display.
[0760] Input: Response data.
[0761] Output: Audio output and screen display (e.g., "Recommended nearby restaurants include (1) Restaurant A and (2) Cafe B.").
[0762] Specific operation: The device will reply using a speech synthesis engine (e.g., tts_engine) and simultaneously display it on the screen.
[0763] Step 7:
[0764] User Selection
[0765] Operation: The user makes a selection regarding the response via voice or touch.
[0766] Input: User selection (e.g., "I want to go to restaurant A (1)").
[0767] Output: Selected data.
[0768] Specific operation: The device collects user selections via voice recognition or screen touch.
[0769] Step 8:
[0770] Route calculation and navigation start
[0771] Operation: The device calculates the route to the selected destination and begins navigation.
[0772] Input: Selected data (e.g., "Restaurant A").
[0773] Output: Route data and navigation guide.
[0774] Specific operation: The device uses a navigation library (e.g., navigation) to calculate the optimal route to the destination and starts navigation.
[0775] Step 9:
[0776] Real-time monitoring of vehicle status
[0777] Operation: Sensors collect vehicle status data in real time and send it to the server.
[0778] Input: Vehicle status data.
[0779] Output: Notification that vehicle status data has been successfully sent to the server.
[0780] Specific operation: Sensors inside the vehicle collect status data in real time and send it to the server.
[0781] Step 10:
[0782] Vehicle abnormality detection
[0783] Operation: The server analyzes the received vehicle status data and detects any abnormalities.
[0784] Input: Vehicle status data.
[0785] Output: Anomaly detection notification (e.g., "Oil change required").
[0786] Specific operation: The analysis engine within the server analyzes the status data and detects anomalies.
[0787] Step 11:
[0788] Generating solutions using generative models
[0789] Operation: The server inputs anomaly detection data into a generation model and generates appropriate countermeasures.
[0790] Input: Anomaly detection data.
[0791] Output: Troubleshooting steps (e.g., "Oil change is needed. Please follow these steps.").
[0792] Specific operation: The server inputs anomaly detection data into the generative model and generates specific countermeasures for the user.
[0793] Step 12:
[0794] Sending the troubleshooting steps to the device
[0795] Operation: The server sends the generated troubleshooting data to the terminal.
[0796] Input: Data on how to handle the situation.
[0797] Output: Notification that the data regarding the solution has been successfully sent to the terminal.
[0798] Specific operation: The generated troubleshooting data is sent to the car navigation terminal using a communication module on the server.
[0799] Step 13:
[0800] Providing users with solutions
[0801] Operation: The device provides the user with received troubleshooting data via voice and screen display.
[0802] Input: Data on how to handle the situation.
[0803] Output: Audio output and screen display (e.g., "Oil change is needed. Please follow the next steps.").
[0804] Specific operation: The device uses a speech synthesis engine to notify the user of the solution via voice, and simultaneously displays it on the screen.
[0805] Step 14:
[0806] User Action
[0807] Operation: The system will address the vehicle malfunction according to the troubleshooting steps provided to the user.
[0808] Input: Display of troubleshooting methods and voice guidance.
[0809] Output: Normal vehicle condition.
[0810] Specific actions: The user follows the provided instructions and performs the actual troubleshooting steps.
[0811] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0812] This invention combines an emotion engine with a car navigation system that interprets voice input from the user using a generative model and generates an appropriate response. The system consists of a car navigation terminal, a server, a generative model, vehicle sensors, and an emotion engine that recognizes the user's emotions. This improves the user's driving experience and enhances safety.
[0813] System configuration and operation
[0814] This system consists of a car navigation terminal, a server on which the generative model runs, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[0815] Car navigation terminal
[0816] A car navigation terminal is installed in a vehicle and handles voice command input, output of responses using a generative model and emotion engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[0817] Generative model
[0818] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[0819] Sensors and Anomaly Detection
[0820] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0821] Emotional Engine
[0822] The emotion engine recognizes the user's emotions from their voice input and the in-car environment. The emotions recognized by the emotion engine are then used to customize the generative model's responses and navigation information. In particular, if signs of stress or fatigue are detected, it includes features that recommend relaxation spots and rest areas.
[0823] User Interface
[0824] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. Settings include language, voice guidance type, and theme color, which can be customized to the user's preferences.
[0825] Explanation of the program's processing
[0826] Here, the system's operation is explained in natural language, describing the program's processing.
[0827] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0828] 2. Terminal: Activates the speech recognition engine to convert the user's voice commands from speech to text. This text data is then sent to the server.
[0829] 3. Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[0830] 4. Generative Model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0831] 5. Server: Sends the generated response data to the car navigation terminal.
[0832] 6. Terminal: Provides the user with the received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[0833] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[0834] 8. Terminal: Calculates the selected route and starts guidance. Updates route information in real time.
[0835] 9. Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[0836] 10. Server: Inputs data from the emotion engine into the generative model to generate a customized response based on emotion. Example: "Would you recommend a relaxation spot?"
[0837] 11. Sensors: Monitor and collect various vehicle status data in real time.
[0838] 12. Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0839] 13. Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[0840] 14. Server: Sends response data, including the solution created by the generative model, to the car navigation terminal.
[0841] 15. Device: Notify the user of how to resolve the issue via voice and on-screen display. Provide specific instructions.
[0842] 16. User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the user will be guided to the nearest gas station or repair shop.
[0843] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users, and further enables quick and specific responses in the event of vehicle malfunctions, thereby enhancing safety and comfort. By combining it with an emotional engine, flexible responses tailored to the user's psychological state become possible, further improving the driving experience.
[0844] The following describes the processing flow.
[0845] Step 1:
[0846] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[0847] Step 2:
[0848] Terminal: Activates the speech recognition engine and converts the user's voice commands into text data. This text data is then sent to the server.
[0849] Step 3:
[0850] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[0851] Step 4:
[0852] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[0853] Step 5:
[0854] Server: Sends the generated response data to the car navigation terminal.
[0855] Step 6:
[0856] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[0857] Step 7:
[0858] User: Select from the suggested options. Example: "I want to go to spot (1)."
[0859] Step 8:
[0860] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[0861] Step 9:
[0862] Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[0863] Step 10:
[0864] Server: Passes emotional data from the emotion engine to the generative model, which then generates an emotionally appropriate response. Example: "Would you recommend a relaxation spot?"
[0865] Step 11:
[0866] Generative model: Generates customized responses based on emotions. Example: "You seem stressed, so I recommend the following relaxation spots: (1)~~ (2)~~ (3)~~"
[0867] Step 12:
[0868] Server: Sends the generated response data to the car navigation terminal.
[0869] Step 13:
[0870] Terminal: Provides the user with received response data via voice output and screen display. Example: "We recommend the following relaxation spots. (1)~~ (2)~~ (3)~~ Where would you like us to take you?"
[0871] Step 14:
[0872] User: Select from the suggested relaxation spots. Example: "I want to go to relaxation spot (2)."
[0873] Step 15:
[0874] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[0875] Step 16:
[0876] Sensors: Monitor and collect various vehicle status data in real time.
[0877] Step 17:
[0878] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0879] Step 18:
[0880] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[0881] Step 19:
[0882] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[0883] Step 20:
[0884] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[0885] Step 21:
[0886] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[0887] This process allows the system to provide an interactive navigation experience while taking the user's emotional state into consideration, and enables quick and specific action in the event of a vehicle malfunction.
[0888] (Example 2)
[0889] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0890] Conventional car navigation systems simply interpret user voice input to provide route guidance, making it difficult to respond in accordance with the user's emotions or the vehicle's malfunction. Furthermore, the lack of customized responses tailored to the user's psychological state meant that stress and fatigue during driving could not be reduced. Additionally, there was a problem with the inability to quickly provide specific solutions in the event of a vehicle malfunction.
[0891] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, means for dynamically updating information displayed on the navigation terminal based on the user's voice instructions, sensors for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensors to the generation model and providing an appropriate countermeasure, an emotion engine for recognizing the user's emotions from the user's voice input and the situation inside the vehicle and reflecting this in the response of the generation model, and means for customizing navigation information based on the user's emotions. This enables flexible responses according to the user's psychological state and quick and specific countermeasures in the event of a vehicle abnormality.
[0892] A "generative model" is a type of artificial intelligence that analyzes input data and generates appropriate responses or predictions based on that data.
[0893] A "navigation terminal" is a device installed inside a vehicle that provides map information and route guidance.
[0894] "Voice instructions" refer to instructions or commands spoken by the user, which serve as input for the navigation system to recognize and process.
[0895] A "sensor" is a device that collects vehicle status data in real time and detects specific conditions or abnormalities.
[0896] An "emotion engine" is software or an algorithm that analyzes the user's emotions from their voice input and the situation inside the vehicle, and reflects that emotional information in the response of a generative model.
[0897] "Handling instructions" refer to specific actions and procedures that the user should take when they detect a problem or abnormality in the vehicle.
[0898] "Customization" refers to modifying the interface and responses based on the user's settings and status to meet individual needs.
[0899] Modes for carrying out the invention
[0900] The car navigation system of the present invention enhances the user's driving experience and improves safety by combining a generative model, a voice recognition engine, an emotion engine, and various sensors. The present invention consists of a navigation terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[0901] Navigation terminal
[0902] The navigation terminal is installed in the vehicle, takes user voice commands as input, outputs responses using a generative model and emotion engine, and displays route guidance. When the user inputs a voice command such as "I want to go for a drive to clear my head," the terminal uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the voice into text data and sends it to the server.
[0903] Generative model
[0904] The server passes the text data received from the user to a generative model (e.g., GPT-3) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is sent to the navigation terminal via the server.
[0905] Sensors and Anomaly Detection
[0906] The vehicle is equipped with various sensors, including tire pressure sensors and an engine diagnostic system, to collect vehicle status data in real time. This data is sent to a server, which analyzes it and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[0907] Emotional Engine
[0908] The emotion engine recognizes the user's emotions from their voice input and the environment inside the vehicle. For example, if the user is feeling stressed, the emotion engine detects this and sends the data to the generative model. The generative model then generates a customized response based on the emotion data, such as recommending a relaxation spot to the user.
[0909] User Interface
[0910] The user interface consists of a screen and audio for the navigation terminal and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, which can be customized to the user's preferences.
[0911] Specific example
[0912] When a user inputs a voice command such as "I want to go for a drive to clear my head," the voice recognition engine converts the voice into text and sends it to the server. The server passes this text data to a generative model, which generates a response such as "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~." The generated response is sent to the navigation terminal, where the user confirms the recommended spots on the screen and by voice and selects a destination. The navigation terminal then calculates the selected route and begins guidance.
[0913] Thus, by utilizing a generative AI model, a speech recognition engine, an emotion engine, and sensors, the present invention enables flexible responses that respond to the user's intentions and emotions, realizing a car navigation system that achieves both safety and comfort.
[0914] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0915] Program processing flow
[0916] Step 1:
[0917] User: The user enters a voice command. Specifically, they say, "I want to go for a drive to clear my head."
[0918] Input: User's voice
[0919] Output: Raw data audio file
[0920] Step 2:
[0921] Device: Uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert user voice commands from speech to text.
[0922] Input: Raw audio file
[0923] Output: Text data (Example: "I want to go for a drive to clear my head")
[0924] Step 3:
[0925] Terminal: Sends text data to the server.
[0926] Input: Text data
[0927] Output: Sending text data to the server
[0928] Step 4:
[0929] Server: Passes the received text data to a generative model (e.g., GPT-3) and starts the analysis.
[0930] Input: Text data
[0931] Output: User intent (e.g., "Recommendations for destinations suitable for a change of pace")
[0932] Step 5:
[0933] Generative model: Generates appropriate responses based on the user's intent. For example, it might create a response like, "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~".
[0934] Input: User intent
[0935] Output: Response text (Example: "Since it's nighttime now, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~")
[0936] Step 6:
[0937] Server: Sends the generated response text data to the navigation terminal.
[0938] Input: Response text data
[0939] Output: Sending response data to the navigation terminal
[0940] Step 7:
[0941] Terminal: Provides the user with the received response text via voice output (e.g., Amazon Polly) and screen display. Specifically, it would say, "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[0942] Input: Response text data
[0943] Output: Voice guidance and screen display
[0944] Step 8:
[0945] User: The user selects from the provided options and responds, "I want to go to spot (1)."
[0946] Input: User Selection
[0947] Output: Selection (Example: "(1) Spot")
[0948] Step 9:
[0949] Terminal: Calculates the selected route using the car navigation system's route calculation engine (e.g., Google Maps API) and begins guidance. Route information is updated in real time.
[0950] Input: Selection
[0951] Output: Route guidance and real-time updates
[0952] Step 10:
[0953] Emotion Engine: Analyzes user voice input and in-car environment to recognize user emotions. Specifically, it determines whether the user is experiencing stress.
[0954] Input: User voice, in-vehicle status data
[0955] Output: Recognized emotion (e.g., "I am feeling stressed")
[0956] Step 11:
[0957] Server: Inputs data from the emotion engine into the generative model to generate customized responses based on the user's emotions. For example, it might create a response such as, "Would you recommend a relaxation spot?"
[0958] Input: Sentiment data
[0959] Output: Response text data (e.g., "Do you recommend any relaxation spots?")
[0960] Step 12:
[0961] Sensors: Collect various status data (tire pressure, engine diagnostics, etc.) in real time.
[0962] Input: Vehicle condition
[0963] Output: Status data
[0964] Step 13:
[0965] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[0966] Input: Status data from the sensor
[0967] Output: Abnormal data and notifications
[0968] Step 14:
[0969] Generative Model: Based on abnormal data, it generates appropriate countermeasures and creates instructions such as, "The oil change light is on. Please take the following steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[0970] Input: Abnormal data
[0971] Output: Text data of the solution.
[0972] Step 15:
[0973] Server: Sends response data containing the solution generated by the generative model to the navigation terminal.
[0974] Input: Text data of the solution
[0975] Output: Send to navigation terminal
[0976] Step 16:
[0977] Terminal: Notifies the user of how to resolve the issue via voice and on-screen display, and guides them through specific steps. For example, it might say, "The oil change light is on. Please follow these steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[0978] Input: Text data of the solution
[0979] Output: Voice guidance and screen display
[0980] Step 17:
[0981] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[0982] Input: Instructions on how to proceed
[0983] Output: Appropriate countermeasures
[0984] (Application Example 2)
[0985] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0986] In recent years, the demand for food delivery services has increased, requiring drivers to perform delivery duties efficiently and safely. However, excessive stress and fatigue during driving can reduce drivers' attention span and increase the risk of accidents. Furthermore, current car navigation systems cannot provide responses that take into account the driver's emotional state, and improvements are needed to enhance the driver's driving experience. Therefore, a system is needed that analyzes the driver's emotional state in real time and generates appropriate responses.
[0987] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, an in-vehicle information terminal, means for dynamically updating the information displayed on the in-vehicle information terminal based on the user's voice instructions, a sensor for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensor to the generation model and providing an appropriate countermeasure, emotion analysis means for analyzing the user's emotional state and generating a response corresponding to that emotion, and means for suggesting rest locations such as relaxation spots according to the user's emotional state recognized by the emotion analysis means. This makes it possible to detect the stress and fatigue of the driver and provide appropriate responses and rest suggestions, thereby enabling safe and efficient delivery operations.
[0988] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[0989] An "in-vehicle information terminal" is a device installed in a vehicle that allows for the input of voice commands, the output of responses from a generative model and sentiment analysis engine, and the display of route guidance.
[0990] A "dynamically updating method" refers to a method of changing the information displayed on the in-vehicle information terminal in real time based on the user's voice commands.
[0991] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[0992] "Means of providing appropriate countermeasures" refers to a method of transmitting anomaly data collected from sensors to a generation model and then presenting specific countermeasures based on that data.
[0993] "Emotion analysis means" refers to a method that recognizes the emotional state from the user's voice input or the situation inside the vehicle and provides that information to a generative model.
[0994] A "relaxation spot" is a place where users can rest and refresh themselves when they feel stressed or tired.
[0995] This invention is a system that supports food delivery drivers in performing their duties safely and efficiently, and consists of the following components.
[0996] System configuration
[0997] This system consists of an in-vehicle information terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion analysis engine.
[0998] In-vehicle information terminal
[0999] The in-vehicle information terminal is installed in the vehicle and handles voice command input, output of responses generated by a generative model and sentiment analysis engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to the server. For example, if a driver voice-inputs "I'm tired, I want to take a short break," the terminal converts that voice into text and sends it to the server.
[1000] Generative model
[1001] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and emotional state to generate a specific and appropriate response. The generated response is returned to the in-vehicle information terminal via the server and provided to the user. For example, in response to input such as "I'm tired, I want to take a break," a response like "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?" is generated.
[1002] Sensors and Anomaly Detection
[1003] Sensors within the vehicle collect vehicle status data in real time and send it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps. For example, if it detects that an oil change is needed, it will provide countermeasures such as, "The oil change lamp is lit. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[1004] Emotion analysis engine
[1005] The emotion analysis engine recognizes the user's emotional state from their voice input and the environment inside the vehicle. The user's emotional state recognized by the emotion analysis engine is reflected in the generative model's responses and the customization of navigation information. In particular, if the system detects that the user is feeling "tired" or "stressed," it includes a function that suggests relaxation spots and rest areas. For example, if the user inputs "tired," the system detects fatigue and suggests nearby relaxation spots.
[1006] User Interface
[1007] The user interface consists of the in-car information terminal's screen and voice, and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, all of which can be customized to the user's preferences. For example, simply saying "I want to change the voice guidance voice" will change it to the preferred voice type.
[1008] Specific example
[1009] 1. User input: "I'm tired, I want to take a break."
[1010] 2. Example of a prompt:
[1011] Emotion: Fatigue, Input: I'm tired and want to take a break. Can you suggest any relaxation spots?
[1012] 3. Response generated: "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[1013] This configuration allows for the detection of driver stress and fatigue, and enables safe and efficient delivery operations by providing appropriate responses and rest suggestions.
[1014] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1015] Step 1:
[1016] The user inputs a voice command. For example, they might say, "I'm tired, I want to take a break." This voice input becomes the starting point for the system's processing.
[1017] Step 2:
[1018] The device activates a speech recognition engine to convert the user's voice commands from speech to text. The input is the user's voice data, and the output is the text data of that voice. The device sends this text data to the server.
[1019] Step 3:
[1020] The server passes the received text data to the generative model. The input is text data sent from the terminal, and the generative model is used to analyze the user's intent and emotions. The output is data containing the analysis results.
[1021] Step 4:
[1022] The generative model generates appropriate responses based on the user's intentions and emotions. The input is the parsed data obtained in the previous step, and the output is the response sentence presented to the user. For example, it might generate a response such as, "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[1023] Step 5:
[1024] The server transmits the generated response data to the in-vehicle infotainment terminal. The input is the response data from the generative model, and the output is the communication data to the in-vehicle infotainment terminal.
[1025] Step 6:
[1026] The terminal provides the user with the received response data through voice output and screen display. The input is the response data from the server, and the output is voice guidance and screen display to the user. For example, it might provide voice guidance such as, "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[1027] Step 7:
[1028] The user selects an option from the suggested relaxation spots and gives a voice command. For example, they might respond, "I want to go to Cafe B." This voice data becomes the input for the next process.
[1029] Step 8:
[1030] The device restarts its speech recognition engine and converts the user's selection from speech to text. The input is the user's voice data, and the output is the text data of the selections. The device sends this text data to the server.
[1031] Step 9:
[1032] The server calculates the route to the selected relaxation spot based on the received text data. The input is the text data of the selected spot, and the output is the data of the optimal route.
[1033] Step 10:
[1034] The server transmits the calculated route to the in-vehicle information terminal. The input is the route calculation result data, and the output is the communication data to the in-vehicle information terminal.
[1035] Step 11:
[1036] The terminal starts navigation based on the received route data and updates the route information in real time. The input is route data from the server, and the output is navigation guidance and screen display.
[1037] Step 12:
[1038] The emotion analysis engine analyzes the user's voice input and the conditions inside the vehicle to recognize the user's emotions. The input is the user's voice data and sensor information, and the output is emotional state data.
[1039] Step 13:
[1040] The server inputs data obtained from the emotion analysis engine into a generative model to generate a customized response based on the emotion. The input is emotional state data, and the output is customized response data. For example, it can generate a response such as, "Would you recommend a relaxation spot?"
[1041] Step 14:
[1042] The sensors monitor and collect various vehicle status data in real time. The input is vehicle status information, and the output is data signals from the sensors.
[1043] Step 15:
[1044] The server continuously receives data from the sensor and, if it detects an anomaly, sends that information to the generative model. The input is the data signal from the sensor, and the output is the anomaly data sent to the generative model.
[1045] Step 16:
[1046] The generative model generates appropriate countermeasures based on abnormal data and creates specific procedures. The input is abnormal data, and the output is procedure data for the countermeasures. For example, it generates procedures such as "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[1047] Step 17:
[1048] The server sends response data, including the countermeasures created by the generative model, to the in-vehicle information terminal. The input is the procedural data of the countermeasures, and the output is the communication data to the in-vehicle information terminal.
[1049] Step 18:
[1050] The device notifies the user of troubleshooting methods via voice and on-screen display, guiding them through specific steps. Input is troubleshooting data from the server, and output is voice guidance and on-screen display to the user.
[1051] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1052] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1053] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1054] [Third Embodiment]
[1055] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1056] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1057] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1058] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1059] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1060] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1061] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1062] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1063] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1064] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1065] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1066] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1067] This invention relates to a car navigation system that uses a generative model to achieve advanced interactive functions. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination.
[1068] System Configuration
[1069] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[1070] Car navigation terminal
[1071] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[1072] Generative model
[1073] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1074] Sensors and Anomaly Detection
[1075] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1076] User Interface
[1077] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[1078] Explanation of the program's processing
[1079] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1080] 2. Terminal: Activates the speech recognition engine to convert voice commands into text data. Then sends that text data to the server.
[1081] 3. Server: Passes text data to the generative model and generates an appropriate response.
[1082] 4. Generative Model: Analyzes user intent and generates appropriate responses. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1083] 5. Server: Sends the response data from the generated model to the car navigation terminal.
[1084] 6. Device: Provides responses to the user via voice and screen display.
[1085] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[1086] 8. Terminal: Calculates the selected route and starts the guidance.
[1087] 9. Sensors: Perform continuous monitoring of the vehicle's condition.
[1088] 10. Server: If an anomaly is detected, it sends that information to the generation model to generate an appropriate countermeasure.
[1089] 11. Generative Model: Provide specific solutions tailored to the type of anomaly. Example: "The oil change light is on. Please follow these steps: (1) (2) (3)"
[1090] 12. Server: Sends the response data from the generated model to the car navigation terminal.
[1091] 13. Device: Notifies the user via voice and display and guides them through the instructed course of action.
[1092] 14. User: Follow the provided procedures to address any vehicle malfunctions.
[1093] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users and enables quick and specific responses in the event of vehicle malfunctions.
[1094] The following describes the processing flow.
[1095] Step 1:
[1096] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1097] Step 2:
[1098] Terminal: Activates the speech recognition engine to convert user voice commands from speech to text, and then sends it to the server.
[1099] Step 3:
[1100] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[1101] Step 4:
[1102] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1103] Step 5:
[1104] Server: Sends the generated response data to the car navigation terminal.
[1105] Step 6:
[1106] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[1107] Step 7:
[1108] User: Choose from the recommended options. Example: "I want to go to spot (1)."
[1109] Step 8:
[1110] Terminal: Calculates the selected route and starts navigation. Updates route information in real time.
[1111] Step 9:
[1112] Sensors: Monitor and collect various vehicle status data in real time.
[1113] Step 10:
[1114] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1115] Step 11:
[1116] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[1117] Step 12:
[1118] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[1119] Step 13:
[1120] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[1121] Step 14:
[1122] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[1123] (Example 1)
[1124] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1125] Conventional car navigation systems had problems with responding quickly and appropriately to user voice input and handling vehicle malfunctions. In particular, it was difficult to provide real-time responses to voice commands and to effectively coordinate vehicle status monitoring and anomaly detection. Furthermore, they lacked an intuitive interface that users could easily operate. As a result, the user experience was significantly reduced, and there was a risk of compromising driving safety.
[1126] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1127] In this invention, the server includes means for converting user voice into text data and transmitting it, means for passing the received text data to a generation model to generate a response, and means for transmitting the generated response to a car navigation terminal and providing it to the user in both voice and screen display. This enables real-time responses to user voice commands, vehicle status monitoring, and provision of appropriate countermeasures.
[1128] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[1129] A "car navigation terminal" is a device installed in a vehicle that takes voice commands as input, outputs generated responses, and displays route guidance.
[1130] A "speech recognition engine" refers to software or hardware that converts a user's voice into text data.
[1131] A "server" is a computing system on which the generative model operates and processes various types of data in conjunction with car navigation terminals and sensors.
[1132] A "sensor" is a device used to monitor the vehicle's condition in real time and collect data.
[1133] "Text data" refers to the character data generated by a speech recognition engine after analyzing speech.
[1134] "User interface" refers to the means, such as screens and sounds, that users use to operate a system.
[1135] "Abnormal data" refers to information about abnormal conditions in a vehicle detected by sensors.
[1136] "Handling procedures" refers to the steps and instructions for taking appropriate action in response to a vehicle malfunction.
[1137] "Response" refers to the content of the reply generated by the generative model based on the user's voice input.
[1138] This invention relates to a car navigation system that achieves advanced interactive functions using a generative model. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination. The configuration and specific operation of this system are described in detail below.
[1139] System Configuration
[1140] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[1141] Car navigation terminal
[1142] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and send it to a server.
[1143] Generative model
[1144] The server passes the text data received from the user to a generative model (e.g., OpenAI GPT-4) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1145] Specific example
[1146] When a user enters a voice command into the car navigation system, such as "I want to go for a drive to clear my head," the navigation system uses its voice recognition engine to convert the voice command into text data. This text data is then sent to a server via the internet.
[1147] The server passes the received text data to the generative model, which generates a response such as, "It's nighttime now, so we recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~". The generated response is then sent back to the car navigation terminal via the server.
[1148] The car navigation terminal uses a speech synthesis engine (e.g., Amazon Polly) and a screen display to provide responses to the user. The user selects "I want to go to spot (1)" from the recommended options, and the car navigation terminal calculates the selected route using a GPS navigation system (e.g., Google Maps API) and begins providing guidance via voice and screen.
[1149] Sensors and Anomaly Detection
[1150] Sensors installed inside the vehicle constantly monitor the vehicle's status. Vehicle status data is collected in real time and transmitted to a server. The server analyzes this data and, for example, if the oil lamp illuminates, sends that information to a generating model. The generating model then generates specific instructions for an oil change, such as "The oil change lamp is illuminated. Please follow these steps: (1) ~~ (2) ~~ (3) ~~," and transmits this information back to the car navigation terminal via the server.
[1151] The car navigation system notifies the user through voice and display, guiding them through the instructed course of action. The user then follows the suggested steps, for example, heading to the nearest gas station for an oil change.
[1152] Through the configuration and specific operation described above, this system realizes the interactive navigation experience that users desire and enables quick and specific responses in the event of a vehicle malfunction.
[1153] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1154] Step 1:
[1155] The user enters a voice command.
[1156] The specific action involves the user issuing a voice command to the car navigation terminal, such as "I want to go for a drive to clear my head." At this stage, the input is the user's voice.
[1157] Step 2:
[1158] The terminal converts voice commands into text data.
[1159] A speech recognition engine (e.g., Google Cloud Speech-to-Text) works to analyze the user's voice commands and generate text data. The input is the user's voice data, and the output is the corresponding text data.
[1160] Step 3:
[1161] The terminal sends text data to the server.
[1162] Text data converted from speech is sent from the terminal to the server via the network. The input is the text data generated by the speech recognition engine, and the output is the completion of the transmission of the text data to the server.
[1163] Step 4:
[1164] The server passes text data to the generative model.
[1165] The server provides the received text data to a generative AI model (e.g., OpenAI GPT-4). The generative model analyzes the text data and understands the user's intent. The input is the text data received by the server, and the output is the analysis request to the generative model.
[1166] Step 5:
[1167] The generative model generates the response.
[1168] The generative model generates an appropriate response based on the analysis results. For example, it might generate a response like, "Since it's nighttime, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~". The input is text data passed from the server, and the output is the generated response.
[1169] Step 6:
[1170] The server sends the generative model response to the terminal.
[1171] The server sends the response message obtained from the generative model to the car navigation terminal. The input is the response message generated by the generative model, and the output is the completion of sending that response message to the terminal.
[1172] Step 7:
[1173] The device provides responses to the user through voice and on-screen displays.
[1174] The device uses a speech synthesis engine (e.g., Amazon Polly) and a display to provide the user with generated responses as audio and visual information. The input is the response text sent from the server, and the output is the audio and screen display response to the user.
[1175] Step 8:
[1176] The user selects from the recommended options.
[1177] The user makes selections using voice or a touch panel, for example, "I want to go to spot (1)." Input is provided by the terminal display and voice guidance, while output is the user's selection.
[1178] Step 9:
[1179] The device calculates the selected route and begins providing directions.
[1180] The device uses a GPS navigation system (e.g., Google Maps API) to calculate the selected route. The input is the user's selection, and the output is the selected route and the start of navigation.
[1181] Step 10:
[1182] Sensors monitor the vehicle's status.
[1183] Sensors installed inside the vehicle monitor engine temperature, tire pressure, and other parameters in real time. Inputs are various vehicle status data, and outputs are data collected by the sensors.
[1184] Step 11:
[1185] If the server detects an anomaly, it sends information to the generative model.
[1186] The server analyzes data from the sensors and, if an anomaly is detected, sends that information to the generative model. The input is state data from the sensors, and the output is the transmission of anomaly data to the generative model.
[1187] Step 12:
[1188] The generative model generates methods for dealing with anomalies.
[1189] The generative model generates specific corrective actions based on the type of anomaly. The input is the anomaly data, and the output is a response message describing the corrective action. For example, "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[1190] Step 13:
[1191] The server sends the generative model response to the terminal.
[1192] The server sends the solution obtained from the generative model to the car navigation terminal. The input is the response message of the solution generated by the generative model, and the output is the completion of sending that response message to the terminal.
[1193] Step 14:
[1194] The device notifies the user and provides instructions.
[1195] The terminal uses voice and screen display to notify and guide the user of the solutions derived from the generative model. Input is the response text of the solution sent from the server, and output is voice and display guidance to the user.
[1196] Step 15:
[1197] The user will take action by following the provided instructions.
[1198] The user follows the instructions provided, for example, by going to the nearest gas station for an oil change. Input is the guidance from the terminal, and output is the action taken to address the vehicle problem.
[1199] (Application Example 1)
[1200] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1201] Conventional car navigation systems can only provide static responses to user voice commands, making real-time responses to vehicle abnormalities difficult. Furthermore, there has been a lack of systems capable of sophisticated user interaction in the operation of autonomous vehicles. Solving these problems is essential.
[1202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1203] In this invention, the server includes means for interpreting voice input from a user using a generative model and generating an appropriate response; means for dynamically updating information displayed on a car navigation terminal based on the user's voice instructions; sensors for collecting vehicle status data in real time and detecting anomalies; means for transmitting anomaly data collected from the sensors to the generative model and providing appropriate countermeasures; means for calculating a route to a destination through navigation using the generative model and providing guidance to the autonomous vehicle; and means for monitoring the vehicle status in real time and guiding the user to appropriate countermeasures via the generative model when an anomaly occurs. This enables the provision of dynamic and appropriate responses to the user's voice instructions, and allows for navigation of the autonomous vehicle and rapid response to anomalies.
[1204] A "generative model" is a machine learning model that interprets voice input from a user and generates an appropriate response.
[1205] A "car navigation terminal" is a device installed in a vehicle that displays information based on the user's voice commands.
[1206] "Means" refer to methods or devices used to achieve a specific function or role.
[1207] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[1208] An "autonomous vehicle" is a vehicle that can drive autonomously without the intervention of a human driver.
[1209] "Guidance" refers to information and instructions that provide route information to a destination and guide a vehicle to its destination.
[1210] A "user interface" refers to equipment such as screens and audio output devices that allow users to operate a system and obtain information.
[1211] "Real-time" refers to actions or processes that occur immediately or with a very short delay.
[1212] This invention is a system that uses a generative model to realize advanced interactive functions. The system consists of a car navigation terminal, a server on which the generative model operates, vehicle sensors, and various interfaces.
[1213] System Configuration
[1214] Car navigation terminal
[1215] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[1216] Generative model
[1217] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1218] Sensors and Anomaly Detection
[1219] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1220] User Interface
[1221] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[1222] Operation overview
[1223] 1. Receiving and analyzing voice commands:
[1224] The car navigation system uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice commands into text data.
[1225] This text data will be sent to the server.
[1226] 2. Response generation using generative models:
[1227] The server passes the transformed text data to a generation model (e.g., the text generation model in the transformers library) to generate an appropriate response.
[1228] The generated response is sent back to the car navigation terminal via the server and provided to the user via voice and screen display.
[1229] 3. Route calculation and navigation:
[1230] The car navigation terminal calculates the optimal route to the destination and performs navigation based on the response of the generative model.
[1231] 4. Detection and response to vehicle abnormalities:
[1232] The sensors collect vehicle status data in real time and transmit it to the server.
[1233] When an anomaly is detected, the server passes that information to the generative model, which then generates an appropriate response.
[1234] The generative model provides specific solutions tailored to the type of anomaly and notifies the user via voice and display.
[1235] Specific example
[1236] For example, when a user enters a voice command on a smartphone, the following prompt message is used:
[1237] "Hey GPS, tell me some good restaurants around here."
[1238] In response to this, the system provides the following answer:
[1239] "We have two recommended restaurants nearby: (1) Restaurant A and (2) Cafe B. Which one would you like to go to?"
[1240] Once the user makes a selection, the system calculates the route to the specified destination and provides directions to the autonomous vehicle.
[1241] In this way, we can realize the interactive navigation experience that users desire and enable quick and specific responses in the event of a vehicle malfunction.
[1242] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1243] Step 1:
[1244] Receiving voice commands
[1245] Operation: The user enters voice commands into the car navigation terminal.
[1246] Input: User voice command (e.g., "Tell me some good restaurants nearby").
[1247] Output: Audio data.
[1248] Specific operation: The car navigation terminal collects the user's voice through the microphone.
[1249] Step 2:
[1250] Converting audio data to text
[1251] Operation: The device uses a speech recognition engine to convert speech data into text data.
[1252] Input: Audio data.
[1253] Output: Text data (e.g., "Tell me some good restaurants nearby").
[1254] Specific operation: Using a speech recognition library (e.g., speech_recognition), the system analyzes audio data and converts it into text data.
[1255] Step 3:
[1256] Sending text data to the server
[1257] Operation: The terminal sends the converted text data to the server.
[1258] Input: Text data.
[1259] Output: Notification that the text data has been successfully sent to the server.
[1260] Specific operation: Use the communication module in the terminal to send text data to the server.
[1261] Step 4:
[1262] Response generation using generative models
[1263] Operation: The server inputs the received text data into a generative model and generates an appropriate response.
[1264] Input: Text data.
[1265] Output: Generated response (e.g., "Recommended nearby restaurants are (1) Restaurant A and (2) Cafe B.").
[1266] Specific operation: The server inputs text data into a generative model (e.g., the text generation model in the transformers library), analyzes the user's intent, and generates an appropriate response.
[1267] Step 5:
[1268] Sending response data to the terminal
[1269] Operation: The server sends the generated response data to the terminal.
[1270] Input: Generated response data.
[1271] Output: Notification that the response data has been successfully sent to the terminal.
[1272] Specific operation: The generated response data is sent to the car navigation terminal using the communication module within the server.
[1273] Step 6:
[1274] Providing responses to users
[1275] Operation: The terminal provides the user with received response data via voice and screen display.
[1276] Input: Response data.
[1277] Output: Audio output and screen display (e.g., "Recommended nearby restaurants include (1) Restaurant A and (2) Cafe B.").
[1278] Specific operation: The device will reply using a speech synthesis engine (e.g., tts_engine) and simultaneously display it on the screen.
[1279] Step 7:
[1280] User Selection
[1281] Operation: The user makes a selection regarding the response via voice or touch.
[1282] Input: User selection (e.g., "I want to go to restaurant A (1)").
[1283] Output: Selected data.
[1284] Specific operation: The device collects user selections via voice recognition or screen touch.
[1285] Step 8:
[1286] Route calculation and navigation start
[1287] Operation: The device calculates the route to the selected destination and begins navigation.
[1288] Input: Selected data (e.g., "Restaurant A").
[1289] Output: Route data and navigation guide.
[1290] Specific operation: The device uses a navigation library (e.g., navigation) to calculate the optimal route to the destination and starts navigation.
[1291] Step 9:
[1292] Real-time monitoring of vehicle status
[1293] Operation: Sensors collect vehicle status data in real time and send it to the server.
[1294] Input: Vehicle status data.
[1295] Output: Notification that vehicle status data has been successfully sent to the server.
[1296] Specific operation: Sensors inside the vehicle collect status data in real time and send it to the server.
[1297] Step 10:
[1298] Vehicle abnormality detection
[1299] Operation: The server analyzes the received vehicle status data and detects any abnormalities.
[1300] Input: Vehicle status data.
[1301] Output: Anomaly detection notification (e.g., "Oil change required").
[1302] Specific operation: The analysis engine within the server analyzes the status data and detects anomalies.
[1303] Step 11:
[1304] Generating solutions using generative models
[1305] Operation: The server inputs anomaly detection data into a generation model and generates appropriate countermeasures.
[1306] Input: Anomaly detection data.
[1307] Output: Troubleshooting steps (e.g., "Oil change is needed. Please follow these steps.").
[1308] Specific operation: The server inputs anomaly detection data into the generative model and generates specific countermeasures for the user.
[1309] Step 12:
[1310] Sending the troubleshooting steps to the device
[1311] Operation: The server sends the generated troubleshooting data to the terminal.
[1312] Input: Data on how to handle the situation.
[1313] Output: Notification that the data regarding the solution has been successfully sent to the terminal.
[1314] Specific operation: The generated troubleshooting data is sent to the car navigation terminal using a communication module on the server.
[1315] Step 13:
[1316] Providing users with solutions
[1317] Operation: The device provides the user with received troubleshooting data via voice and screen display.
[1318] Input: Data on how to handle the situation.
[1319] Output: Audio output and screen display (e.g., "Oil change is needed. Please follow the next steps.").
[1320] Specific operation: The device uses a speech synthesis engine to notify the user of the solution via voice, and simultaneously displays it on the screen.
[1321] Step 14:
[1322] User Action
[1323] Operation: The system will address the vehicle malfunction according to the troubleshooting steps provided to the user.
[1324] Input: Display of troubleshooting methods and voice guidance.
[1325] Output: Normal vehicle condition.
[1326] Specific actions: The user follows the provided instructions and performs the actual troubleshooting steps.
[1327] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1328] This invention combines an emotion engine with a car navigation system that interprets voice input from the user using a generative model and generates an appropriate response. The system consists of a car navigation terminal, a server, a generative model, vehicle sensors, and an emotion engine that recognizes the user's emotions. This improves the user's driving experience and enhances safety.
[1329] System configuration and operation
[1330] This system consists of a car navigation terminal, a server on which the generative model runs, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[1331] Car navigation terminal
[1332] A car navigation terminal is installed in a vehicle and handles voice command input, output of responses using a generative model and emotion engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[1333] Generative model
[1334] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1335] Sensors and Anomaly Detection
[1336] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1337] Emotional Engine
[1338] The emotion engine recognizes the user's emotions from their voice input and the in-car environment. The emotions recognized by the emotion engine are then used to customize the generative model's responses and navigation information. In particular, if signs of stress or fatigue are detected, it includes features that recommend relaxation spots and rest areas.
[1339] User Interface
[1340] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. Settings include language, voice guidance type, and theme color, which can be customized to the user's preferences.
[1341] Explanation of the program's processing
[1342] Here, the system's operation is explained in natural language, describing the program's processing.
[1343] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1344] 2. Terminal: Activates the speech recognition engine to convert the user's voice commands from speech to text. This text data is then sent to the server.
[1345] 3. Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[1346] 4. Generative Model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1347] 5. Server: Sends the generated response data to the car navigation terminal.
[1348] 6. Terminal: Provides the user with the received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[1349] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[1350] 8. Terminal: Calculates the selected route and starts guidance. Updates route information in real time.
[1351] 9. Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[1352] 10. Server: Inputs data from the emotion engine into the generative model to generate a customized response based on emotion. Example: "Would you recommend a relaxation spot?"
[1353] 11. Sensors: Monitor and collect various vehicle status data in real time.
[1354] 12. Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1355] 13. Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[1356] 14. Server: Sends response data, including the solution created by the generative model, to the car navigation terminal.
[1357] 15. Device: Notify the user of how to resolve the issue via voice and on-screen display. Provide specific instructions.
[1358] 16. User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the user will be guided to the nearest gas station or repair shop.
[1359] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users, and further enables quick and specific responses in the event of vehicle malfunctions, thereby enhancing safety and comfort. By combining it with an emotional engine, flexible responses tailored to the user's psychological state become possible, further improving the driving experience.
[1360] The following describes the processing flow.
[1361] Step 1:
[1362] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1363] Step 2:
[1364] Terminal: Activates the speech recognition engine and converts the user's voice commands into text data. This text data is then sent to the server.
[1365] Step 3:
[1366] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[1367] Step 4:
[1368] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1369] Step 5:
[1370] Server: Sends the generated response data to the car navigation terminal.
[1371] Step 6:
[1372] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[1373] Step 7:
[1374] User: Select from the suggested options. Example: "I want to go to spot (1)."
[1375] Step 8:
[1376] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[1377] Step 9:
[1378] Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[1379] Step 10:
[1380] Server: Passes emotional data from the emotion engine to the generative model, which then generates an emotionally appropriate response. Example: "Would you recommend a relaxation spot?"
[1381] Step 11:
[1382] Generative model: Generates customized responses based on emotions. Example: "You seem stressed, so I recommend the following relaxation spots: (1)~~ (2)~~ (3)~~"
[1383] Step 12:
[1384] Server: Sends the generated response data to the car navigation terminal.
[1385] Step 13:
[1386] Terminal: Provides the user with received response data via voice output and screen display. Example: "We recommend the following relaxation spots. (1)~~ (2)~~ (3)~~ Where would you like us to take you?"
[1387] Step 14:
[1388] User: Select from the suggested relaxation spots. Example: "I want to go to relaxation spot (2)."
[1389] Step 15:
[1390] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[1391] Step 16:
[1392] Sensors: Monitor and collect various vehicle status data in real time.
[1393] Step 17:
[1394] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1395] Step 18:
[1396] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[1397] Step 19:
[1398] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[1399] Step 20:
[1400] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[1401] Step 21:
[1402] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[1403] This process allows the system to provide an interactive navigation experience while taking the user's emotional state into consideration, and enables quick and specific action in the event of a vehicle malfunction.
[1404] (Example 2)
[1405] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1406] Conventional car navigation systems simply interpret user voice input to provide route guidance, making it difficult to respond in accordance with the user's emotions or the vehicle's malfunction. Furthermore, the lack of customized responses tailored to the user's psychological state meant that stress and fatigue during driving could not be reduced. Additionally, there was a problem with the inability to quickly provide specific solutions in the event of a vehicle malfunction.
[1407] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, means for dynamically updating information displayed on the navigation terminal based on the user's voice instructions, sensors for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensors to the generation model and providing an appropriate countermeasure, an emotion engine for recognizing the user's emotions from the user's voice input and the situation inside the vehicle and reflecting this in the response of the generation model, and means for customizing navigation information based on the user's emotions. This enables flexible responses according to the user's psychological state and quick and specific countermeasures in the event of a vehicle abnormality.
[1408] A "generative model" is a type of artificial intelligence that analyzes input data and generates appropriate responses or predictions based on that data.
[1409] A "navigation terminal" is a device installed inside a vehicle that provides map information and route guidance.
[1410] "Voice instructions" refer to instructions or commands spoken by the user, which serve as input for the navigation system to recognize and process.
[1411] A "sensor" is a device that collects vehicle status data in real time and detects specific conditions or abnormalities.
[1412] An "emotion engine" is software or an algorithm that analyzes the user's emotions from their voice input and the situation inside the vehicle, and reflects that emotional information in the response of a generative model.
[1413] "Handling instructions" refer to specific actions and procedures that the user should take when they detect a problem or abnormality in the vehicle.
[1414] "Customization" refers to modifying the interface and responses based on the user's settings and status to meet individual needs.
[1415] Modes for carrying out the invention
[1416] The car navigation system of the present invention enhances the user's driving experience and improves safety by combining a generative model, a voice recognition engine, an emotion engine, and various sensors. The present invention consists of a navigation terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[1417] Navigation terminal
[1418] The navigation terminal is installed in the vehicle, takes user voice commands as input, outputs responses using a generative model and emotion engine, and displays route guidance. When the user inputs a voice command such as "I want to go for a drive to clear my head," the terminal uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the voice into text data and sends it to the server.
[1419] Generative model
[1420] The server passes the text data received from the user to a generative model (e.g., GPT-3) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is sent to the navigation terminal via the server.
[1421] Sensors and Anomaly Detection
[1422] The vehicle is equipped with various sensors, including tire pressure sensors and an engine diagnostic system, to collect vehicle status data in real time. This data is sent to a server, which analyzes it and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1423] Emotional Engine
[1424] The emotion engine recognizes the user's emotions from their voice input and the environment inside the vehicle. For example, if the user is feeling stressed, the emotion engine detects this and sends the data to the generative model. The generative model then generates a customized response based on the emotion data, such as recommending a relaxation spot to the user.
[1425] User Interface
[1426] The user interface consists of a screen and audio for the navigation terminal and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, which can be customized to the user's preferences.
[1427] Specific example
[1428] When a user inputs a voice command such as "I want to go for a drive to clear my head," the voice recognition engine converts the voice into text and sends it to the server. The server passes this text data to a generative model, which generates a response such as "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~." The generated response is sent to the navigation terminal, where the user confirms the recommended spots on the screen and by voice and selects a destination. The navigation terminal then calculates the selected route and begins guidance.
[1429] Thus, by utilizing a generative AI model, a speech recognition engine, an emotion engine, and sensors, the present invention enables flexible responses that respond to the user's intentions and emotions, realizing a car navigation system that achieves both safety and comfort.
[1430] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1431] Program processing flow
[1432] Step 1:
[1433] User: The user enters a voice command. Specifically, they say, "I want to go for a drive to clear my head."
[1434] Input: User's voice
[1435] Output: Raw data audio file
[1436] Step 2:
[1437] Device: Uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert user voice commands from speech to text.
[1438] Input: Raw audio file
[1439] Output: Text data (Example: "I want to go for a drive to clear my head")
[1440] Step 3:
[1441] Terminal: Sends text data to the server.
[1442] Input: Text data
[1443] Output: Sending text data to the server
[1444] Step 4:
[1445] Server: Passes the received text data to a generative model (e.g., GPT-3) and starts the analysis.
[1446] Input: Text data
[1447] Output: User intent (e.g., "Recommendations for destinations suitable for a change of pace")
[1448] Step 5:
[1449] Generative model: Generates appropriate responses based on the user's intent. For example, it might create a response like, "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~".
[1450] Input: User intent
[1451] Output: Response text (Example: "Since it's nighttime now, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~")
[1452] Step 6:
[1453] Server: Sends the generated response text data to the navigation terminal.
[1454] Input: Response text data
[1455] Output: Sending response data to the navigation terminal
[1456] Step 7:
[1457] Terminal: Provides the user with the received response text via voice output (e.g., Amazon Polly) and screen display. Specifically, it would say, "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[1458] Input: Response text data
[1459] Output: Voice guidance and screen display
[1460] Step 8:
[1461] User: The user selects from the provided options and responds, "I want to go to spot (1)."
[1462] Input: User Selection
[1463] Output: Selection (Example: "(1) Spot")
[1464] Step 9:
[1465] Terminal: Calculates the selected route using the car navigation system's route calculation engine (e.g., Google Maps API) and begins guidance. Route information is updated in real time.
[1466] Input: Selection
[1467] Output: Route guidance and real-time updates
[1468] Step 10:
[1469] Emotion Engine: Analyzes user voice input and in-car environment to recognize user emotions. Specifically, it determines whether the user is experiencing stress.
[1470] Input: User voice, in-vehicle status data
[1471] Output: Recognized emotion (e.g., "I am feeling stressed")
[1472] Step 11:
[1473] Server: Inputs data from the emotion engine into the generative model to generate customized responses based on the user's emotions. For example, it might create a response such as, "Would you recommend a relaxation spot?"
[1474] Input: Sentiment data
[1475] Output: Response text data (e.g., "Do you recommend any relaxation spots?")
[1476] Step 12:
[1477] Sensors: Collect various status data (tire pressure, engine diagnostics, etc.) in real time.
[1478] Input: Vehicle condition
[1479] Output: Status data
[1480] Step 13:
[1481] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1482] Input: Status data from the sensor
[1483] Output: Abnormal data and notifications
[1484] Step 14:
[1485] Generative Model: Based on abnormal data, it generates appropriate countermeasures and creates instructions such as, "The oil change light is on. Please take the following steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[1486] Input: Abnormal data
[1487] Output: Text data of the solution.
[1488] Step 15:
[1489] Server: Sends response data containing the solution generated by the generative model to the navigation terminal.
[1490] Input: Text data of the solution
[1491] Output: Send to navigation terminal
[1492] Step 16:
[1493] Terminal: Notifies the user of how to resolve the issue via voice and on-screen display, and guides them through specific steps. For example, it might say, "The oil change light is on. Please follow these steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[1494] Input: Text data of the solution
[1495] Output: Voice guidance and screen display
[1496] Step 17:
[1497] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[1498] Input: Instructions on how to proceed
[1499] Output: Appropriate countermeasures
[1500] (Application Example 2)
[1501] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1502] In recent years, the demand for food delivery services has increased, requiring drivers to perform delivery duties efficiently and safely. However, excessive stress and fatigue during driving can reduce drivers' attention span and increase the risk of accidents. Furthermore, current car navigation systems cannot provide responses that take into account the driver's emotional state, and improvements are needed to enhance the driver's driving experience. Therefore, a system is needed that analyzes the driver's emotional state in real time and generates appropriate responses.
[1503] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, an in-vehicle information terminal, means for dynamically updating the information displayed on the in-vehicle information terminal based on the user's voice instructions, a sensor for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensor to the generation model and providing an appropriate countermeasure, emotion analysis means for analyzing the user's emotional state and generating a response corresponding to that emotion, and means for suggesting rest locations such as relaxation spots according to the user's emotional state recognized by the emotion analysis means. This makes it possible to detect the stress and fatigue of the driver and provide appropriate responses and rest suggestions, thereby enabling safe and efficient delivery operations.
[1504] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[1505] An "in-vehicle information terminal" is a device installed in a vehicle that allows for the input of voice commands, the output of responses from a generative model and sentiment analysis engine, and the display of route guidance.
[1506] A "dynamically updating method" refers to a method of changing the information displayed on the in-vehicle information terminal in real time based on the user's voice commands.
[1507] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[1508] "Means of providing appropriate countermeasures" refers to a method of transmitting anomaly data collected from sensors to a generation model and then presenting specific countermeasures based on that data.
[1509] "Emotion analysis means" refers to a method that recognizes the emotional state from the user's voice input or the situation inside the vehicle and provides that information to a generative model.
[1510] A "relaxation spot" is a place where users can rest and refresh themselves when they feel stressed or tired.
[1511] This invention is a system that supports food delivery drivers in performing their duties safely and efficiently, and consists of the following components.
[1512] System configuration
[1513] This system consists of an in-vehicle information terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion analysis engine.
[1514] In-vehicle information terminal
[1515] The in-vehicle information terminal is installed in the vehicle and handles voice command input, output of responses generated by a generative model and sentiment analysis engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to the server. For example, if a driver voice-inputs "I'm tired, I want to take a short break," the terminal converts that voice into text and sends it to the server.
[1516] Generative model
[1517] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and emotional state to generate a specific and appropriate response. The generated response is returned to the in-vehicle information terminal via the server and provided to the user. For example, in response to input such as "I'm tired, I want to take a break," a response like "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?" is generated.
[1518] Sensors and Anomaly Detection
[1519] Sensors within the vehicle collect vehicle status data in real time and send it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps. For example, if it detects that an oil change is needed, it will provide countermeasures such as, "The oil change lamp is lit. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[1520] Emotion analysis engine
[1521] The emotion analysis engine recognizes the user's emotional state from their voice input and the environment inside the vehicle. The user's emotional state recognized by the emotion analysis engine is reflected in the generative model's responses and the customization of navigation information. In particular, if the system detects that the user is feeling "tired" or "stressed," it includes a function that suggests relaxation spots and rest areas. For example, if the user inputs "tired," the system detects fatigue and suggests nearby relaxation spots.
[1522] User Interface
[1523] The user interface consists of the in-car information terminal's screen and voice, and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, all of which can be customized to the user's preferences. For example, simply saying "I want to change the voice guidance voice" will change it to the preferred voice type.
[1524] Specific example
[1525] 1. User input: "I'm tired, I want to take a break."
[1526] 2. Example of a prompt:
[1527] Emotion: Fatigue, Input: I'm tired and want to take a break. Can you suggest any relaxation spots?
[1528] 3. Response generated: "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[1529] This configuration allows for the detection of driver stress and fatigue, and enables safe and efficient delivery operations by providing appropriate responses and rest suggestions.
[1530] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1531] Step 1:
[1532] The user inputs a voice command. For example, they might say, "I'm tired, I want to take a break." This voice input becomes the starting point for the system's processing.
[1533] Step 2:
[1534] The device activates a speech recognition engine to convert the user's voice commands from speech to text. The input is the user's voice data, and the output is the text data of that voice. The device sends this text data to the server.
[1535] Step 3:
[1536] The server passes the received text data to the generative model. The input is text data sent from the terminal, and the generative model is used to analyze the user's intent and emotions. The output is data containing the analysis results.
[1537] Step 4:
[1538] The generative model generates appropriate responses based on the user's intentions and emotions. The input is the parsed data obtained in the previous step, and the output is the response sentence presented to the user. For example, it might generate a response such as, "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[1539] Step 5:
[1540] The server transmits the generated response data to the in-vehicle infotainment terminal. The input is the response data from the generative model, and the output is the communication data to the in-vehicle infotainment terminal.
[1541] Step 6:
[1542] The terminal provides the user with the received response data through voice output and screen display. The input is the response data from the server, and the output is voice guidance and screen display to the user. For example, it might provide voice guidance such as, "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[1543] Step 7:
[1544] The user selects an option from the suggested relaxation spots and gives a voice command. For example, they might respond, "I want to go to Cafe B." This voice data becomes the input for the next process.
[1545] Step 8:
[1546] The device restarts its speech recognition engine and converts the user's selection from speech to text. The input is the user's voice data, and the output is the text data of the selections. The device sends this text data to the server.
[1547] Step 9:
[1548] The server calculates the route to the selected relaxation spot based on the received text data. The input is the text data of the selected spot, and the output is the data of the optimal route.
[1549] Step 10:
[1550] The server transmits the calculated route to the in-vehicle information terminal. The input is the route calculation result data, and the output is the communication data to the in-vehicle information terminal.
[1551] Step 11:
[1552] The terminal starts navigation based on the received route data and updates the route information in real time. The input is route data from the server, and the output is navigation guidance and screen display.
[1553] Step 12:
[1554] The emotion analysis engine analyzes the user's voice input and the conditions inside the vehicle to recognize the user's emotions. The input is the user's voice data and sensor information, and the output is emotional state data.
[1555] Step 13:
[1556] The server inputs data obtained from the emotion analysis engine into a generative model to generate a customized response based on the emotion. The input is emotional state data, and the output is customized response data. For example, it can generate a response such as, "Would you recommend a relaxation spot?"
[1557] Step 14:
[1558] The sensors monitor and collect various vehicle status data in real time. The input is vehicle status information, and the output is data signals from the sensors.
[1559] Step 15:
[1560] The server continuously receives data from the sensor and, if it detects an anomaly, sends that information to the generative model. The input is the data signal from the sensor, and the output is the anomaly data sent to the generative model.
[1561] Step 16:
[1562] The generative model generates appropriate countermeasures based on abnormal data and creates specific procedures. The input is abnormal data, and the output is procedure data for the countermeasures. For example, it generates procedures such as "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[1563] Step 17:
[1564] The server sends response data, including the countermeasures created by the generative model, to the in-vehicle information terminal. The input is the procedural data of the countermeasures, and the output is the communication data to the in-vehicle information terminal.
[1565] Step 18:
[1566] The device notifies the user of troubleshooting methods via voice and on-screen display, guiding them through specific steps. Input is troubleshooting data from the server, and output is voice guidance and on-screen display to the user.
[1567] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1568] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1569] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1570] [Fourth Embodiment]
[1571] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1572] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1573] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1574] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1575] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1576] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1577] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1578] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1579] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1580] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1581] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1582] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1583] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1584] This invention relates to a car navigation system that uses a generative model to achieve advanced interactive functions. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination.
[1585] System Configuration
[1586] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[1587] Car navigation terminal
[1588] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[1589] Generative model
[1590] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1591] Sensors and Anomaly Detection
[1592] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1593] User Interface
[1594] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[1595] Explanation of the program's processing
[1596] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1597] 2. Terminal: Activates the speech recognition engine to convert voice commands into text data. Then sends that text data to the server.
[1598] 3. Server: Passes text data to the generative model and generates an appropriate response.
[1599] 4. Generative Model: Analyzes user intent and generates appropriate responses. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1600] 5. Server: Sends the response data from the generated model to the car navigation terminal.
[1601] 6. Device: Provides responses to the user via voice and screen display.
[1602] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[1603] 8. Terminal: Calculates the selected route and starts the guidance.
[1604] 9. Sensors: Perform continuous monitoring of the vehicle's condition.
[1605] 10. Server: If an anomaly is detected, it sends that information to the generation model to generate an appropriate countermeasure.
[1606] 11. Generative Model: Provide specific solutions tailored to the type of anomaly. Example: "The oil change light is on. Please follow these steps: (1) (2) (3)"
[1607] 12. Server: Sends the response data from the generated model to the car navigation terminal.
[1608] 13. Device: Notifies the user via voice and display and guides them through the instructed course of action.
[1609] 14. User: Follow the provided procedures to address any vehicle malfunctions.
[1610] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users and enables quick and specific responses in the event of vehicle malfunctions.
[1611] The following describes the processing flow.
[1612] Step 1:
[1613] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1614] Step 2:
[1615] Terminal: Activates the speech recognition engine to convert user voice commands from speech to text, and then sends it to the server.
[1616] Step 3:
[1617] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[1618] Step 4:
[1619] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1620] Step 5:
[1621] Server: Sends the generated response data to the car navigation terminal.
[1622] Step 6:
[1623] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[1624] Step 7:
[1625] User: Choose from the recommended options. Example: "I want to go to spot (1)."
[1626] Step 8:
[1627] Terminal: Calculates the selected route and starts navigation. Updates route information in real time.
[1628] Step 9:
[1629] Sensors: Monitor and collect various vehicle status data in real time.
[1630] Step 10:
[1631] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1632] Step 11:
[1633] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[1634] Step 12:
[1635] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[1636] Step 13:
[1637] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[1638] Step 14:
[1639] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[1640] (Example 1)
[1641] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1642] Conventional car navigation systems had problems with responding quickly and appropriately to user voice input and handling vehicle malfunctions. In particular, it was difficult to provide real-time responses to voice commands and to effectively coordinate vehicle status monitoring and anomaly detection. Furthermore, they lacked an intuitive interface that users could easily operate. As a result, the user experience was significantly reduced, and there was a risk of compromising driving safety.
[1643] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1644] In this invention, the server includes means for converting user voice into text data and transmitting it, means for passing the received text data to a generation model to generate a response, and means for transmitting the generated response to a car navigation terminal and providing it to the user in both voice and screen display. This enables real-time responses to user voice commands, vehicle status monitoring, and provision of appropriate countermeasures.
[1645] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[1646] A "car navigation terminal" is a device installed in a vehicle that takes voice commands as input, outputs generated responses, and displays route guidance.
[1647] A "speech recognition engine" refers to software or hardware that converts a user's voice into text data.
[1648] A "server" is a computing system on which the generative model operates and processes various types of data in conjunction with car navigation terminals and sensors.
[1649] A "sensor" is a device used to monitor the vehicle's condition in real time and collect data.
[1650] "Text data" refers to the character data generated by a speech recognition engine after analyzing speech.
[1651] "User interface" refers to the means, such as screens and sounds, that users use to operate a system.
[1652] "Abnormal data" refers to information about abnormal conditions in a vehicle detected by sensors.
[1653] "Handling procedures" refers to the steps and instructions for taking appropriate action in response to a vehicle malfunction.
[1654] "Response" refers to the content of the reply generated by the generative model based on the user's voice input.
[1655] This invention relates to a car navigation system that achieves advanced interactive functions using a generative model. For the invention to be implemented, the car navigation terminal, generative model, vehicle sensors, and various interfaces must work together in coordination. The configuration and specific operation of this system are described in detail below.
[1656] System Configuration
[1657] This system consists of a car navigation terminal, a server on which the generation model runs, sensors that monitor the vehicle's status, and a user interface.
[1658] Car navigation terminal
[1659] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice into text data and send it to a server.
[1660] Generative model
[1661] The server passes the text data received from the user to a generative model (e.g., OpenAI GPT-4) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1662] Specific example
[1663] When a user enters a voice command into the car navigation system, such as "I want to go for a drive to clear my head," the navigation system uses its voice recognition engine to convert the voice command into text data. This text data is then sent to a server via the internet.
[1664] The server passes the received text data to the generative model, which generates a response such as, "It's nighttime now, so we recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~". The generated response is then sent back to the car navigation terminal via the server.
[1665] The car navigation terminal uses a speech synthesis engine (e.g., Amazon Polly) and a screen display to provide responses to the user. The user selects "I want to go to spot (1)" from the recommended options, and the car navigation terminal calculates the selected route using a GPS navigation system (e.g., Google Maps API) and begins providing guidance via voice and screen.
[1666] Sensors and Anomaly Detection
[1667] Sensors installed inside the vehicle constantly monitor the vehicle's status. Vehicle status data is collected in real time and transmitted to a server. The server analyzes this data and, for example, if the oil lamp illuminates, sends that information to a generating model. The generating model then generates specific instructions for an oil change, such as "The oil change lamp is illuminated. Please follow these steps: (1) ~~ (2) ~~ (3) ~~," and transmits this information back to the car navigation terminal via the server.
[1668] The car navigation system notifies the user through voice and display, guiding them through the instructed course of action. The user then follows the suggested steps, for example, heading to the nearest gas station for an oil change.
[1669] Through the configuration and specific operation described above, this system realizes the interactive navigation experience that users desire and enables quick and specific responses in the event of a vehicle malfunction.
[1670] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1671] Step 1:
[1672] The user enters a voice command.
[1673] The specific action involves the user issuing a voice command to the car navigation terminal, such as "I want to go for a drive to clear my head." At this stage, the input is the user's voice.
[1674] Step 2:
[1675] The terminal converts voice commands into text data.
[1676] A speech recognition engine (e.g., Google Cloud Speech-to-Text) works to analyze the user's voice commands and generate text data. The input is the user's voice data, and the output is the corresponding text data.
[1677] Step 3:
[1678] The terminal sends text data to the server.
[1679] Text data converted from speech is sent from the terminal to the server via the network. The input is the text data generated by the speech recognition engine, and the output is the completion of the transmission of the text data to the server.
[1680] Step 4:
[1681] The server passes text data to the generative model.
[1682] The server provides the received text data to a generative AI model (e.g., OpenAI GPT-4). The generative model analyzes the text data and understands the user's intent. The input is the text data received by the server, and the output is the analysis request to the generative model.
[1683] Step 5:
[1684] The generative model generates the response.
[1685] The generative model generates an appropriate response based on the analysis results. For example, it might generate a response like, "Since it's nighttime, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~". The input is text data passed from the server, and the output is the generated response.
[1686] Step 6:
[1687] The server sends the generative model response to the terminal.
[1688] The server sends the response message obtained from the generative model to the car navigation terminal. The input is the response message generated by the generative model, and the output is the completion of sending that response message to the terminal.
[1689] Step 7:
[1690] The device provides responses to the user through voice and on-screen displays.
[1691] The device uses a speech synthesis engine (e.g., Amazon Polly) and a display to provide the user with generated responses as audio and visual information. The input is the response text sent from the server, and the output is the audio and screen display response to the user.
[1692] Step 8:
[1693] The user selects from the recommended options.
[1694] The user makes selections using voice or a touch panel, for example, "I want to go to spot (1)." Input is provided by the terminal display and voice guidance, while output is the user's selection.
[1695] Step 9:
[1696] The device calculates the selected route and begins providing directions.
[1697] The device uses a GPS navigation system (e.g., Google Maps API) to calculate the selected route. The input is the user's selection, and the output is the selected route and the start of navigation.
[1698] Step 10:
[1699] Sensors monitor the vehicle's status.
[1700] Sensors installed inside the vehicle monitor engine temperature, tire pressure, and other parameters in real time. Inputs are various vehicle status data, and outputs are data collected by the sensors.
[1701] Step 11:
[1702] If the server detects an anomaly, it sends information to the generative model.
[1703] The server analyzes data from the sensors and, if an anomaly is detected, sends that information to the generative model. The input is state data from the sensors, and the output is the transmission of anomaly data to the generative model.
[1704] Step 12:
[1705] The generative model generates methods for dealing with anomalies.
[1706] The generative model generates specific corrective actions based on the type of anomaly. The input is the anomaly data, and the output is a response message describing the corrective action. For example, "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[1707] Step 13:
[1708] The server sends the generative model response to the terminal.
[1709] The server sends the solution obtained from the generative model to the car navigation terminal. The input is the response message of the solution generated by the generative model, and the output is the completion of sending that response message to the terminal.
[1710] Step 14:
[1711] The device notifies the user and provides instructions.
[1712] The terminal uses voice and screen display to notify and guide the user of the solutions derived from the generative model. Input is the response text of the solution sent from the server, and output is voice and display guidance to the user.
[1713] Step 15:
[1714] The user will take action by following the provided instructions.
[1715] The user follows the instructions provided, for example, by going to the nearest gas station for an oil change. Input is the guidance from the terminal, and output is the action taken to address the vehicle problem.
[1716] (Application Example 1)
[1717] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1718] Conventional car navigation systems can only provide static responses to user voice commands, making real-time responses to vehicle abnormalities difficult. Furthermore, there has been a lack of systems capable of sophisticated user interaction in the operation of autonomous vehicles. Solving these problems is essential.
[1719] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1720] In this invention, the server includes means for interpreting voice input from a user using a generative model and generating an appropriate response; means for dynamically updating information displayed on a car navigation terminal based on the user's voice instructions; sensors for collecting vehicle status data in real time and detecting anomalies; means for transmitting anomaly data collected from the sensors to the generative model and providing appropriate countermeasures; means for calculating a route to a destination through navigation using the generative model and providing guidance to the autonomous vehicle; and means for monitoring the vehicle status in real time and guiding the user to appropriate countermeasures via the generative model when an anomaly occurs. This enables the provision of dynamic and appropriate responses to the user's voice instructions, and allows for navigation of the autonomous vehicle and rapid response to anomalies.
[1721] A "generative model" is a machine learning model that interprets voice input from a user and generates an appropriate response.
[1722] A "car navigation terminal" is a device installed in a vehicle that displays information based on the user's voice commands.
[1723] "Means" refer to methods or devices used to achieve a specific function or role.
[1724] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[1725] An "autonomous vehicle" is a vehicle that can drive autonomously without the intervention of a human driver.
[1726] "Guidance" refers to information and instructions that provide route information to a destination and guide a vehicle to its destination.
[1727] A "user interface" refers to equipment such as screens and audio output devices that allow users to operate a system and obtain information.
[1728] "Real-time" refers to actions or processes that occur immediately or with a very short delay.
[1729] This invention is a system that uses a generative model to realize advanced interactive functions. The system consists of a car navigation terminal, a server on which the generative model operates, vehicle sensors, and various interfaces.
[1730] System Configuration
[1731] Car navigation terminal
[1732] A car navigation terminal is a device installed in a vehicle that takes voice commands as input, outputs responses using a generative model, and displays route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[1733] Generative model
[1734] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1735] Sensors and Anomaly Detection
[1736] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1737] User Interface
[1738] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. It is also customizable to the user's preferences, with settings including language, voice guidance type, and theme color.
[1739] Operation overview
[1740] 1. Receiving and analyzing voice commands:
[1741] The car navigation system uses a speech recognition engine (e.g., the speech_recognition library) to convert the user's voice commands into text data.
[1742] This text data will be sent to the server.
[1743] 2. Response generation using generative models:
[1744] The server passes the transformed text data to a generation model (e.g., the text generation model in the transformers library) to generate an appropriate response.
[1745] The generated response is sent back to the car navigation terminal via the server and provided to the user via voice and screen display.
[1746] 3. Route calculation and navigation:
[1747] The car navigation terminal calculates the optimal route to the destination and performs navigation based on the response of the generative model.
[1748] 4. Detection and response to vehicle abnormalities:
[1749] The sensors collect vehicle status data in real time and transmit it to the server.
[1750] When an anomaly is detected, the server passes that information to the generative model, which then generates an appropriate response.
[1751] The generative model provides specific solutions tailored to the type of anomaly and notifies the user via voice and display.
[1752] Specific example
[1753] For example, when a user enters a voice command on a smartphone, the following prompt message is used:
[1754] "Hey GPS, tell me some good restaurants around here."
[1755] In response to this, the system provides the following answer:
[1756] "We have two recommended restaurants nearby: (1) Restaurant A and (2) Cafe B. Which one would you like to go to?"
[1757] Once the user makes a selection, the system calculates the route to the specified destination and provides directions to the autonomous vehicle.
[1758] In this way, we can realize the interactive navigation experience that users desire and enable quick and specific responses in the event of a vehicle malfunction.
[1759] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1760] Step 1:
[1761] Receiving voice commands
[1762] Operation: The user enters voice commands into the car navigation terminal.
[1763] Input: User voice command (e.g., "Tell me some good restaurants nearby").
[1764] Output: Audio data.
[1765] Specific operation: The car navigation terminal collects the user's voice through the microphone.
[1766] Step 2:
[1767] Converting audio data to text
[1768] Operation: The device uses a speech recognition engine to convert speech data into text data.
[1769] Input: Audio data.
[1770] Output: Text data (e.g., "Tell me some good restaurants nearby").
[1771] Specific operation: Using a speech recognition library (e.g., speech_recognition), the system analyzes audio data and converts it into text data.
[1772] Step 3:
[1773] Sending text data to the server
[1774] Operation: The terminal sends the converted text data to the server.
[1775] Input: Text data.
[1776] Output: Notification that the text data has been successfully sent to the server.
[1777] Specific operation: Use the communication module in the terminal to send text data to the server.
[1778] Step 4:
[1779] Response generation using generative models
[1780] Operation: The server inputs the received text data into a generative model and generates an appropriate response.
[1781] Input: Text data.
[1782] Output: Generated response (e.g., "Recommended nearby restaurants are (1) Restaurant A and (2) Cafe B.").
[1783] Specific operation: The server inputs text data into a generative model (e.g., the text generation model in the transformers library), analyzes the user's intent, and generates an appropriate response.
[1784] Step 5:
[1785] Sending response data to the terminal
[1786] Operation: The server sends the generated response data to the terminal.
[1787] Input: Generated response data.
[1788] Output: Notification that the response data has been successfully sent to the terminal.
[1789] Specific operation: The generated response data is sent to the car navigation terminal using the communication module within the server.
[1790] Step 6:
[1791] Providing responses to users
[1792] Operation: The terminal provides the user with received response data via voice and screen display.
[1793] Input: Response data.
[1794] Output: Audio output and screen display (e.g., "Recommended nearby restaurants include (1) Restaurant A and (2) Cafe B.").
[1795] Specific operation: The device will reply using a speech synthesis engine (e.g., tts_engine) and simultaneously display it on the screen.
[1796] Step 7:
[1797] User Selection
[1798] Operation: The user makes a selection regarding the response via voice or touch.
[1799] Input: User selection (e.g., "I want to go to restaurant A (1)").
[1800] Output: Selected data.
[1801] Specific operation: The device collects user selections via voice recognition or screen touch.
[1802] Step 8:
[1803] Route calculation and navigation start
[1804] Operation: The device calculates the route to the selected destination and begins navigation.
[1805] Input: Selected data (e.g., "Restaurant A").
[1806] Output: Route data and navigation guide.
[1807] Specific operation: The device uses a navigation library (e.g., navigation) to calculate the optimal route to the destination and starts navigation.
[1808] Step 9:
[1809] Real-time monitoring of vehicle status
[1810] Operation: Sensors collect vehicle status data in real time and send it to the server.
[1811] Input: Vehicle status data.
[1812] Output: Notification that vehicle status data has been successfully sent to the server.
[1813] Specific operation: Sensors inside the vehicle collect status data in real time and send it to the server.
[1814] Step 10:
[1815] Vehicle abnormality detection
[1816] Operation: The server analyzes the received vehicle status data and detects any abnormalities.
[1817] Input: Vehicle status data.
[1818] Output: Anomaly detection notification (e.g., "Oil change required").
[1819] Specific operation: The analysis engine within the server analyzes the status data and detects anomalies.
[1820] Step 11:
[1821] Generating solutions using generative models
[1822] Operation: The server inputs anomaly detection data into a generation model and generates appropriate countermeasures.
[1823] Input: Anomaly detection data.
[1824] Output: Troubleshooting steps (e.g., "Oil change is needed. Please follow these steps.").
[1825] Specific operation: The server inputs anomaly detection data into the generative model and generates specific countermeasures for the user.
[1826] Step 12:
[1827] Sending the troubleshooting steps to the device
[1828] Operation: The server sends the generated troubleshooting data to the terminal.
[1829] Input: Data on how to handle the situation.
[1830] Output: Notification that the data regarding the solution has been successfully sent to the terminal.
[1831] Specific operation: The generated troubleshooting data is sent to the car navigation terminal using a communication module on the server.
[1832] Step 13:
[1833] Providing users with solutions
[1834] Operation: The device provides the user with received troubleshooting data via voice and screen display.
[1835] Input: Data on how to handle the situation.
[1836] Output: Audio output and screen display (e.g., "Oil change is needed. Please follow the next steps.").
[1837] Specific operation: The device uses a speech synthesis engine to notify the user of the solution via voice, and simultaneously displays it on the screen.
[1838] Step 14:
[1839] User Action
[1840] Operation: The system will address the vehicle malfunction according to the troubleshooting steps provided to the user.
[1841] Input: Display of troubleshooting methods and voice guidance.
[1842] Output: Normal vehicle condition.
[1843] Specific actions: The user follows the provided instructions and performs the actual troubleshooting steps.
[1844] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1845] This invention combines an emotion engine with a car navigation system that interprets voice input from the user using a generative model and generates an appropriate response. The system consists of a car navigation terminal, a server, a generative model, vehicle sensors, and an emotion engine that recognizes the user's emotions. This improves the user's driving experience and enhances safety.
[1846] System configuration and operation
[1847] This system consists of a car navigation terminal, a server on which the generative model runs, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[1848] Car navigation terminal
[1849] A car navigation terminal is installed in a vehicle and handles voice command input, output of responses using a generative model and emotion engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to a server.
[1850] Generative model
[1851] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is returned to the car navigation terminal via the server and provided to the user.
[1852] Sensors and Anomaly Detection
[1853] Sensors within the vehicle collect vehicle status data in real time and transmit it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1854] Emotional Engine
[1855] The emotion engine recognizes the user's emotions from their voice input and the in-car environment. The emotions recognized by the emotion engine are then used to customize the generative model's responses and navigation information. In particular, if signs of stress or fatigue are detected, it includes features that recommend relaxation spots and rest areas.
[1856] User Interface
[1857] The user interface consists of the car navigation terminal's screen and voice guidance, and is designed for easy operation. Settings include language, voice guidance type, and theme color, which can be customized to the user's preferences.
[1858] Explanation of the program's processing
[1859] Here, the system's operation is explained in natural language, describing the program's processing.
[1860] 1. User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1861] 2. Terminal: Activates the speech recognition engine to convert the user's voice commands from speech to text. This text data is then sent to the server.
[1862] 3. Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[1863] 4. Generative Model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1864] 5. Server: Sends the generated response data to the car navigation terminal.
[1865] 6. Terminal: Provides the user with the received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[1866] 7. User: Choose from the recommended options. Example: "I want to go to spot (1)."
[1867] 8. Terminal: Calculates the selected route and starts guidance. Updates route information in real time.
[1868] 9. Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[1869] 10. Server: Inputs data from the emotion engine into the generative model to generate a customized response based on emotion. Example: "Would you recommend a relaxation spot?"
[1870] 11. Sensors: Monitor and collect various vehicle status data in real time.
[1871] 12. Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1872] 13. Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[1873] 14. Server: Sends response data, including the solution created by the generative model, to the car navigation terminal.
[1874] 15. Device: Notify the user of how to resolve the issue via voice and on-screen display. Provide specific instructions.
[1875] 16. User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the user will be guided to the nearest gas station or repair shop.
[1876] Through the configuration and processing described above, this system realizes the interactive navigation experience desired by users, and further enables quick and specific responses in the event of vehicle malfunctions, thereby enhancing safety and comfort. By combining it with an emotional engine, flexible responses tailored to the user's psychological state become possible, further improving the driving experience.
[1877] The following describes the processing flow.
[1878] Step 1:
[1879] User: Enter a voice command. Example: "I want to go for a drive to clear my head."
[1880] Step 2:
[1881] Terminal: Activates the speech recognition engine and converts the user's voice commands into text data. This text data is then sent to the server.
[1882] Step 3:
[1883] Server: Passes the received text data to the generative model. The generative model analyzes the text data and understands the user's intent.
[1884] Step 4:
[1885] Generative model: Generates appropriate responses based on the user's intent. Example: "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~"
[1886] Step 5:
[1887] Server: Sends the generated response data to the car navigation terminal.
[1888] Step 6:
[1889] Terminal: Provides the user with received response data via voice output and screen display. Example: "It's nighttime now, so I recommend a night view. Here are some night view spots: (1) (2) (3) Where would you like me to take you?"
[1890] Step 7:
[1891] User: Select from the suggested options. Example: "I want to go to spot (1)."
[1892] Step 8:
[1893] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[1894] Step 9:
[1895] Emotion Engine: Analyzes user voice input and in-car conditions to recognize the user's emotions. Example: "The user is feeling stressed."
[1896] Step 10:
[1897] Server: Passes emotional data from the emotion engine to the generative model, which then generates an emotionally appropriate response. Example: "Would you recommend a relaxation spot?"
[1898] Step 11:
[1899] Generative model: Generates customized responses based on emotions. Example: "You seem stressed, so I recommend the following relaxation spots: (1)~~ (2)~~ (3)~~"
[1900] Step 12:
[1901] Server: Sends the generated response data to the car navigation terminal.
[1902] Step 13:
[1903] Terminal: Provides the user with received response data via voice output and screen display. Example: "We recommend the following relaxation spots. (1)~~ (2)~~ (3)~~ Where would you like us to take you?"
[1904] Step 14:
[1905] User: Select from the suggested relaxation spots. Example: "I want to go to relaxation spot (2)."
[1906] Step 15:
[1907] Terminal: Calculates a route based on the selected spot and starts navigation. Updates route information in real time.
[1908] Step 16:
[1909] Sensors: Monitor and collect various vehicle status data in real time.
[1910] Step 17:
[1911] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1912] Step 18:
[1913] Generative Model: Generates appropriate countermeasures based on abnormal data and creates specific procedures. Example: "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~"
[1914] Step 19:
[1915] Server: Sends response data containing the solution created by the generative model to the car navigation terminal.
[1916] Step 20:
[1917] Device: Notifies the user of how to resolve the issue via voice and on-screen display. Provides specific instructions.
[1918] Step 21:
[1919] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[1920] This process allows the system to provide an interactive navigation experience while taking the user's emotional state into consideration, and enables quick and specific action in the event of a vehicle malfunction.
[1921] (Example 2)
[1922] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1923] Conventional car navigation systems simply interpret user voice input to provide route guidance, making it difficult to respond in accordance with the user's emotions or the vehicle's malfunction. Furthermore, the lack of customized responses tailored to the user's psychological state meant that stress and fatigue during driving could not be reduced. Additionally, there was a problem with the inability to quickly provide specific solutions in the event of a vehicle malfunction.
[1924] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, means for dynamically updating information displayed on the navigation terminal based on the user's voice instructions, sensors for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensors to the generation model and providing an appropriate countermeasure, an emotion engine for recognizing the user's emotions from the user's voice input and the situation inside the vehicle and reflecting this in the response of the generation model, and means for customizing navigation information based on the user's emotions. This enables flexible responses according to the user's psychological state and quick and specific countermeasures in the event of a vehicle abnormality.
[1925] A "generative model" is a type of artificial intelligence that analyzes input data and generates appropriate responses or predictions based on that data.
[1926] A "navigation terminal" is a device installed inside a vehicle that provides map information and route guidance.
[1927] "Voice instructions" refer to instructions or commands spoken by the user, which serve as input for the navigation system to recognize and process.
[1928] A "sensor" is a device that collects vehicle status data in real time and detects specific conditions or abnormalities.
[1929] An "emotion engine" is software or an algorithm that analyzes the user's emotions from their voice input and the situation inside the vehicle, and reflects that emotional information in the response of a generative model.
[1930] "Handling instructions" refer to specific actions and procedures that the user should take when they detect a problem or abnormality in the vehicle.
[1931] "Customization" refers to modifying the interface and responses based on the user's settings and status to meet individual needs.
[1932] Modes for carrying out the invention
[1933] The car navigation system of the present invention enhances the user's driving experience and improves safety by combining a generative model, a voice recognition engine, an emotion engine, and various sensors. The present invention consists of a navigation terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion engine.
[1934] Navigation terminal
[1935] The navigation terminal is installed in the vehicle, takes user voice commands as input, outputs responses using a generative model and emotion engine, and displays route guidance. When the user inputs a voice command such as "I want to go for a drive to clear my head," the terminal uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the voice into text data and sends it to the server.
[1936] Generative model
[1937] The server passes the text data received from the user to a generative model (e.g., GPT-3) to generate an appropriate response. The generative model analyzes the user's intent and generates a specific and appropriate response. The generated response is sent to the navigation terminal via the server.
[1938] Sensors and Anomaly Detection
[1939] The vehicle is equipped with various sensors, including tire pressure sensors and an engine diagnostic system, to collect vehicle status data in real time. This data is sent to a server, which analyzes it and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps.
[1940] Emotional Engine
[1941] The emotion engine recognizes the user's emotions from their voice input and the environment inside the vehicle. For example, if the user is feeling stressed, the emotion engine detects this and sends the data to the generative model. The generative model then generates a customized response based on the emotion data, such as recommending a relaxation spot to the user.
[1942] User Interface
[1943] The user interface consists of a screen and audio for the navigation terminal and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, which can be customized to the user's preferences.
[1944] Specific example
[1945] When a user inputs a voice command such as "I want to go for a drive to clear my head," the voice recognition engine converts the voice into text and sends it to the server. The server passes this text data to a generative model, which generates a response such as "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~." The generated response is sent to the navigation terminal, where the user confirms the recommended spots on the screen and by voice and selects a destination. The navigation terminal then calculates the selected route and begins guidance.
[1946] Thus, by utilizing a generative AI model, a speech recognition engine, an emotion engine, and sensors, the present invention enables flexible responses that respond to the user's intentions and emotions, realizing a car navigation system that achieves both safety and comfort.
[1947] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1948] Program processing flow
[1949] Step 1:
[1950] User: The user enters a voice command. Specifically, they say, "I want to go for a drive to clear my head."
[1951] Input: User's voice
[1952] Output: Raw data audio file
[1953] Step 2:
[1954] Device: Uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert user voice commands from speech to text.
[1955] Input: Raw audio file
[1956] Output: Text data (Example: "I want to go for a drive to clear my head")
[1957] Step 3:
[1958] Terminal: Sends text data to the server.
[1959] Input: Text data
[1960] Output: Sending text data to the server
[1961] Step 4:
[1962] Server: Passes the received text data to a generative model (e.g., GPT-3) and starts the analysis.
[1963] Input: Text data
[1964] Output: User intent (e.g., "Recommendations for destinations suitable for a change of pace")
[1965] Step 5:
[1966] Generative model: Generates appropriate responses based on the user's intent. For example, it might create a response like, "It's nighttime now, so I recommend night views. Here are some night view spots: (1)~~ (2)~~ (3)~~~".
[1967] Input: User intent
[1968] Output: Response text (Example: "Since it's nighttime now, I recommend the night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~")
[1969] Step 6:
[1970] Server: Sends the generated response text data to the navigation terminal.
[1971] Input: Response text data
[1972] Output: Sending response data to the navigation terminal
[1973] Step 7:
[1974] Terminal: Provides the user with the received response text via voice output (e.g., Amazon Polly) and screen display. Specifically, it would say, "It's nighttime now, so I recommend a night view. Here are some night view spots: (1)~~ (2)~~ (3)~~~ Where would you like me to take you?"
[1975] Input: Response text data
[1976] Output: Voice guidance and screen display
[1977] Step 8:
[1978] User: The user selects from the provided options and responds, "I want to go to spot (1)."
[1979] Input: User Selection
[1980] Output: Selection (Example: "(1) Spot")
[1981] Step 9:
[1982] Terminal: Calculates the selected route using the car navigation system's route calculation engine (e.g., Google Maps API) and begins guidance. Route information is updated in real time.
[1983] Input: Selection
[1984] Output: Route guidance and real-time updates
[1985] Step 10:
[1986] Emotion Engine: Analyzes user voice input and in-car environment to recognize user emotions. Specifically, it determines whether the user is experiencing stress.
[1987] Input: User voice, in-vehicle status data
[1988] Output: Recognized emotion (e.g., "I am feeling stressed")
[1989] Step 11:
[1990] Server: Inputs data from the emotion engine into the generative model to generate customized responses based on the user's emotions. For example, it might create a response such as, "Would you recommend a relaxation spot?"
[1991] Input: Sentiment data
[1992] Output: Response text data (e.g., "Do you recommend any relaxation spots?")
[1993] Step 12:
[1994] Sensors: Collect various status data (tire pressure, engine diagnostics, etc.) in real time.
[1995] Input: Vehicle condition
[1996] Output: Status data
[1997] Step 13:
[1998] Server: Continuously receives data from sensors and, if an anomaly is detected, sends that information to the generative model.
[1999] Input: Status data from the sensor
[2000] Output: Abnormal data and notifications
[2001] Step 14:
[2002] Generative Model: Based on abnormal data, it generates appropriate countermeasures and creates instructions such as, "The oil change light is on. Please take the following steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[2003] Input: Abnormal data
[2004] Output: Text data of the solution.
[2005] Step 15:
[2006] Server: Sends response data containing the solution generated by the generative model to the navigation terminal.
[2007] Input: Text data of the solution
[2008] Output: Send to navigation terminal
[2009] Step 16:
[2010] Terminal: Notifies the user of how to resolve the issue via voice and on-screen display, and guides them through specific steps. For example, it might say, "The oil change light is on. Please follow these steps: (1) Check the oil (2) Go to the nearest gas station (3) Change the oil."
[2011] Input: Text data of the solution
[2012] Output: Voice guidance and screen display
[2013] Step 17:
[2014] User: Follow the provided instructions to address any vehicle malfunctions. If necessary, the system will search for and navigate you to the nearest gas station or repair shop.
[2015] Input: Instructions on how to proceed
[2016] Output: Appropriate countermeasures
[2017] (Application Example 2)
[2018] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2019] In recent years, the demand for food delivery services has increased, requiring drivers to perform delivery duties efficiently and safely. However, excessive stress and fatigue during driving can reduce drivers' attention span and increase the risk of accidents. Furthermore, current car navigation systems cannot provide responses that take into account the driver's emotional state, and improvements are needed to enhance the driver's driving experience. Therefore, a system is needed that analyzes the driver's emotional state in real time and generates appropriate responses.
[2020] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for interpreting voice input from the user using a generation model and generating an appropriate response, an in-vehicle information terminal, means for dynamically updating the information displayed on the in-vehicle information terminal based on the user's voice instructions, a sensor for collecting vehicle status data in real time and detecting abnormalities, means for transmitting abnormal data collected from the sensor to the generation model and providing an appropriate countermeasure, emotion analysis means for analyzing the user's emotional state and generating a response corresponding to that emotion, and means for suggesting rest locations such as relaxation spots according to the user's emotional state recognized by the emotion analysis means. This makes it possible to detect the stress and fatigue of the driver and provide appropriate responses and rest suggestions, thereby enabling safe and efficient delivery operations.
[2021] A "generative model" is an artificial intelligence model that interprets voice input from a user and generates an appropriate response.
[2022] An "in-vehicle information terminal" is a device installed in a vehicle that allows for the input of voice commands, the output of responses from a generative model and sentiment analysis engine, and the display of route guidance.
[2023] A "dynamically updating method" refers to a method of changing the information displayed on the in-vehicle information terminal in real time based on the user's voice commands.
[2024] A "sensor" is a device that collects vehicle status data in real time and detects abnormalities.
[2025] "Means of providing appropriate countermeasures" refers to a method of transmitting anomaly data collected from sensors to a generation model and then presenting specific countermeasures based on that data.
[2026] "Emotion analysis means" refers to a method that recognizes the emotional state from the user's voice input or the situation inside the vehicle and provides that information to a generative model.
[2027] A "relaxation spot" is a place where users can rest and refresh themselves when they feel stressed or tired.
[2028] This invention is a system that supports food delivery drivers in performing their duties safely and efficiently, and consists of the following components.
[2029] System configuration
[2030] This system consists of an in-vehicle information terminal, a server on which the generative model operates, sensors that monitor the vehicle's status, a user interface, and an emotion analysis engine.
[2031] In-vehicle information terminal
[2032] The in-vehicle information terminal is installed in the vehicle and handles voice command input, output of responses generated by a generative model and sentiment analysis engine, and display of route guidance. When a user inputs a voice command, the terminal converts the voice into text data and sends it to the server. For example, if a driver voice-inputs "I'm tired, I want to take a short break," the terminal converts that voice into text and sends it to the server.
[2033] Generative model
[2034] The server passes the text data received from the user to a generative model, which generates an appropriate response. The generative model analyzes the user's intent and emotional state to generate a specific and appropriate response. The generated response is returned to the in-vehicle information terminal via the server and provided to the user. For example, in response to input such as "I'm tired, I want to take a break," a response like "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?" is generated.
[2035] Sensors and Anomaly Detection
[2036] Sensors within the vehicle collect vehicle status data in real time and send it to a server. The server analyzes this data and, if an anomaly is detected, sends the information to a generative model. The generative model generates appropriate countermeasures according to the type of anomaly and guides the user through the specific steps. For example, if it detects that an oil change is needed, it will provide countermeasures such as, "The oil change lamp is lit. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[2037] Emotion analysis engine
[2038] The emotion analysis engine recognizes the user's emotional state from their voice input and the environment inside the vehicle. The user's emotional state recognized by the emotion analysis engine is reflected in the generative model's responses and the customization of navigation information. In particular, if the system detects that the user is feeling "tired" or "stressed," it includes a function that suggests relaxation spots and rest areas. For example, if the user inputs "tired," the system detects fatigue and suggests nearby relaxation spots.
[2039] User Interface
[2040] The user interface consists of the in-car information terminal's screen and voice, and is designed for easy operation. Settings include language, voice type for voice guidance, and theme color, all of which can be customized to the user's preferences. For example, simply saying "I want to change the voice guidance voice" will change it to the preferred voice type.
[2041] Specific example
[2042] 1. User input: "I'm tired, I want to take a break."
[2043] 2. Example of a prompt:
[2044] Emotion: Fatigue, Input: I'm tired and want to take a break. Can you suggest any relaxation spots?
[2045] 3. Response generated: "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[2046] This configuration allows for the detection of driver stress and fatigue, and enables safe and efficient delivery operations by providing appropriate responses and rest suggestions.
[2047] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2048] Step 1:
[2049] The user inputs a voice command. For example, they might say, "I'm tired, I want to take a break." This voice input becomes the starting point for the system's processing.
[2050] Step 2:
[2051] The device activates a speech recognition engine to convert the user's voice commands from speech to text. The input is the user's voice data, and the output is the text data of that voice. The device sends this text data to the server.
[2052] Step 3:
[2053] The server passes the received text data to the generative model. The input is text data sent from the terminal, and the generative model is used to analyze the user's intent and emotions. The output is data containing the analysis results.
[2054] Step 4:
[2055] The generative model generates appropriate responses based on the user's intentions and emotions. The input is the parsed data obtained in the previous step, and the output is the response sentence presented to the user. For example, it might generate a response such as, "There are three nearby relaxation spots: Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[2056] Step 5:
[2057] The server transmits the generated response data to the in-vehicle infotainment terminal. The input is the response data from the generative model, and the output is the communication data to the in-vehicle infotainment terminal.
[2058] Step 6:
[2059] The terminal provides the user with the received response data through voice output and screen display. The input is the response data from the server, and the output is voice guidance and screen display to the user. For example, it might provide voice guidance such as, "Nearby relaxation spots include Park A, Cafe B, and Hot Spring C. Which one would you like to go to?"
[2060] Step 7:
[2061] The user selects an option from the suggested relaxation spots and gives a voice command. For example, they might respond, "I want to go to Cafe B." This voice data becomes the input for the next process.
[2062] Step 8:
[2063] The device restarts its speech recognition engine and converts the user's selection from speech to text. The input is the user's voice data, and the output is the text data of the selections. The device sends this text data to the server.
[2064] Step 9:
[2065] The server calculates the route to the selected relaxation spot based on the received text data. The input is the text data of the selected spot, and the output is the data of the optimal route.
[2066] Step 10:
[2067] The server transmits the calculated route to the in-vehicle information terminal. The input is the route calculation result data, and the output is the communication data to the in-vehicle information terminal.
[2068] Step 11:
[2069] The terminal starts navigation based on the received route data and updates the route information in real time. The input is route data from the server, and the output is navigation guidance and screen display.
[2070] Step 12:
[2071] The emotion analysis engine analyzes the user's voice input and the conditions inside the vehicle to recognize the user's emotions. The input is the user's voice data and sensor information, and the output is emotional state data.
[2072] Step 13:
[2073] The server inputs data obtained from the emotion analysis engine into a generative model to generate a customized response based on the emotion. The input is emotional state data, and the output is customized response data. For example, it can generate a response such as, "Would you recommend a relaxation spot?"
[2074] Step 14:
[2075] The sensors monitor and collect various vehicle status data in real time. The input is vehicle status information, and the output is data signals from the sensors.
[2076] Step 15:
[2077] The server continuously receives data from the sensor and, if it detects an anomaly, sends that information to the generative model. The input is the data signal from the sensor, and the output is the anomaly data sent to the generative model.
[2078] Step 16:
[2079] The generative model generates appropriate countermeasures based on abnormal data and creates specific procedures. The input is abnormal data, and the output is procedure data for the countermeasures. For example, it generates procedures such as "The oil change light is on. Please take the following steps: (1) ~~ (2) ~~ (3) ~~".
[2080] Step 17:
[2081] The server sends response data, including the countermeasures created by the generative model, to the in-vehicle information terminal. The input is the procedural data of the countermeasures, and the output is the communication data to the in-vehicle information terminal.
[2082] Step 18:
[2083] The device notifies the user of troubleshooting methods via voice and on-screen display, guiding them through specific steps. Input is troubleshooting data from the server, and output is voice guidance and on-screen display to the user.
[2084] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2085] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2086] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[2087] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2088] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[2089] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[2090] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[2091] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[2092] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[2093] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2094] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2095] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2096] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[2097] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2098] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[2099] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[2100] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[2101] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[2102] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, ...
Claims
1. A model that uses a generative model to interpret voice input from the user and generate an appropriate response, Car navigation terminal and Means for dynamically updating the information displayed on the car navigation terminal based on the user's voice commands, A sensor that collects vehicle status data in real time and detects abnormalities, A means for transmitting abnormal data collected from the aforementioned sensor to a generation model and providing an appropriate countermeasure, A system that includes this.
2. The system according to claim 1, further comprising means for appropriately setting and updating the vehicle's navigation route in real time.
3. The system according to claim 1, further comprising means for providing a customizable interface based on user settings.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A