System
A voice-activated navigation system processes voice commands for real-time information on destinations and services, addressing safety and convenience issues in conventional navigation systems by allowing safe and efficient route adjustments.
Patent Information
- Application Number
- JP2024122701
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Conventional car navigation systems and smartphone apps require significant time for setting destinations and stopovers and are unsafe to operate while driving, lacking real-time information on tourist spots, restaurants, and hotels, compromising safety and comfort.
A voice-activated car navigation system that uses artificial intelligence to process voice commands, providing real-time information on destinations, tourist spots, restaurants, and hotels, allowing for safe and easy adjustments to routes and schedules while driving.
Enables safe and efficient destination setting and schedule adjustments during travel, offering comprehensive information and reservations through voice commands, enhancing user safety and comfort.
Smart Images

Figure 2026021019000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When traveling by car, using conventional car navigation systems or smartphone apps has the problem of requiring a lot of time to set destinations and stopovers in advance. Furthermore, because it is dangerous to operate a smartphone or navigation system while driving, it is difficult to quickly respond to sudden weather changes or unexpected schedule changes. Furthermore, the lack of accurate information on tourist spots, restaurants, hotels, and other information near the destination compromises the comfort and safety of the trip. There is a need for a system that can solve these issues and enable safe and easy route adjustments and information provision while driving. [Means for solving the problem]
[0005] The present invention provides a means for activating a voice input engine and receiving a user's voice command, followed by a processing means using an artificial intelligence model to convert the voice data into text. It then includes a means for extracting a destination and desired arrival time from the converted text data, transmitting this information to a server and receiving an estimated arrival time. It also includes a means for providing audio feedback on the estimated arrival time to the user. The system includes a means for searching for nearby tourist destination candidates based on current location and route information, acquiring and transmitting popularity, word-of-mouth reviews, photos, and parking information. It is possible to audio-notify the user of the received information, receive feedback, recalculate the route to the selected tourist destination, and adjust the estimated arrival time. It also includes a means for searching for nearby restaurant candidates based on current location and route, acquiring word-of-mouth reviews, budget, and parking information, audio-notify the user of the received information, and, if necessary, make a reservation call to receive feedback. It also includes a means for recalculating the route to the selected restaurant. It also includes a means for searching for nearby accommodation candidates based on current location and destination, acquiring word-of-mouth reviews, popularity, budget, and parking information, audio-notify the user of the received information, receive feedback, and proceed with the reservation process as needed. The route to the selected accommodation can be recalculated and the estimated arrival time adjusted, allowing users to plan and adjust their journey safely and comfortably, all by voice.
[0006] A "voice input engine" is a software or hardware device that converts voice into a digital signal, analyzes the signal, and converts it into text data.
[0007] "User voice command" refers to a voice input by a user to give instructions to the system through voice.
[0008] An "artificial intelligence model" is an algorithm or neural network technology used to learn from large amounts of data and perform complex tasks.
[0009] "Converting to text" refers to the process of analyzing voice data and converting it into text data as character information.
[0010] "Destination" refers to the point or location where the user ultimately wishes to arrive.
[0011] "Desired arrival time" is the time when a user desires to arrive at a particular location.
[0012] A "server" is a computer system used on a network to store and process data.
[0013] "Estimated arrival time" refers to the estimated time it will take to arrive at the specified destination.
[0014] "Voice feedback" refers to the process in which the system responds to the user with information by voice.
[0015] "Current location" refers to the place or location where the user is currently located.
[0016] "Route information" is information about the route or path from the departure point to the destination.
[0017] "Candidate tourist destinations" is a list of candidate tourist destinations that the user may visit.
[0018] "Popularity" is an indicator of how many people support a particular place or item.
[0019] "Word of mouth reputation" is information that represents the evaluations and opinions of other users.
[0020] A "photograph" is a still image that depicts a visual image of a particular place or object.
[0021] "Parking lot information" is information about parking of vehicles at a particular location.
[0022] "Feedback" refers to responses or reactions from users.
[0023] "Recalculating the route" refers to the process of calculating a new route based on existing route information.
[0024] "Restaurant candidates" is a list of potential dining locations that the user may visit.
[0025] A "budget" is the amount a user plans to spend on a particular service or product.
[0026] "Making a reservation call" refers to a process of making a reservation for a specific destination by voice call based on instructions from the user.
[0027] "Potential accommodation" is a list of potential accommodations where the user may stay.
[0028] "Reservation Process" means the process by which a particular destination or service is reserved in advance.
[0029] "Estimated time of arrival" is the expected time to arrive at a specified location. [Brief explanation of the drawings]
[0030] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0031] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0032] First, the terms used in the following description will be explained.
[0033] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0034] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0035] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0036] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0037] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0038] [First embodiment]
[0039] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0040] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0041] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0042] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0043] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0044] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0045] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0046] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0047] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0048] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0049] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0050] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0051] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time.
[0052] Functionality Overview
[0053] 1. Pre-departure settings
[0054] Device:
[0055] It starts the voice input engine and receives the user's voice commands.
[0056] The voice data is sent to an artificial intelligence model and converted into text.
[0057] Information about the final destination and desired arrival time set by the user is sent to the server.
[0058] The destination information returned from the server is checked and audio feedback is given to the user.
[0059] 2. Tourist information while driving
[0060] server:
[0061] Potential tourist spots in the area are selected based on the user's route and current location.
[0062] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[0063] This information is sent to the terminal.
[0064] Device:
[0065] Tourist information from the server is announced to the user by voice.
[0066] It receives user feedback and sends it to the server.
[0067] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[0068] 3. Restaurant directions while driving
[0069] server:
[0070] Select nearby restaurant options based on the user's current location and route.
[0071] Obtain information such as restaurant reviews, budget, and whether parking is available.
[0072] If necessary, the reservation telephone number is transmitted to the terminal.
[0073] Device:
[0074] Restaurant information sent from the server is announced to the user by voice.
[0075] Receive user feedback and send selections to the server.
[0076] Make a reservation if necessary.
[0077] It provides the functionality to receive user feedback, send it to the server, and make reservation calls.
[0078] 4. Hotel reservation information while driving
[0079] server:
[0080] Select nearby hotel options based on the user's current location and destination.
[0081] Obtain hotel reviews, popularity, budget, parking information, etc.
[0082] If necessary, information for reservation procedures is sent to the terminal.
[0083] Device:
[0084] Hotel information sent from the server is announced to the user by voice.
[0085] Receive user feedback and send the selection results to the server.
[0086] Proceed with the booking process as needed.
[0087] Once the booking is completed, the user is given audio feedback.
[0088] Recalculate the route to the user's selected hotel and adjust the estimated arrival time.
[0089] Specific examples
[0090] 1. Pre-departure settings
[0091] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[0092] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[0093] Server: Calculates the estimated arrival time and returns it to the device.
[0094] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[0095] 2. Tourist information
[0096] User: "What are some nearby tourist spots?"
[0097] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[0098] User: "I'll be there."
[0099] Terminal: "Route set to museum. Arrival time 15 minutes."
[0100] 3. Restaurant Information
[0101] User: "I'm hungry, can you find a restaurant nearby?"
[0102] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[0103] User: "Yes, I'd like to make a reservation."
[0104] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[0105] 4. Hotel Reservation Information
[0106] User: "Find a hotel tonight."
[0107] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[0108] User: "Yes, please make a reservation."
[0109] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[0110] In this way, the system of the present invention allows all operations to be performed by voice, allowing you to safely set your destination or change your schedule while driving, providing a comfortable travel experience.
[0111] The processing flow will be explained below.
[0112] 1. Pre-departure settings
[0113] Step 1:
[0114] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[0115] Step 2:
[0116] The device activates the voice input engine and receives the user's voice command.
[0117] Step 3:
[0118] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0119] Step 4:
[0120] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[0121] Step 5:
[0122] The terminal transmits the extracted information to the server.
[0123] Step 6:
[0124] The server calculates the estimated arrival time and returns the result to the terminal.
[0125] Step 7:
[0126] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[0127] 2. Tourist information while driving
[0128] Step 1:
[0129] The user issues a voice command: "Tell me about nearby tourist attractions."
[0130] Step 2:
[0131] The device activates the voice input engine and receives the user's voice command.
[0132] Step 3:
[0133] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0134] Step 4:
[0135] The terminal transmits the converted text data to the server.
[0136] Step 5:
[0137] The server selects nearby tourist spots based on the user's route and current location.
[0138] Step 6:
[0139] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[0140] Step 7:
[0141] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[0142] Step 8:
[0143] The user gives verbal feedback: "I'm going there."
[0144] Step 9:
[0145] The device sends the user's feedback to the server.
[0146] Step 10:
[0147] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[0148] Step 11:
[0149] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[0150] 3. Restaurant directions while driving
[0151] Step 1:
[0152] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[0153] Step 2:
[0154] The device activates the voice input engine and receives the user's voice command.
[0155] Step 3:
[0156] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0157] Step 4:
[0158] The terminal transmits the converted text data to the server.
[0159] Step 5:
[0160] The server searches for nearby restaurant options based on the user's current location and route.
[0161] Step 6:
[0162] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[0163] Step 7:
[0164] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[0165] Step 8:
[0166] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[0167] Step 9:
[0168] The terminal sends the user's feedback to the server and activates the reservation call function.
[0169] Step 10:
[0170] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[0171] Step 11:
[0172] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[0173] 4. Hotel reservation information while driving
[0174] Step 1:
[0175] A user issues a voice command: "Find a hotel for tonight."
[0176] Step 2:
[0177] The device activates the voice input engine and receives the user's voice command.
[0178] Step 3:
[0179] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0180] Step 4:
[0181] The terminal transmits the converted text data to the server.
[0182] Step 5:
[0183] The server searches for nearby hotel options based on the user's current location and destination.
[0184] Step 6:
[0185] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[0186] Step 7:
[0187] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[0188] Step 8:
[0189] The user gives verbal feedback: "Yes, please book."
[0190] Step 9:
[0191] The device sends the user's feedback to the server and proceeds with the reservation process.
[0192] Step 10:
[0193] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[0194] Step 11:
[0195] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[0196] Example 1
[0197] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0198] Current car navigation systems have the challenge of making it difficult for users to safely and efficiently set destinations and adjust schedules while driving. Furthermore, they lack the functionality to provide comprehensive, real-time information on tourist spots, restaurants, hotels, and other information, preventing users from making appropriate choices quickly. This often results in a loss of convenience and safety while driving.
[0199] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0200] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for suggesting tourist spots, restaurants, and hotels based on the user's current location and destination and acquiring information, means for notifying the user of the acquired information audio-visually and receiving feedback, means for recalculating a route to the selected destination and adjusting the estimated arrival time, and means for making reservations if necessary. This allows destination setting and schedule adjustments to be performed safely and efficiently even while driving, providing a comfortable travel experience.
[0201] A "voice input engine" is a device or software that receives a user's voice commands and converts them into digital data.
[0202] "Voice Data" means human speech information converted into digital form by a speech input engine.
[0203] "Artificial intelligence model" refers to technologies such as machine learning algorithms and neural networks used to convert voice data into text data.
[0204] "Text data" is character string information converted from voice data by an artificial intelligence model.
[0205] "Destination" is information indicating the place or location to which the user wishes to travel.
[0206] The "desired arrival time" is information indicating a specific time at which the user desires to arrive at the destination.
[0207] A "server" is a central system that receives requests on a computer network, processes data, and sends and receives information.
[0208] "Estimated arrival time" is information indicating the estimated arrival time from the current location to the destination.
[0209] "Tourist destination" refers to tourist spots and famous places that users aim to visit.
[0210] "Restaurant" means an eating and drinking establishment selected by a patron for dining.
[0211] "Hotel" means the accommodation facility selected for the Guest's stay.
[0212] "Means of acquisition" refers to the methods and technologies used to collect and receive the required information or data.
[0213] "Means of receiving feedback" refers to the methods and techniques for receiving responses or reactions from users.
[0214] "Means for recalculating a route" refers to a method or technology for recalculating a new route based on specified conditions.
[0215] "Means for completing reservation procedures" refers to the methods and technologies used to make reservations for the service selected by the user (such as restaurant reservations or hotel reservations).
[0216] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, adjusting routes and arrival times in real time.
[0217] The main components of the system are the "terminals" used by users and the "servers" that process data. The specific configuration and operation are explained below.
[0218] The system's devices, which include smartphones and car navigation systems, are equipped with a voice input engine that receives users' voice commands and converts the voice data into a digital format. This digital voice data is then converted into text data using voice recognition software such as Google Cloud Speech-to-Text.
[0219] The converted text data is sent to the server and used to extract the destination and desired arrival time. The server calculates the estimated arrival time based on this information and returns the result to the terminal. The terminal then provides this information as feedback to the user via voice. Specifically, it notifies the user by saying something like, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[0220] In the case of tourist attraction guidance while driving, the device sends the user's current location and route information to the server, and the server searches for nearby tourist attractions. The server obtains the tourist attraction's popularity, reviews, photos, and parking information and sends them to the device. The device then announces this information to the user by voice (e.g., "Yes, there's a famous museum nearby. It has a 4.5 rating and there's parking available. What would you like to do?"). Based on the user's feedback, the server recalculates the route and calculates a new estimated arrival time.
[0221] Regarding restaurant guidance, the system searches for nearby restaurants based on the current location and route, and obtains information such as reviews, budget, and parking information. If necessary, it obtains a reservation phone number and sends it to the terminal. The terminal then announces this information to the user by voice (e.g., "There's a highly rated Chinese restaurant nearby. The reviews are 4.7, and your budget is around 1,500 yen for lunch. Would you like to make a reservation?") and proceeds with the reservation process based on the user's feedback.
[0222] Additionally, the hotel guide searches for nearby hotels based on the current location and destination, and obtains information such as reviews, popularity, budget, and parking information. If a reservation is required, the server sends the necessary information to the terminal, which then proceeds with the reservation process (e.g., "Looking for a hotel to stay at tonight," "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?").
[0223] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set their destination or change their schedule while driving, providing a comfortable travel experience.
[0224] Specific prompt examples
[0225] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[0226] "Tell me about nearby tourist spots."
[0227] I'm hungry, so I'm looking for a nearby restaurant.
[0228] "Find a hotel to stay at tonight."
[0229] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0230] The flow of this system's program processing
[0231] 1. Pre-departure settings
[0232] Step 1:
[0233] Subject: User
[0234] Description: The user speaks a voice command into the device.
[0235] Specific actions: For example, say, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 PM."
[0236] Step 2:
[0237] Subject: Device
[0238] Description: Starts the voice input engine and receives voice commands.
[0239] Input: Voice command from the user
[0240] Output: Digital audio data
[0241] What it does: The voice input engine receives the voice and converts it into a digital format.
[0242] Step 3:
[0243] Subject: Device
[0244] Description: Uses artificial intelligence models to convert voice data into text.
[0245] Input: Digital audio data
[0246] Output: Text data
[0247] What it does: Converts speech to text using a service like Google Cloud Speech-to-Text.
[0248] Step 4:
[0249] Subject: Device
[0250] Description: Sends the converted text data to the server.
[0251] Input: Text data (e.g., "Set a route to Tokyo Station, with a desired arrival time of 3:00 PM.")
[0252] Output: Request to server
[0253] Specific operation: Sends an HTTP request to the server.
[0254] Step 5:
[0255] Subject: Server
[0256] Description: Extracts the destination and desired arrival time and calculates the estimated arrival time.
[0257] Input: Text data
[0258] Output: Estimated arrival time
[0259] Specific operation: Uses natural language processing technology to analyze text and calculate arrival times by referencing traffic information and road conditions.
[0260] Step 6:
[0261] Subject: Server
[0262] Description: Sends the calculation result back to the terminal.
[0263] Input: Estimated arrival time
[0264] Output: Feedback data
[0265] Specific operation: Returns the calculation result to the terminal.
[0266] Step 7:
[0267] Subject: Device
[0268] Description: Provides audio feedback data from the server to the user.
[0269] Input: Feedback data
[0270] Output: Audio notification to the user
[0271] Specific operation: The text data is converted into speech, and a message such as "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM" is spoken to the user.
[0272] 2. Tourist information while driving
[0273] Step 1:
[0274] Subject: User
[0275] Description: Request "Tell me about nearby tourist spots."
[0276] Specific action: Speak a voice command into the device.
[0277] Step 2:
[0278] Subject: Device
[0279] Description: Converts voice commands into text and sends it to the server.
[0280] Input: Voice command
[0281] Output: Text data
[0282] Specific operation: The audio is converted into text using Google Cloud Speech-to-Text or similar and sent to the server.
[0283] Step 3:
[0284] Subject: Server
[0285] Description: Searches for nearby tourist spots based on the user's current location and route.
[0286] Input: Current location and route information
[0287] Output: List of tourist destination candidates
[0288] Specific behavior: Generate a list of tourist destinations by retrieving information from a database or external API.
[0289] Step 4:
[0290] Subject: Server
[0291] Description: Get tourist attraction popularity, reviews, photos, and parking information.
[0292] Input: Tourist destination candidate list
[0293] Output: Detailed information
[0294] Specific actions: Collect and list detailed information about each tourist destination.
[0295] Step 5:
[0296] Subject: Server
[0297] Description: Sends the acquired information to the device.
[0298] Input: More information
[0299] Output: Feedback data
[0300] Specific operation: Sends collected information to the device.
[0301] Step 6:
[0302] Subject: Device
[0303] Description: Presents information to the user audibly.
[0304] Input: Feedback data
[0305] Output: Audio notification to the user
[0306] Action: The audio message is, "There's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[0307] Step 7:
[0308] Subject: User
[0309] Description: Gives feedback saying "There you go."
[0310] Specific actions: Select a tourist spot specified by voice.
[0311] Step 8:
[0312] Subject: Device
[0313] Description: Recalculates route based on user selection.
[0314] Input: User's choice
[0315] Output: Recalculated route
[0316] Specific behavior: Calculate a new route and send a notification such as "Route to the museum has been set. It will take 15 minutes to arrive."
[0317] 3. Restaurant directions while driving
[0318] Step 1:
[0319] Subject: User
[0320] Description: "I'm hungry, find me a nearby restaurant."
[0321] Specific action: Speak a voice command into the device.
[0322] Step 2:
[0323] Subject: Device
[0324] Description: Sends a voice command to the server.
[0325] Input: Voice command
[0326] Output: Text data
[0327] Specific operation: Converts speech into text and sends it to the server.
[0328] Step 3:
[0329] Subject: Server
[0330] Description: Find nearby restaurant suggestions based on your current location and route.
[0331] Input: Current location and route information
[0332] Output: Restaurant candidate list
[0333] Specific operation: Retrieves restaurant information from a database or external API and creates a list.
[0334] Step 4:
[0335] Subject: Server
[0336] Description: Get detailed information like reviews, budget, parking info, etc.
[0337] Input: Restaurant candidate list
[0338] Output: Detailed information
[0339] What it does: Collect and list detailed information about each restaurant.
[0340] Step 5:
[0341] Subject: Server
[0342] Description: Sends the acquired information to the device.
[0343] Input: More information
[0344] Output: Feedback data
[0345] Specific operation: Sends collected information to the device.
[0346] Step 6:
[0347] Subject: Device
[0348] Description: Presents information to the user audibly.
[0349] Input: Feedback data
[0350] Output: Audio notification to the user
[0351] Specific action: Say, "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[0352] Step 7:
[0353] Subject: User
[0354] Instructions: Answer "Yes, I'd like to make a reservation as well."
[0355] Specific action: Indicate your intention to make a reservation by voice.
[0356] Step 8:
[0357] Subject: Device
[0358] Description: Proceed with the booking process.
[0359] Input: User's booking request
[0360] Output: Reservation completion notification
[0361] Specific operation: Make a reservation using the OpenTable API or similar, and notify the customer, "Your reservation has been arranged. You will arrive in 10 minutes."
[0362] 4. Hotel reservation information while driving
[0363] Step 1:
[0364] Subject: User
[0365] Description: "Find me a hotel tonight."
[0366] Specific action: Speak a voice command into the device.
[0367] Step 2:
[0368] Subject: Device
[0369] Description: Sends a voice command to the server.
[0370] Input: Voice command
[0371] Output: Text data
[0372] Specific operation: Converts voice commands into text and sends it to the server.
[0373] Step 3:
[0374] Subject: Server
[0375] Description: Search for nearby hotel suggestions based on your current location and destination.
[0376] Input: Current location and destination information
[0377] Output: Hotel candidate list
[0378] Specific operation: Retrieves hotel information from a database or external API and creates a list.
[0379] Step 4:
[0380] Subject: Server
[0381] Description: Get detailed information like reviews, popularity, budget, parking information, and more.
[0382] Input: Hotel candidate list
[0383] Output: Detailed information
[0384] What to do: Collect and list detailed information about each hotel.
[0385] Step 5:
[0386] Subject: Server
[0387] Description: Sends the acquired information to the device.
[0388] Input: More information
[0389] Output: Feedback data
[0390] Specific operation: Sends collected information to the device.
[0391] Step 6:
[0392] Subject: Device
[0393] Description: Presents information to the user audibly.
[0394] Input: Feedback data
[0395] Output: Audio notification to the user
[0396] Action: The voice will say, "There's a highly rated hotel nearby with a 4.8 rating and parking. Would you like to make a reservation?"
[0397] Step 7:
[0398] Subject: User
[0399] Instructions: "Yes, please make a reservation."
[0400] Specific action: Indicate your intention to make a reservation by voice.
[0401] Step 8:
[0402] Subject: Device
[0403] Description: Proceed with the booking process.
[0404] Input: User's booking request
[0405] Output: Reservation completion notification
[0406] Specific behavior: Make a reservation using Expedia API or Booking.com API and notify the user, "Hotel has been booked. Estimated arrival time is 6:00 PM."
[0407] (Application example 1)
[0408] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0409] When traveling by car, there is a demand for voice control to select and reserve destinations, tourist attractions, restaurants, and hotels along the way, and for integration with autonomous driving systems to provide a safe and comfortable travel experience. However, with conventional systems, it can be difficult to provide sufficient information, complete reservations, and adjust routes using voice control alone. Furthermore, there is a lack of integration with autonomous driving systems, and manual operation by the user is often required. This can make operations while driving cumbersome, potentially compromising safety and efficiency.
[0410] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0411] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for adding control means compatible with the autonomous driving system, and means for automating automatic route adjustment to selected facilities and reservation procedures. This makes it possible for an autonomous vehicle to utilize voice operation and AI technology to select and reserve tourist spots, restaurants, and hotels in real time based on user instructions, and to automatically adjust the route and provide feedback on the arrival time.
[0412] A "voice input engine" is a device or software that receives voice commands from a user.
[0413] "Artificial intelligence model" is a machine learning algorithm used to convert voice data into text.
[0414] "Text data" is voice data converted into text format.
[0415] A "server" is a computer system that provides services to clients over a computer network.
[0416] The "destination" is the destination point set by the user.
[0417] "Desired arrival time" is the time the user desires to arrive at the destination.
[0418] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the server.
[0419] An "autonomous driving system" is a technology that enables vehicles to perform driving operations autonomously.
[0420] "Route adjustment" refers to recalculating the route to a selected destination.
[0421] "Tourist destination candidates" is a list of tourist destinations suggested to the user.
[0422] "Popularity" is an indicator that shows the evaluation of tourist destinations and facilities.
[0423] "Word of mouth" refers to reviews and feedback from users.
[0424] "Photos" are image data that provide visual images of tourist spots and facilities.
[0425] "Parking information" is information about parking spaces at tourist spots and facilities.
[0426] "Restaurant candidates" is a list of restaurants suggested to the user.
[0427] The "budget" is an estimate of the cost for the user to use the service.
[0428] "Reservation phone" refers to a means of making a reservation for a facility by telephone.
[0429] In one embodiment of the present invention, the system first activates a voice input engine to receive a user's voice command. The voice input engine may be, for example, the Google Speech Recognition API or a similar voice recognition engine. The device that receives the voice data converts it into text data using an artificial intelligence model (e.g., a generative AI model).
[0430] Natural language processing (NLP) techniques are used to extract the destination and desired arrival time from the converted text data. The text data is analyzed to extract specific information (in this case, the destination and desired arrival time). This information is sent to a server, which calculates the estimated arrival time. The calculation result is sent back to the device, which then communicates the result to the user as voice feedback.
[0431] By adding control means compatible with the autonomous driving system, it becomes possible to automatically adjust routes to selected facilities and make reservations. Specifically, the terminal provides route information to the autonomous driving system based on information received from the server and uses an API for reservation procedures. This allows users to set destinations and make reservations using only voice commands without manual operation.
[0432] For example, if a user voice-inputs "My destination is Tokyo Station, and I would like to arrive at 3:00 PM," the device converts the voice data into text data and extracts the destination and desired arrival time. This information is sent to the server, which calculates the estimated arrival time and receives feedback. Furthermore, if the user issues the command "Find a nearby restaurant," the server searches for restaurant candidates based on the current location and route information, obtains reviews and budget information, and sends it to the device. The device notifies the user of this by voice, and if the user responds "Make a reservation," the device will automatically complete the reservation procedure.
[0433] An example of a prompt sentence might be:
[0434] Please let us know your destination and desired arrival time.
[0435] Could you tell me about nearby tourist spots?
[0436] Find a restaurant near you.
[0437] Find a hotel to stay in tonight.
[0438] In this way, the present invention combines voice control with an automated driving system to provide users with a safe and convenient travel experience.
[0439] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0440] Step 1:
[0441] The device starts a voice input engine to receive the user's voice command. The voice input engine used here is a speech recognition engine such as the Google Speech Recognition API. The voice input engine captures the user's voice data and processes it as digital voice data.
[0442] Input: User's voice command
[0443] Output: Digital audio data
[0444] Step 2:
[0445] The voice data received by the device is converted into text data using a generative AI model (a speech recognition algorithm). This process involves analyzing the voice signal and generating the corresponding text.
[0446] Input: Digital audio data
[0447] Output: Text data
[0448] Step 3:
[0449] The device extracts the destination and desired arrival time from the generated text data, using natural language processing (NLP) techniques to identify keywords and phrases related to the destination and desired arrival time.
[0450] Input: Text data
[0451] Output: Destination and desired arrival time information
[0452] Step 4:
[0453] The device sends the extracted information to a server, which calculates and returns an estimated arrival time. The server then uses a map database and traffic data to calculate a route from the current location to the destination.
[0454] Input: Destination and desired arrival time information
[0455] Output: Estimated arrival time
[0456] Step 5:
[0457] The terminal receives the estimated arrival time from the server and provides the user with audio feedback. Here, a synthetic speech engine (e.g., pyttsx3) is used to generate speech from text and communicate it to the user.
[0458] Input: Estimated arrival time
[0459] Output: Audio feedback
[0460] Step 6:
[0461] The user inputs an additional voice command (e.g., "Tell me about nearby tourist spots," "Find nearby restaurants," etc.). Based on this prompt, the device queries the server.
[0462] Input: User's additional voice command
[0463] Output: prompt statement
[0464] Step 7:
[0465] The server searches for potential tourist spots and restaurants based on the user's current location and route information, and then retrieves and sends the information to the device. The retrieved information includes popularity, reviews, photos, parking information, budget, etc.
[0466] Input: User's current location and route information
[0467] Output: Information on tourist spots and restaurants
[0468] Step 8:
[0469] The terminal announces the information received from the server to the user by voice and receives user feedback. When the user makes a selection, the selection is sent to the server, which then processes the new route and reservation.
[0470] Input: Information about tourist attractions and restaurants
[0471] Output: Audio feedback and user selection results
[0472] Step 9:
[0473] Based on the user's feedback, the device issues instructions to the autonomous driving system, automatically adjusting the route to the selected facility and making reservations, allowing the user to automatically head to the next destination without manual intervention.
[0474] Input: User selection
[0475] Output: Instructions to the autonomous driving system and reservation procedures
[0476] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0477] The system of this invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice commands, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience.
[0478] Functionality Overview
[0479] 1. Pre-departure settings
[0480] Device:
[0481] It starts the voice input engine and receives the user's voice commands.
[0482] The voice data is sent to an artificial intelligence model and converted into text.
[0483] Information about the final destination and desired arrival time set by the user is sent to the server.
[0484] The destination information returned from the server is checked and audio feedback is given to the user.
[0485] The emotion engine recognizes emotions from the user's voice data and adjusts the feedback content according to the user's emotions.
[0486] 2. Tourist information while driving
[0487] server:
[0488] Potential tourist spots in the area are selected based on the user's route and current location.
[0489] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[0490] This information is sent to the terminal.
[0491] Device:
[0492] Tourist information from the server is announced to the user by voice.
[0493] The emotion engine recognizes the user's emotions and optimizes tourist destination suggestions based on the results.
[0494] It receives user feedback via voice and sends it to the server.
[0495] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[0496] 3. Restaurant directions while driving
[0497] server:
[0498] Select nearby restaurant options based on the user's current location and route.
[0499] Get restaurant reviews, budget and parking information.
[0500] If necessary, the reservation telephone number is transmitted to the terminal.
[0501] Device:
[0502] Restaurant information sent from the server is announced to the user by voice.
[0503] The emotion engine recognizes the user's emotions and optimizes restaurant suggestions based on the results.
[0504] Receive user feedback and send selections to the server.
[0505] Make a reservation if necessary.
[0506] Once the booking is completed, the user is given audio feedback.
[0507] Recalculate your route to the restaurant and adjust your estimated arrival time at your final destination.
[0508] 4. Hotel reservation information while driving
[0509] server:
[0510] Search for nearby hotel suggestions based on the user's current location and destination.
[0511] Get hotel reviews, popularity, budget, and parking information.
[0512] If necessary, information regarding the reservation procedure is sent to the terminal.
[0513] Device:
[0514] Hotel information sent from the server is announced to the user by voice.
[0515] The emotion engine recognizes the user's emotions and optimizes hotel recommendations based on the results.
[0516] Receive user feedback and send the selections to the server.
[0517] Proceed with the booking process as needed.
[0518] Once the booking is completed, the user is given audio feedback.
[0519] Recalculate your route to the hotel and adjust your estimated arrival time at your final destination.
[0520] Specific examples
[0521] 1. Pre-departure settings
[0522] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[0523] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[0524] Server: Calculates the estimated arrival time and returns it to the device.
[0525] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[0526] On the device: The emotion engine recognizes emotions from the user's voice data and provides additional advice and information depending on the user's mood.
[0527] 2. Tourist information
[0528] User: "What are some nearby tourist spots?"
[0529] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[0530] User: "I'll be there."
[0531] Terminal: "Route set to museum. Arrival time 15 minutes."
[0532] On the device: The emotion engine recognizes the user's emotions and suggests additional tourist destinations that may be of interest.
[0533] 3. Restaurant Information
[0534] User: "I'm hungry, can you find a restaurant nearby?"
[0535] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[0536] User: "Yes, I'd like to make a reservation."
[0537] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[0538] On the device: The emotion engine recognizes the user's emotions and provides adaptive responses to suggested restaurant choices.
[0539] 4. Hotel Reservation Information
[0540] User: "Find a hotel tonight."
[0541] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[0542] User: "Yes, please make a reservation."
[0543] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[0544] On the device: The emotion engine takes into account the user's emotions and provides additional information and suggestions to create a relaxing atmosphere.
[0545] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set destinations and change schedules while driving, and also recognizes the user's emotions through an emotion engine, providing a more personalized and comfortable travel experience.
[0546] The processing flow will be explained below.
[0547] 1. Pre-departure settings
[0548] Step 1:
[0549] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[0550] Step 2:
[0551] The device activates the voice input engine and receives the user's voice command.
[0552] Step 3:
[0553] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0554] Step 4:
[0555] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[0556] Step 5:
[0557] The terminal transmits the extracted information to the server.
[0558] Step 6:
[0559] The server calculates the estimated arrival time and returns the result to the terminal.
[0560] Step 7:
[0561] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[0562] Step 8:
[0563] The device uses an emotion engine to recognize emotions from the user's voice data and provides additional feedback based on the user's emotions, such as "You seem to be in a good mood. Enjoy your trip!"
[0564] 2. Tourist information while driving
[0565] Step 1:
[0566] The user issues a voice command: "Tell me about nearby tourist attractions."
[0567] Step 2:
[0568] The device activates the voice input engine and receives the user's voice command.
[0569] Step 3:
[0570] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0571] Step 4:
[0572] The terminal transmits the converted text data to the server.
[0573] Step 5:
[0574] The server selects nearby tourist spots based on the user's route and current location.
[0575] Step 6:
[0576] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[0577] Step 7:
[0578] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[0579] Step 8:
[0580] The device uses an emotion engine to recognize the user's emotions and adjusts the tourist attraction suggestions accordingly: "Since you seem to be in a good mood, we'll also give you more information about this museum."
[0581] Step 9:
[0582] The user gives verbal feedback: "I'm going there."
[0583] Step 10:
[0584] The device sends the user's feedback to the server.
[0585] Step 11:
[0586] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[0587] Step 12:
[0588] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[0589] 3. Restaurant directions while driving
[0590] Step 1:
[0591] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[0592] Step 2:
[0593] The device activates the voice input engine and receives the user's voice command.
[0594] Step 3:
[0595] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0596] Step 4:
[0597] The terminal transmits the converted text data to the server.
[0598] Step 5:
[0599] The server searches for nearby restaurant options based on the user's current location and route.
[0600] Step 6:
[0601] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[0602] Step 7:
[0603] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[0604] Step 8:
[0605] The device uses an emotion engine to recognize the user's emotions and adjusts restaurant suggestions accordingly: "You seem hungry, so I highly recommend this!"
[0606] Step 9:
[0607] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[0608] Step 10:
[0609] The terminal sends the user's feedback to the server and activates the reservation call function.
[0610] Step 11:
[0611] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[0612] Step 12:
[0613] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[0614] 4. Hotel reservation information while driving
[0615] Step 1:
[0616] A user issues a voice command: "Find a hotel for tonight."
[0617] Step 2:
[0618] The device activates the voice input engine and receives the user's voice command.
[0619] Step 3:
[0620] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0621] Step 4:
[0622] The terminal transmits the converted text data to the server.
[0623] Step 5:
[0624] The server searches for nearby hotel options based on the user's current location and destination.
[0625] Step 6:
[0626] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[0627] Step 7:
[0628] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[0629] Step 8:
[0630] The device uses an emotion engine to recognize the user's emotions and adjusts hotel suggestions accordingly: "I was looking for a place with a relaxing atmosphere."
[0631] Step 9:
[0632] The user gives verbal feedback: "Yes, please book."
[0633] Step 10:
[0634] The device sends the user's feedback to the server and proceeds with the reservation process.
[0635] Step 11:
[0636] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[0637] Step 12:
[0638] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[0639] Example 2
[0640] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0641] While traveling by car, it can be difficult for drivers to safely obtain information on tourist spots, restaurants, and hotels, and to plan optimal routes and make reservations. Furthermore, conventional systems have had difficulty providing personalized information based on the driver's emotions and mood. To solve this problem, a smarter car navigation application system that supports emotion recognition is needed.
[0642] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0643] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model for converting voice data into text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing voice feedback on the estimated arrival time to the user, and means for recognizing emotions from the user's voice data and adjusting the feedback content based on the emotions. This allows the user to safely set a destination and desired arrival time and receive personalized feedback according to their emotions.
[0644] A "voice input engine" is software or hardware that converts voice into a digital signal and analyzes it.
[0645] "Voice command" is a method by which a user gives instructions to a system through voice.
[0646] An "artificial intelligence model" is an algorithm or system that analyzes and processes data to automatically perform a specific task.
[0647] "Text data" is voice data converted into a character string format.
[0648] A "destination" is a location that a user intends to reach using the system.
[0649] "Desired arrival time" is the time at which the user wishes to arrive at the destination.
[0650] A "server" is a computer system that provides functions and data to clients over a network.
[0651] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the system.
[0652] "Feedback" is the response or information that a system provides to a user.
[0653] An "emotion engine" is an algorithm or system that analyzes and recognizes a user's emotions and adjusts responses based on the results.
[0654] "Current location" refers to the location where the user is currently located.
[0655] "Route information" is information about the route to the destination.
[0656] "Candidate tourist destinations" are tourist destination options that the system suggests to users.
[0657] "Popularity" is an indicator that shows how highly a tourist destination or facility is rated by many people.
[0658] "Word of mouth reputation" is review information based on user ratings and impressions.
[0659] "Parking information" is information about where you can park at tourist spots and facilities.
[0660] "Restaurant candidates" are restaurant options that the system suggests to users.
[0661] "Budget" is the amount the user plans to pay.
[0662] A "reservation call" is a telephone means of contact for making restaurant or hotel reservations.
[0663] "Facilities" are places and buildings that users visit, such as tourist attractions, restaurants, and hotels.
[0664] The system of the present invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience. A specific embodiment of this system is described below.
[0665] Hardware and Software Use
[0666] Device: Smartphone or dedicated car navigation device (e.g., Android device)
[0667] Server: Cloud-based backend server (e.g., AWS or Google Cloud)
[0668] software:
[0669] Speech recognition engine: Google Cloud Speech-to-Text
[0670] Artificial intelligence model: GPT-4
[0671] Emotion recognition engine: Microsoft Azure Emotion API
[0672] Route Calculation Engine
[0673] Database: tourist attractions, restaurants, and hotel information
[0674] Processing flow
[0675] 1. Voice to Text
[0676] The user enters a voice command.
[0677] The device activates a voice input engine and converts the voice data into text using Google Cloud Speech-to-Text.
[0678] The converted text data is sent to an artificial intelligence model (GPT-4) for analysis.
[0679] 2. Set your destination and desired arrival time
[0680] The terminal transmits information about the final destination and desired arrival time set by the user to the server.
[0681] The server uses a route calculation engine to calculate the estimated arrival time and returns the result to the terminal.
[0682] The terminal provides the user with audio feedback on the estimated arrival time.
[0683] An emotion engine is used to recognize emotions from the user's voice data and adjust the feedback content.
[0684] 3. Tourist information
[0685] The user inputs a voice command such as "Tell me about nearby tourist spots."
[0686] The device converts the voice command into text and sends the current location information to the server.
[0687] The server searches for potential tourist destinations and obtains popularity, reviews, photos, and parking information.
[0688] The terminal notifies the user of the received information by voice and receives feedback.
[0689] Recalculate the route to the tourist spot selected by the user and adjust the estimated arrival time.
[0690] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[0691] 4. Restaurant Information
[0692] The user enters a voice command such as "I'm hungry, find a nearby restaurant."
[0693] The device converts the voice command into text and sends the current location information to the server.
[0694] The server searches for restaurant candidates and retrieves reviews, budget, and parking information.
[0695] The terminal notifies the user of the received information by voice and receives feedback.
[0696] Make a reservation call if necessary.
[0697] Recalculate your route to the selected restaurant and adjust your estimated arrival time.
[0698] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[0699] 5. Hotel Reservation Information
[0700] The user enters the voice command "Find a hotel tonight."
[0701] The device converts the voice command into text and sends the current location information to the server.
[0702] The server searches for hotel options and retrieves reviews, popularity, budget, and parking information.
[0703] The terminal notifies the user of the received information by voice and receives feedback.
[0704] Make a reservation call if necessary.
[0705] Recalculate your route to the selected hotel and adjust your estimated arrival time.
[0706] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[0707] Prompt Sentence Examples
[0708] As a concrete example, the following prompt sentence will be used.
[0709] Pre-departure setup:
[0710] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[0711] Tourist Information:
[0712] "Tell me about nearby tourist spots."
[0713] Restaurant Information:
[0714] I'm hungry, so I'm looking for a nearby restaurant.
[0715] Hotel Reservation Information:
[0716] "Find a hotel to stay at tonight."
[0717] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set destinations and change schedules while driving. It also recognizes the user's emotions through an emotion engine, providing a more personalized and comfortable travel experience.
[0718] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0719] Step 1:
[0720] User: Enters a voice command.
[0721] Specific operation: The user gives voice instructions such as, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 p.m."
[0722] Input: Voice command
[0723] Output: Audio data
[0724] Step 2:
[0725] Device: Activates the voice input engine and converts the voice data into text.
[0726] Specific operation: Converts audio data into text data using Google Cloud Speech-to-Text.
[0727] Input: Audio data
[0728] Output: Text data
[0729] Step 3:
[0730] Terminal: The converted text data is sent to the artificial intelligence model for analysis.
[0731] What it does: It uses GPT-4 to parse text and extract the destination and desired arrival time.
[0732] Input: Text data
[0733] Output: Analysis results (destination and desired arrival time)
[0734] Step 4:
[0735] Terminal: Sends the extracted information to the server.
[0736] Specific operation: Send information to the server about the final destination "Tokyo Station" and the desired arrival time "3:00 PM".
[0737] Input: Analysis results (destination and desired arrival time)
[0738] Output: Request to server
[0739] Step 5:
[0740] Server: Calculates the estimated arrival time using a route calculation engine and returns the result to the terminal.
[0741] Specific operation: Uses a route calculation engine to calculate the optimal route to the destination and derives the estimated arrival time.
[0742] Input: Destination and desired arrival time
[0743] Output: Estimated arrival time
[0744] Step 6:
[0745] Terminal: Provides audio feedback to the user on the estimated arrival time.
[0746] Specific operation: Using a speech synthesis engine, the system notifies the user, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[0747] Input: Estimated arrival time
[0748] Output: Feedback audio
[0749] Step 7:
[0750] Terminal: Activates the emotion engine and recognizes emotions from the user's voice data.
[0751] Specific behavior: Analyzes user emotions using the Microsoft Azure Emotion API.
[0752] Input: Audio data
[0753] Output: Emotion data
[0754] Step 8:
[0755] Device: Adjusts feedback content based on user emotional data.
[0756] Specific behavior: Based on the results of the emotion engine, adjust the feedback content and provide appropriate additional information.
[0757] Input: Emotion data
[0758] Output: Adjusted feedback content
[0759] Step 9:
[0760] User: Enters voice commands for nearby tourist attractions.
[0761] Specific operation: The user gives a voice command such as "Tell me about nearby tourist spots."
[0762] Input: Voice command
[0763] Output: Audio data
[0764] Step 10:
[0765] Device: Converts voice commands into text and sends location information to the server.
[0766] Specific operation: The voice data is converted into text data using Google Cloud Speech-to-Text and sent to the server along with GPS information.
[0767] Input: Audio data
[0768] Output: Text data and current location information
[0769] Step 11:
[0770] Server: Search for potential tourist destinations and obtain popularity, reviews, photos, and parking information.
[0771] Specific operation: Filters tourist destination candidates based on the current location from the database and obtains information about each location.
[0772] Input: Current location information
[0773] Output: Tourist destination information
[0774] Step 12:
[0775] Server: Sends the acquired tourist spot information to the terminal.
[0776] Specific operation: Send tourist spot information to the terminal.
[0777] Input: Tourist destination information
[0778] Output: Response to terminal
[0779] Step 13:
[0780] Terminal: Announces the received tourist spot information to the user by voice and receives feedback.
[0781] Specific operation: Uses a speech synthesis engine to notify the user of tourist information and receive feedback.
[0782] Input: Tourist destination information
[0783] Output: Feedback speech and user feedback
[0784] Step 14:
[0785] Terminal: Recalculate the route to the selected tourist spot and adjust the estimated arrival time.
[0786] What it does: Uses the route calculation engine to calculate a new route and adjust the estimated arrival time.
[0787] Input: Selected tourist destination
[0788] Output: New route and estimated arrival time
[0789] Step 15:
[0790] Device: The emotion engine recognizes the user's emotions and optimizes the suggestions.
[0791] What it does: It uses an emotion engine to analyze the user's emotions and adjusts suggestions based on the results.
[0792] Input: Emotion data
[0793] Output: Optimized proposals
[0794] (Application example 2)
[0795] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0796] While traveling by car, users need to obtain real-time information on tourist spots, restaurants, hotels, etc., and make reservations. However, conventional car navigation systems do not provide personalized suggestions based on the user's emotions. Another issue is that it is difficult to ensure safety when users set destinations or make reservations while driving. To solve these issues, a new car navigation system that combines voice input and emotion recognition functions is needed.
[0797] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0798] In this invention, the server includes means for activating a voice input engine and receiving voice commands from the user, processing means using an artificial intelligence model to convert voice data into text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing voice feedback on the estimated arrival time to the user, and means including an emotion engine for recognizing emotions from the user's voice data and adjusting the feedback content. This makes it possible to make personalized suggestions based on the user's emotions, providing a safe and comfortable travel experience.
[0799] "Voice input engine" is a general term for devices and software for receiving voice commands from users.
[0800] "Artificial Intelligence Model" means the machine learning algorithms and techniques used to convert voice data into text.
[0801] "Destination and desired arrival time" refers to the final destination set by the user and the desired arrival time at that destination.
[0802] A "server" is a computer system that processes information entered by a user and provides the necessary information and estimated time.
[0803] The "estimated arrival time" is the estimated time required to arrive at the destination specified by the user.
[0804] The "emotion engine" is a technology and algorithm that recognizes emotions from the user's voice data and optimizes the feedback content based on the results.
[0805] "Candidate tourist destinations" are potential travel destinations suggested based on the user's current location and route information.
[0806] "Popularity, word-of-mouth reputation, photos and parking information" is a general term for ratings and reviews of tourist spots and restaurants, as well as images and information about parking.
[0807] "Feedback" refers to the response or guidance provided by the system to the user.
[0808] "Recalculating the route" means recalculating the travel route based on the user's selection and the situation.
[0809] The system of the present invention is a car navigation application system that supports travel in autonomous vehicles, and uses the following main hardware and software:
[0810] 1. Voice Input Engine
[0811] A voice input engine is a device or software that receives voice commands from users, specifically voice recognition engines such as Google Voice Recognition and Apple's Siri, allowing users to input commands using only their voice without using their hands.
[0812] 2. Artificial Intelligence Model
[0813] Artificial intelligence models are used to convert the speech data into text, such as Google Cloud Speech-to-Text API and IBM Watson Speech to Text. This process converts the speech data into text.
[0814] 3. Emotion Engine
[0815] The emotion engine is a technology and algorithm that recognizes emotions from users' voice data and optimizes feedback content. Specifically, it uses Microsoft Azure Emotion API and Affectiva's emotion recognition technology. This enables personalized suggestions to be provided based on the user's emotions.
[0816] 4. Navigation system
[0817] The navigation system uses the Google Maps API, Here Maps API, etc. to calculate the route to the destination, locate the current location, and estimate the arrival time. It also obtains information on nearby tourist attractions, restaurants, hotels, etc.
[0818] 5. Recommendation Services
[0819] The recommendation service provides information on tourist spots, restaurants, and hotels using APIs such as FourSquare and Yelp, and obtains detailed information such as popularity, reviews, photos, and parking information, and makes suggestions to users.
[0820] Specific examples
[0821] Example of user voice input and system response
[0822] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[0823] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[0824] Server: Calculates the estimated arrival time and returns it to the device.
[0825] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[0826] On the device: The emotion engine recognizes emotions from the user's voice data and provides additional advice and information depending on the user's mood.
[0827] Example of tourist information
[0828] User: "What are some nearby tourist spots?"
[0829] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[0830] User: "I'll be there."
[0831] Terminal: "Route set to museum. Arrival time 15 minutes."
[0832] On the device: The emotion engine recognizes the user's emotions and suggests additional tourist destinations that may be of interest.
[0833] This allows users to have a comfortable and personalized travel experience.
[0834] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0835] Step 1:
[0836] Receiving voice commands
[0837] The user issues a voice command.
[0838] The terminal activates a voice input engine and receives voice commands from the user.
[0839] Input: User's voice command
[0840] Output: Audio data
[0841] Step 2:
[0842] Converting audio data to text
[0843] The device converts the voice data into text using an artificial intelligence model (e.g., Google Cloud Speech-to-Text API).
[0844] Input: Audio data
[0845] Output: Text data
[0846] Step 3:
[0847] Extracting destination and desired arrival time
[0848] The device uses natural language processing technology to extract the destination and desired arrival time from the text data.
[0849] Input: Text data
[0850] Output: Extracted information about destination and desired arrival time
[0851] Step 4:
[0852] Destination and desired arrival time sent to server
[0853] The terminal transmits the extracted information to the server and requests it to calculate the estimated arrival time.
[0854] Input: Extract information about destination and desired arrival time
[0855] Output: Request to calculate estimated arrival time
[0856] Step 5:
[0857] Calculating and receiving estimated arrival times
[0858] The server uses a navigation system (e.g., Google Maps API) to calculate the estimated arrival time and sends this information back to the device.
[0859] Input: Extract information about destination and desired arrival time
[0860] Output: Estimated arrival time
[0861] Step 6:
[0862] emotion recognition
[0863] The device uses an emotion engine (e.g., Microsoft Azure Emotion API) to recognize emotions from the user's voice data.
[0864] Input: User's voice data
[0865] Output: Emotion data
[0866] Step 7:
[0867] Personalized Feedback
[0868] The terminal adjusts the feedback content of the estimated arrival time based on the emotion data and provides the feedback to the user by voice.
[0869] Input: Estimated arrival time, emotion data
[0870] Output: Audio feedback
[0871] Step 8:
[0872] Search and submit tourist destination suggestions
[0873] The server searches for nearby tourist spots based on the current location and route information using a recommendation service (e.g., FourSquare API) and sends detailed information to the device.
[0874] Input: Current location and route information
[0875] Output: Tourist destination information
[0876] Step 9:
[0877] To announce tourist destination information and receive feedback
[0878] The terminal notifies the user of tourist spot information by voice and receives feedback from the user.
[0879] Input: Tourist destination information
[0880] Output: User feedback
[0881] Step 10:
[0882] Recalculate route to selected tourist spot
[0883] The server recalculates the route to the tourist spot selected by the user and adjusts the estimated arrival time.
[0884] Input: User feedback
[0885] Output: Recalculated route and estimated arrival time
[0886] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0887] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0888] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0889] [Second embodiment]
[0890] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0891] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0892] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0893] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0894] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0895] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0896] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0897] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0898] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0899] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0900] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0901] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0902] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time.
[0903] Functionality Overview
[0904] 1. Pre-departure settings
[0905] Device:
[0906] It starts the voice input engine and receives the user's voice commands.
[0907] The voice data is sent to an artificial intelligence model and converted into text.
[0908] Information about the final destination and desired arrival time set by the user is sent to the server.
[0909] The destination information returned from the server is checked and audio feedback is given to the user.
[0910] 2. Tourist information while driving
[0911] server:
[0912] Potential tourist spots in the area are selected based on the user's route and current location.
[0913] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[0914] This information is sent to the terminal.
[0915] Device:
[0916] Tourist information from the server is announced to the user by voice.
[0917] It receives user feedback and sends it to the server.
[0918] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[0919] 3. Restaurant directions while driving
[0920] server:
[0921] Select nearby restaurant options based on the user's current location and route.
[0922] Obtain information such as restaurant reviews, budget, and whether parking is available.
[0923] If necessary, the reservation telephone number is transmitted to the terminal.
[0924] Device:
[0925] Restaurant information sent from the server is announced to the user by voice.
[0926] Receive user feedback and send selections to the server.
[0927] Make a reservation if necessary.
[0928] It provides the functionality to receive user feedback, send it to the server, and make reservation calls.
[0929] 4. Hotel reservation information while driving
[0930] server:
[0931] Select nearby hotel options based on the user's current location and destination.
[0932] Obtain hotel reviews, popularity, budget, parking information, etc.
[0933] If necessary, information for reservation procedures is sent to the terminal.
[0934] Device:
[0935] Hotel information sent from the server is announced to the user by voice.
[0936] Receive user feedback and send the selection results to the server.
[0937] Proceed with the booking process as needed.
[0938] Once the booking is completed, the user is given audio feedback.
[0939] Recalculate the route to the user's selected hotel and adjust the estimated arrival time.
[0940] Specific examples
[0941] 1. Pre-departure settings
[0942] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[0943] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[0944] Server: Calculates the estimated arrival time and returns it to the device.
[0945] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[0946] 2. Tourist information
[0947] User: "What are some nearby tourist spots?"
[0948] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[0949] User: "I'll be there."
[0950] Terminal: "Route set to museum. Arrival time 15 minutes."
[0951] 3. Restaurant Information
[0952] User: "I'm hungry, can you find a restaurant nearby?"
[0953] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[0954] User: "Yes, I'd like to make a reservation."
[0955] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[0956] 4. Hotel Reservation Information
[0957] User: "Find a hotel tonight."
[0958] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[0959] User: "Yes, please make a reservation."
[0960] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[0961] In this way, the system of the present invention allows all operations to be performed by voice, allowing you to safely set your destination or change your schedule while driving, providing a comfortable travel experience.
[0962] The processing flow will be explained below.
[0963] 1. Pre-departure settings
[0964] Step 1:
[0965] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[0966] Step 2:
[0967] The device activates the voice input engine and receives the user's voice command.
[0968] Step 3:
[0969] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0970] Step 4:
[0971] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[0972] Step 5:
[0973] The terminal transmits the extracted information to the server.
[0974] Step 6:
[0975] The server calculates the estimated arrival time and returns the result to the terminal.
[0976] Step 7:
[0977] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[0978] 2. Tourist information while driving
[0979] Step 1:
[0980] The user issues a voice command: "Tell me about nearby tourist attractions."
[0981] Step 2:
[0982] The device activates the voice input engine and receives the user's voice command.
[0983] Step 3:
[0984] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[0985] Step 4:
[0986] The terminal transmits the converted text data to the server.
[0987] Step 5:
[0988] The server selects nearby tourist spots based on the user's route and current location.
[0989] Step 6:
[0990] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[0991] Step 7:
[0992] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[0993] Step 8:
[0994] The user gives verbal feedback: "I'm going there."
[0995] Step 9:
[0996] The device sends the user's feedback to the server.
[0997] Step 10:
[0998] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[0999] Step 11:
[1000] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[1001] 3. Restaurant directions while driving
[1002] Step 1:
[1003] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[1004] Step 2:
[1005] The device activates the voice input engine and receives the user's voice command.
[1006] Step 3:
[1007] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1008] Step 4:
[1009] The terminal transmits the converted text data to the server.
[1010] Step 5:
[1011] The server searches for nearby restaurant options based on the user's current location and route.
[1012] Step 6:
[1013] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[1014] Step 7:
[1015] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[1016] Step 8:
[1017] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[1018] Step 9:
[1019] The terminal sends the user's feedback to the server and activates the reservation call function.
[1020] Step 10:
[1021] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[1022] Step 11:
[1023] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[1024] 4. Hotel reservation information while driving
[1025] Step 1:
[1026] A user issues a voice command: "Find a hotel for tonight."
[1027] Step 2:
[1028] The device activates the voice input engine and receives the user's voice command.
[1029] Step 3:
[1030] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1031] Step 4:
[1032] The terminal transmits the converted text data to the server.
[1033] Step 5:
[1034] The server searches for nearby hotel options based on the user's current location and destination.
[1035] Step 6:
[1036] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[1037] Step 7:
[1038] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[1039] Step 8:
[1040] The user gives verbal feedback: "Yes, please book."
[1041] Step 9:
[1042] The device sends the user's feedback to the server and proceeds with the reservation process.
[1043] Step 10:
[1044] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[1045] Step 11:
[1046] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[1047] Example 1
[1048] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1049] Current car navigation systems have the challenge of making it difficult for users to safely and efficiently set destinations and adjust schedules while driving. Furthermore, they lack the functionality to provide comprehensive, real-time information on tourist spots, restaurants, hotels, and other information, preventing users from making appropriate choices quickly. This often results in a loss of convenience and safety while driving.
[1050] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1051] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for suggesting tourist spots, restaurants, and hotels based on the user's current location and destination and acquiring information, means for notifying the user of the acquired information audio-visually and receiving feedback, means for recalculating a route to the selected destination and adjusting the estimated arrival time, and means for making reservations if necessary. This allows destination setting and schedule adjustments to be performed safely and efficiently even while driving, providing a comfortable travel experience.
[1052] A "voice input engine" is a device or software that receives a user's voice commands and converts them into digital data.
[1053] "Voice Data" means human speech information converted into digital form by a speech input engine.
[1054] "Artificial intelligence model" refers to technologies such as machine learning algorithms and neural networks used to convert voice data into text data.
[1055] "Text data" is character string information converted from voice data by an artificial intelligence model.
[1056] "Destination" is information indicating the place or location to which the user wishes to travel.
[1057] The "desired arrival time" is information indicating a specific time at which the user desires to arrive at the destination.
[1058] A "server" is a central system that receives requests on a computer network, processes data, and sends and receives information.
[1059] "Estimated arrival time" is information indicating the estimated arrival time from the current location to the destination.
[1060] "Tourist destination" refers to tourist spots and famous places that users aim to visit.
[1061] "Restaurant" means an eating and drinking establishment selected by a patron for dining.
[1062] "Hotel" means the accommodation facility selected for the Guest's stay.
[1063] "Means of acquisition" refers to the methods and technologies used to collect and receive the required information or data.
[1064] "Means of receiving feedback" refers to the methods and techniques for receiving responses or reactions from users.
[1065] "Means for recalculating a route" refers to a method or technology for recalculating a new route based on specified conditions.
[1066] "Means for completing reservation procedures" refers to the methods and technologies used to make reservations for the service selected by the user (such as restaurant reservations or hotel reservations).
[1067] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, adjusting routes and arrival times in real time.
[1068] The main components of the system are the "terminals" used by users and the "servers" that process data. The specific configuration and operation are explained below.
[1069] The system's devices, which include smartphones and car navigation systems, are equipped with a voice input engine that receives users' voice commands and converts the voice data into a digital format. This digital voice data is then converted into text data using voice recognition software such as Google Cloud Speech-to-Text.
[1070] The converted text data is sent to the server and used to extract the destination and desired arrival time. The server calculates the estimated arrival time based on this information and returns the result to the terminal. The terminal then provides this information as feedback to the user via voice. Specifically, it notifies the user by saying something like, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[1071] In the case of tourist attraction guidance while driving, the device sends the user's current location and route information to the server, and the server searches for nearby tourist attractions. The server obtains the tourist attraction's popularity, reviews, photos, and parking information and sends them to the device. The device then announces this information to the user by voice (e.g., "Yes, there's a famous museum nearby. It has a 4.5 rating and there's parking available. What would you like to do?"). Based on the user's feedback, the server recalculates the route and calculates a new estimated arrival time.
[1072] Regarding restaurant guidance, the system searches for nearby restaurants based on the current location and route, and obtains information such as reviews, budget, and parking information. If necessary, it obtains a reservation phone number and sends it to the terminal. The terminal then announces this information to the user by voice (e.g., "There's a highly rated Chinese restaurant nearby. The reviews are 4.7, and your budget is around 1,500 yen for lunch. Would you like to make a reservation?") and proceeds with the reservation process based on the user's feedback.
[1073] Additionally, the hotel guide searches for nearby hotels based on the current location and destination, and obtains information such as reviews, popularity, budget, and parking information. If a reservation is required, the server sends the necessary information to the terminal, which then proceeds with the reservation process (e.g., "Looking for a hotel to stay at tonight," "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?").
[1074] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set their destination or change their schedule while driving, providing a comfortable travel experience.
[1075] Specific prompt examples
[1076] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[1077] "Tell me about nearby tourist spots."
[1078] I'm hungry, so I'm looking for a nearby restaurant.
[1079] "Find a hotel to stay at tonight."
[1080] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1081] The flow of this system's program processing
[1082] 1. Pre-departure settings
[1083] Step 1:
[1084] Subject: User
[1085] Description: The user speaks a voice command into the device.
[1086] Specific actions: For example, say, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 PM."
[1087] Step 2:
[1088] Subject: Device
[1089] Description: Starts the voice input engine and receives voice commands.
[1090] Input: Voice command from the user
[1091] Output: Digital audio data
[1092] What it does: The voice input engine receives the voice and converts it into a digital format.
[1093] Step 3:
[1094] Subject: Device
[1095] Description: Uses artificial intelligence models to convert voice data into text.
[1096] Input: Digital audio data
[1097] Output: Text data
[1098] What it does: Converts speech to text using a service like Google Cloud Speech-to-Text.
[1099] Step 4:
[1100] Subject: Device
[1101] Description: Sends the converted text data to the server.
[1102] Input: Text data (e.g., "Set a route to Tokyo Station, with a desired arrival time of 3:00 PM.")
[1103] Output: Request to server
[1104] Specific operation: Sends an HTTP request to the server.
[1105] Step 5:
[1106] Subject: Server
[1107] Description: Extracts the destination and desired arrival time and calculates the estimated arrival time.
[1108] Input: Text data
[1109] Output: Estimated arrival time
[1110] Specific operation: Uses natural language processing technology to analyze text and calculate arrival times by referencing traffic information and road conditions.
[1111] Step 6:
[1112] Subject: Server
[1113] Description: Sends the calculation result back to the terminal.
[1114] Input: Estimated arrival time
[1115] Output: Feedback data
[1116] Specific operation: Returns the calculation result to the terminal.
[1117] Step 7:
[1118] Subject: Device
[1119] Description: Provides audio feedback data from the server to the user.
[1120] Input: Feedback data
[1121] Output: Audio notification to the user
[1122] Specific operation: The text data is converted into speech, and a message such as "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM" is spoken to the user.
[1123] 2. Tourist information while driving
[1124] Step 1:
[1125] Subject: User
[1126] Description: Request "Tell me about nearby tourist spots."
[1127] Specific action: Speak a voice command into the device.
[1128] Step 2:
[1129] Subject: Device
[1130] Description: Converts voice commands into text and sends it to the server.
[1131] Input: Voice command
[1132] Output: Text data
[1133] Specific operation: The audio is converted into text using Google Cloud Speech-to-Text or similar and sent to the server.
[1134] Step 3:
[1135] Subject: Server
[1136] Description: Searches for nearby tourist spots based on the user's current location and route.
[1137] Input: Current location and route information
[1138] Output: List of tourist destination candidates
[1139] Specific behavior: Generate a list of tourist destinations by retrieving information from a database or external API.
[1140] Step 4:
[1141] Subject: Server
[1142] Description: Get tourist attraction popularity, reviews, photos, and parking information.
[1143] Input: Tourist destination candidate list
[1144] Output: Detailed information
[1145] Specific actions: Collect and list detailed information about each tourist destination.
[1146] Step 5:
[1147] Subject: Server
[1148] Description: Sends the acquired information to the device.
[1149] Input: More information
[1150] Output: Feedback data
[1151] Specific operation: Sends collected information to the device.
[1152] Step 6:
[1153] Subject: Device
[1154] Description: Presents information to the user audibly.
[1155] Input: Feedback data
[1156] Output: Audio notification to the user
[1157] Action: The audio message is, "There's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[1158] Step 7:
[1159] Subject: User
[1160] Description: Gives feedback saying "There you go."
[1161] Specific actions: Select a tourist spot specified by voice.
[1162] Step 8:
[1163] Subject: Device
[1164] Description: Recalculates route based on user selection.
[1165] Input: User's choice
[1166] Output: Recalculated route
[1167] Specific behavior: Calculate a new route and send a notification such as "Route to the museum has been set. It will take 15 minutes to arrive."
[1168] 3. Restaurant directions while driving
[1169] Step 1:
[1170] Subject: User
[1171] Description: "I'm hungry, find me a nearby restaurant."
[1172] Specific action: Speak a voice command into the device.
[1173] Step 2:
[1174] Subject: Device
[1175] Description: Sends a voice command to the server.
[1176] Input: Voice command
[1177] Output: Text data
[1178] Specific operation: Converts speech into text and sends it to the server.
[1179] Step 3:
[1180] Subject: Server
[1181] Description: Find nearby restaurant suggestions based on your current location and route.
[1182] Input: Current location and route information
[1183] Output: Restaurant candidate list
[1184] Specific operation: Retrieves restaurant information from a database or external API and creates a list.
[1185] Step 4:
[1186] Subject: Server
[1187] Description: Get detailed information like reviews, budget, parking info, etc.
[1188] Input: Restaurant candidate list
[1189] Output: Detailed information
[1190] What it does: Collect and list detailed information about each restaurant.
[1191] Step 5:
[1192] Subject: Server
[1193] Description: Sends the acquired information to the device.
[1194] Input: More information
[1195] Output: Feedback data
[1196] Specific operation: Sends collected information to the device.
[1197] Step 6:
[1198] Subject: Device
[1199] Description: Presents information to the user audibly.
[1200] Input: Feedback data
[1201] Output: Audio notification to the user
[1202] Specific action: Say, "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[1203] Step 7:
[1204] Subject: User
[1205] Instructions: Answer "Yes, I'd like to make a reservation as well."
[1206] Specific action: Indicate your intention to make a reservation by voice.
[1207] Step 8:
[1208] Subject: Device
[1209] Description: Proceed with the booking process.
[1210] Input: User's booking request
[1211] Output: Reservation completion notification
[1212] Specific operation: Make a reservation using the OpenTable API or similar, and notify the customer, "Your reservation has been arranged. You will arrive in 10 minutes."
[1213] 4. Hotel reservation information while driving
[1214] Step 1:
[1215] Subject: User
[1216] Description: "Find me a hotel tonight."
[1217] Specific action: Speak a voice command into the device.
[1218] Step 2:
[1219] Subject: Device
[1220] Description: Sends a voice command to the server.
[1221] Input: Voice command
[1222] Output: Text data
[1223] Specific operation: Converts voice commands into text and sends it to the server.
[1224] Step 3:
[1225] Subject: Server
[1226] Description: Search for nearby hotel suggestions based on your current location and destination.
[1227] Input: Current location and destination information
[1228] Output: Hotel candidate list
[1229] Specific operation: Retrieves hotel information from a database or external API and creates a list.
[1230] Step 4:
[1231] Subject: Server
[1232] Description: Get detailed information like reviews, popularity, budget, parking information, and more.
[1233] Input: Hotel candidate list
[1234] Output: Detailed information
[1235] What to do: Collect and list detailed information about each hotel.
[1236] Step 5:
[1237] Subject: Server
[1238] Description: Sends the acquired information to the device.
[1239] Input: More information
[1240] Output: Feedback data
[1241] Specific operation: Sends collected information to the device.
[1242] Step 6:
[1243] Subject: Device
[1244] Description: Presents information to the user audibly.
[1245] Input: Feedback data
[1246] Output: Audio notification to the user
[1247] Action: The voice will say, "There's a highly rated hotel nearby with a 4.8 rating and parking. Would you like to make a reservation?"
[1248] Step 7:
[1249] Subject: User
[1250] Instructions: "Yes, please make a reservation."
[1251] Specific action: Indicate your intention to make a reservation by voice.
[1252] Step 8:
[1253] Subject: Device
[1254] Description: Proceed with the booking process.
[1255] Input: User's booking request
[1256] Output: Reservation completion notification
[1257] Specific behavior: Make a reservation using Expedia API or Booking.com API and notify the user, "Hotel has been booked. Estimated arrival time is 6:00 PM."
[1258] (Application example 1)
[1259] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1260] When traveling by car, there is a demand for voice control to select and reserve destinations, tourist attractions, restaurants, and hotels along the way, and for integration with autonomous driving systems to provide a safe and comfortable travel experience. However, with conventional systems, it can be difficult to provide sufficient information, complete reservations, and adjust routes using voice control alone. Furthermore, there is a lack of integration with autonomous driving systems, and manual operation by the user is often required. This can make operations while driving cumbersome, potentially compromising safety and efficiency.
[1261] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1262] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for adding control means compatible with the autonomous driving system, and means for automating automatic route adjustment to selected facilities and reservation procedures. This makes it possible for an autonomous vehicle to utilize voice operation and AI technology to select and reserve tourist spots, restaurants, and hotels in real time based on user instructions, and to automatically adjust the route and provide feedback on the arrival time.
[1263] A "voice input engine" is a device or software that receives voice commands from a user.
[1264] "Artificial intelligence model" is a machine learning algorithm used to convert voice data into text.
[1265] "Text data" is voice data converted into text format.
[1266] A "server" is a computer system that provides services to clients over a computer network.
[1267] The "destination" is the destination point set by the user.
[1268] "Desired arrival time" is the time the user desires to arrive at the destination.
[1269] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the server.
[1270] An "autonomous driving system" is a technology that enables vehicles to perform driving operations autonomously.
[1271] "Route adjustment" refers to recalculating the route to a selected destination.
[1272] "Tourist destination candidates" is a list of tourist destinations suggested to the user.
[1273] "Popularity" is an indicator that shows the evaluation of tourist destinations and facilities.
[1274] "Word of mouth" refers to reviews and feedback from users.
[1275] "Photos" are image data that provide visual images of tourist spots and facilities.
[1276] "Parking information" is information about parking spaces at tourist spots and facilities.
[1277] "Restaurant candidates" is a list of restaurants suggested to the user.
[1278] The "budget" is an estimate of the cost for the user to use the service.
[1279] "Reservation phone" refers to a means of making a reservation for a facility by telephone.
[1280] In one embodiment of the present invention, the system first activates a voice input engine to receive a user's voice command. The voice input engine may be, for example, the Google Speech Recognition API or a similar voice recognition engine. The device that receives the voice data converts it into text data using an artificial intelligence model (e.g., a generative AI model).
[1281] Natural language processing (NLP) techniques are used to extract the destination and desired arrival time from the converted text data. The text data is analyzed to extract specific information (in this case, the destination and desired arrival time). This information is sent to a server, which calculates the estimated arrival time. The calculation result is sent back to the device, which then communicates the result to the user as voice feedback.
[1282] By adding control means compatible with the autonomous driving system, it becomes possible to automatically adjust routes to selected facilities and make reservations. Specifically, the terminal provides route information to the autonomous driving system based on information received from the server and uses an API for reservation procedures. This allows users to set destinations and make reservations using only voice commands without manual operation.
[1283] For example, if a user voice-inputs "My destination is Tokyo Station, and I would like to arrive at 3:00 PM," the device converts the voice data into text data and extracts the destination and desired arrival time. This information is sent to the server, which calculates the estimated arrival time and receives feedback. Furthermore, if the user issues the command "Find a nearby restaurant," the server searches for restaurant candidates based on the current location and route information, obtains reviews and budget information, and sends it to the device. The device notifies the user of this by voice, and if the user responds "Make a reservation," the device will automatically complete the reservation procedure.
[1284] An example of a prompt sentence might be:
[1285] Please let us know your destination and desired arrival time.
[1286] Could you tell me about nearby tourist spots?
[1287] Find a restaurant near you.
[1288] Find a hotel to stay in tonight.
[1289] In this way, the present invention combines voice control with an automated driving system to provide users with a safe and convenient travel experience.
[1290] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1291] Step 1:
[1292] The device starts a voice input engine to receive the user's voice command. The voice input engine used here is a speech recognition engine such as the Google Speech Recognition API. The voice input engine captures the user's voice data and processes it as digital voice data.
[1293] Input: User's voice command
[1294] Output: Digital audio data
[1295] Step 2:
[1296] The voice data received by the device is converted into text data using a generative AI model (a speech recognition algorithm). This process involves analyzing the voice signal and generating the corresponding text.
[1297] Input: Digital audio data
[1298] Output: Text data
[1299] Step 3:
[1300] The device extracts the destination and desired arrival time from the generated text data, using natural language processing (NLP) techniques to identify keywords and phrases related to the destination and desired arrival time.
[1301] Input: Text data
[1302] Output: Destination and desired arrival time information
[1303] Step 4:
[1304] The device sends the extracted information to a server, which calculates and returns an estimated arrival time. The server then uses a map database and traffic data to calculate a route from the current location to the destination.
[1305] Input: Destination and desired arrival time information
[1306] Output: Estimated arrival time
[1307] Step 5:
[1308] The terminal receives the estimated arrival time from the server and provides the user with audio feedback. Here, a synthetic speech engine (e.g., pyttsx3) is used to generate speech from text and communicate it to the user.
[1309] Input: Estimated arrival time
[1310] Output: Audio feedback
[1311] Step 6:
[1312] The user inputs an additional voice command (e.g., "Tell me about nearby tourist spots," "Find nearby restaurants," etc.). Based on this prompt, the device queries the server.
[1313] Input: User's additional voice command
[1314] Output: prompt statement
[1315] Step 7:
[1316] The server searches for potential tourist spots and restaurants based on the user's current location and route information, and then retrieves and sends the information to the device. The retrieved information includes popularity, reviews, photos, parking information, budget, etc.
[1317] Input: User's current location and route information
[1318] Output: Information on tourist spots and restaurants
[1319] Step 8:
[1320] The terminal announces the information received from the server to the user by voice and receives user feedback. When the user makes a selection, the selection is sent to the server, which then processes the new route and reservation.
[1321] Input: Information about tourist attractions and restaurants
[1322] Output: Audio feedback and user selection results
[1323] Step 9:
[1324] Based on the user's feedback, the device issues instructions to the autonomous driving system, automatically adjusting the route to the selected facility and making reservations, allowing the user to automatically head to the next destination without manual intervention.
[1325] Input: User selection
[1326] Output: Instructions to the autonomous driving system and reservation procedures
[1327] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1328] The system of this invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice commands, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience.
[1329] Functionality Overview
[1330] 1. Pre-departure settings
[1331] Device:
[1332] It starts the voice input engine and receives the user's voice commands.
[1333] The voice data is sent to an artificial intelligence model and converted into text.
[1334] Information about the final destination and desired arrival time set by the user is sent to the server.
[1335] The destination information returned from the server is checked and audio feedback is given to the user.
[1336] The emotion engine recognizes emotions from the user's voice data and adjusts the feedback content according to the user's emotions.
[1337] 2. Tourist information while driving
[1338] server:
[1339] Potential tourist spots in the area are selected based on the user's route and current location.
[1340] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[1341] This information is sent to the terminal.
[1342] Device:
[1343] Tourist information from the server is announced to the user by voice.
[1344] The emotion engine recognizes the user's emotions and optimizes tourist destination suggestions based on the results.
[1345] It receives user feedback via voice and sends it to the server.
[1346] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[1347] 3. Restaurant directions while driving
[1348] server:
[1349] Select nearby restaurant options based on the user's current location and route.
[1350] Get restaurant reviews, budget and parking information.
[1351] If necessary, the reservation telephone number is transmitted to the terminal.
[1352] Device:
[1353] Restaurant information sent from the server is announced to the user by voice.
[1354] The emotion engine recognizes the user's emotions and optimizes restaurant suggestions based on the results.
[1355] Receive user feedback and send selections to the server.
[1356] Make a reservation if necessary.
[1357] Once the booking is completed, the user is given audio feedback.
[1358] Recalculate your route to the restaurant and adjust your estimated arrival time at your final destination.
[1359] 4. Hotel reservation information while driving
[1360] server:
[1361] Search for nearby hotel suggestions based on the user's current location and destination.
[1362] Get hotel reviews, popularity, budget, and parking information.
[1363] If necessary, information regarding the reservation procedure is sent to the terminal.
[1364] Device:
[1365] Hotel information sent from the server is announced to the user by voice.
[1366] The emotion engine recognizes the user's emotions and optimizes hotel recommendations based on the results.
[1367] Receive user feedback and send the selections to the server.
[1368] Proceed with the booking process as needed.
[1369] Once the booking is completed, the user is given audio feedback.
[1370] Recalculate your route to the hotel and adjust your estimated arrival time at your final destination.
[1371] Specific examples
[1372] 1. Pre-departure settings
[1373] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[1374] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[1375] Server: Calculates the estimated arrival time and returns it to the device.
[1376] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[1377] On the device: The emotion engine recognizes emotions from the user's voice data and provides additional advice and information depending on the user's mood.
[1378] 2. Tourist information
[1379] User: "What are some nearby tourist spots?"
[1380] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[1381] User: "I'll be there."
[1382] Terminal: "Route set to museum. Arrival time 15 minutes."
[1383] On the device: The emotion engine recognizes the user's emotions and suggests additional tourist destinations that may be of interest.
[1384] 3. Restaurant Information
[1385] User: "I'm hungry, can you find a restaurant nearby?"
[1386] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[1387] User: "Yes, I'd like to make a reservation."
[1388] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[1389] On the device: The emotion engine recognizes the user's emotions and provides adaptive responses to suggested restaurant choices.
[1390] 4. Hotel Reservation Information
[1391] User: "Find a hotel tonight."
[1392] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[1393] User: "Yes, please make a reservation."
[1394] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[1395] On the device: The emotion engine takes into account the user's emotions and provides additional information and suggestions to create a relaxing atmosphere.
[1396] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set destinations and change schedules while driving, and also recognizes the user's emotions through an emotion engine, providing a more personalized and comfortable travel experience.
[1397] The processing flow will be explained below.
[1398] 1. Pre-departure settings
[1399] Step 1:
[1400] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[1401] Step 2:
[1402] The device activates the voice input engine and receives the user's voice command.
[1403] Step 3:
[1404] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1405] Step 4:
[1406] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[1407] Step 5:
[1408] The terminal transmits the extracted information to the server.
[1409] Step 6:
[1410] The server calculates the estimated arrival time and returns the result to the terminal.
[1411] Step 7:
[1412] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[1413] Step 8:
[1414] The device uses an emotion engine to recognize emotions from the user's voice data and provides additional feedback based on the user's emotions, such as "You seem to be in a good mood. Enjoy your trip!"
[1415] 2. Tourist information while driving
[1416] Step 1:
[1417] The user issues a voice command: "Tell me about nearby tourist attractions."
[1418] Step 2:
[1419] The device activates the voice input engine and receives the user's voice command.
[1420] Step 3:
[1421] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1422] Step 4:
[1423] The terminal transmits the converted text data to the server.
[1424] Step 5:
[1425] The server selects nearby tourist spots based on the user's route and current location.
[1426] Step 6:
[1427] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[1428] Step 7:
[1429] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[1430] Step 8:
[1431] The device uses an emotion engine to recognize the user's emotions and adjusts the tourist attraction suggestions accordingly: "Since you seem to be in a good mood, we'll also give you more information about this museum."
[1432] Step 9:
[1433] The user gives verbal feedback: "I'm going there."
[1434] Step 10:
[1435] The device sends the user's feedback to the server.
[1436] Step 11:
[1437] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[1438] Step 12:
[1439] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[1440] 3. Restaurant directions while driving
[1441] Step 1:
[1442] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[1443] Step 2:
[1444] The device activates the voice input engine and receives the user's voice command.
[1445] Step 3:
[1446] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1447] Step 4:
[1448] The terminal transmits the converted text data to the server.
[1449] Step 5:
[1450] The server searches for nearby restaurant options based on the user's current location and route.
[1451] Step 6:
[1452] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[1453] Step 7:
[1454] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[1455] Step 8:
[1456] The device uses an emotion engine to recognize the user's emotions and adjusts restaurant suggestions accordingly: "You seem hungry, so I highly recommend this!"
[1457] Step 9:
[1458] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[1459] Step 10:
[1460] The terminal sends the user's feedback to the server and activates the reservation call function.
[1461] Step 11:
[1462] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[1463] Step 12:
[1464] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[1465] 4. Hotel reservation information while driving
[1466] Step 1:
[1467] A user issues a voice command: "Find a hotel for tonight."
[1468] Step 2:
[1469] The device activates the voice input engine and receives the user's voice command.
[1470] Step 3:
[1471] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1472] Step 4:
[1473] The terminal transmits the converted text data to the server.
[1474] Step 5:
[1475] The server searches for nearby hotel options based on the user's current location and destination.
[1476] Step 6:
[1477] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[1478] Step 7:
[1479] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[1480] Step 8:
[1481] The device uses an emotion engine to recognize the user's emotions and adjusts hotel suggestions accordingly: "I was looking for a place with a relaxing atmosphere."
[1482] Step 9:
[1483] The user gives verbal feedback: "Yes, please book."
[1484] Step 10:
[1485] The device sends the user's feedback to the server and proceeds with the reservation process.
[1486] Step 11:
[1487] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[1488] Step 12:
[1489] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[1490] Example 2
[1491] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1492] While traveling by car, it can be difficult for drivers to safely obtain information on tourist spots, restaurants, and hotels, and to plan optimal routes and make reservations. Furthermore, conventional systems have had difficulty providing personalized information based on the driver's emotions and mood. To solve this problem, a smarter car navigation application system that supports emotion recognition is needed.
[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1494] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model for converting voice data into text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing voice feedback on the estimated arrival time to the user, and means for recognizing emotions from the user's voice data and adjusting the feedback content based on the emotions. This allows the user to safely set a destination and desired arrival time and receive personalized feedback according to their emotions.
[1495] A "voice input engine" is software or hardware that converts voice into a digital signal and analyzes it.
[1496] "Voice command" is a method by which a user gives instructions to a system through voice.
[1497] An "artificial intelligence model" is an algorithm or system that analyzes and processes data to automatically perform a specific task.
[1498] "Text data" is voice data converted into a character string format.
[1499] A "destination" is a location that a user intends to reach using the system.
[1500] "Desired arrival time" is the time at which the user wishes to arrive at the destination.
[1501] A "server" is a computer system that provides functions and data to clients over a network.
[1502] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the system.
[1503] "Feedback" is the response or information that a system provides to a user.
[1504] An "emotion engine" is an algorithm or system that analyzes and recognizes a user's emotions and adjusts responses based on the results.
[1505] "Current location" refers to the location where the user is currently located.
[1506] "Route information" is information about the route to the destination.
[1507] "Candidate tourist destinations" are tourist destination options that the system suggests to users.
[1508] "Popularity" is an indicator that shows how highly a tourist destination or facility is rated by many people.
[1509] "Word of mouth reputation" is review information based on user ratings and impressions.
[1510] "Parking information" is information about where you can park at tourist spots and facilities.
[1511] "Restaurant candidates" are restaurant options that the system suggests to users.
[1512] "Budget" is the amount the user plans to pay.
[1513] A "reservation call" is a telephone means of contact for making restaurant or hotel reservations.
[1514] "Facilities" are places and buildings that users visit, such as tourist attractions, restaurants, and hotels.
[1515] The system of the present invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience. A specific embodiment of this system is described below.
[1516] Hardware and Software Use
[1517] Device: Smartphone or dedicated car navigation device (e.g., Android device)
[1518] Server: Cloud-based backend server (e.g., AWS or Google Cloud)
[1519] software:
[1520] Speech recognition engine: Google Cloud Speech-to-Text
[1521] Artificial intelligence model: GPT-4
[1522] Emotion recognition engine: Microsoft Azure Emotion API
[1523] Route Calculation Engine
[1524] Database: tourist attractions, restaurants, and hotel information
[1525] Processing flow
[1526] 1. Voice to Text
[1527] The user enters a voice command.
[1528] The device activates a voice input engine and converts the voice data into text using Google Cloud Speech-to-Text.
[1529] The converted text data is sent to an artificial intelligence model (GPT-4) for analysis.
[1530] 2. Set your destination and desired arrival time
[1531] The terminal transmits information about the final destination and desired arrival time set by the user to the server.
[1532] The server uses a route calculation engine to calculate the estimated arrival time and returns the result to the terminal.
[1533] The terminal provides the user with audio feedback on the estimated arrival time.
[1534] An emotion engine is used to recognize emotions from the user's voice data and adjust the feedback content.
[1535] 3. Tourist information
[1536] The user inputs a voice command such as "Tell me about nearby tourist spots."
[1537] The device converts the voice command into text and sends the current location information to the server.
[1538] The server searches for potential tourist destinations and obtains popularity, reviews, photos, and parking information.
[1539] The terminal notifies the user of the received information by voice and receives feedback.
[1540] Recalculate the route to the tourist spot selected by the user and adjust the estimated arrival time.
[1541] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[1542] 4. Restaurant Information
[1543] The user enters a voice command such as "I'm hungry, find a nearby restaurant."
[1544] The device converts the voice command into text and sends the current location information to the server.
[1545] The server searches for restaurant candidates and retrieves reviews, budget, and parking information.
[1546] The terminal notifies the user of the received information by voice and receives feedback.
[1547] Make a reservation call if necessary.
[1548] Recalculate your route to the selected restaurant and adjust your estimated arrival time.
[1549] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[1550] 5. Hotel Reservation Information
[1551] The user enters the voice command "Find a hotel tonight."
[1552] The device converts the voice command into text and sends the current location information to the server.
[1553] The server searches for hotel options and retrieves reviews, popularity, budget, and parking information.
[1554] The terminal notifies the user of the received information by voice and receives feedback.
[1555] Make a reservation call if necessary.
[1556] Recalculate your route to the selected hotel and adjust your estimated arrival time.
[1557] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[1558] Prompt Sentence Examples
[1559] As a concrete example, the following prompt sentence will be used.
[1560] Pre-departure setup:
[1561] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[1562] Tourist Information:
[1563] "Tell me about nearby tourist spots."
[1564] Restaurant Information:
[1565] I'm hungry, so I'm looking for a nearby restaurant.
[1566] Hotel Reservation Information:
[1567] "Find a hotel to stay at tonight."
[1568] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set destinations and change schedules while driving. It also recognizes the user's emotions through an emotion engine, providing a more personalized and comfortable travel experience.
[1569] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1570] Step 1:
[1571] User: Enters a voice command.
[1572] Specific operation: The user gives voice instructions such as, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 p.m."
[1573] Input: Voice command
[1574] Output: Audio data
[1575] Step 2:
[1576] Device: Activates the voice input engine and converts the voice data into text.
[1577] Specific operation: Converts audio data into text data using Google Cloud Speech-to-Text.
[1578] Input: Audio data
[1579] Output: Text data
[1580] Step 3:
[1581] Terminal: The converted text data is sent to the artificial intelligence model for analysis.
[1582] What it does: It uses GPT-4 to parse text and extract the destination and desired arrival time.
[1583] Input: Text data
[1584] Output: Analysis results (destination and desired arrival time)
[1585] Step 4:
[1586] Terminal: Sends the extracted information to the server.
[1587] Specific operation: Send information to the server about the final destination "Tokyo Station" and the desired arrival time "3:00 PM".
[1588] Input: Analysis results (destination and desired arrival time)
[1589] Output: Request to server
[1590] Step 5:
[1591] Server: Calculates the estimated arrival time using a route calculation engine and returns the result to the terminal.
[1592] Specific operation: Uses a route calculation engine to calculate the optimal route to the destination and derives the estimated arrival time.
[1593] Input: Destination and desired arrival time
[1594] Output: Estimated arrival time
[1595] Step 6:
[1596] Terminal: Provides audio feedback to the user on the estimated arrival time.
[1597] Specific operation: Using a speech synthesis engine, the system notifies the user, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[1598] Input: Estimated arrival time
[1599] Output: Feedback audio
[1600] Step 7:
[1601] Terminal: Activates the emotion engine and recognizes emotions from the user's voice data.
[1602] Specific behavior: Analyzes user emotions using the Microsoft Azure Emotion API.
[1603] Input: Audio data
[1604] Output: Emotion data
[1605] Step 8:
[1606] Device: Adjusts feedback content based on user emotional data.
[1607] Specific behavior: Based on the results of the emotion engine, adjust the feedback content and provide appropriate additional information.
[1608] Input: Emotion data
[1609] Output: Adjusted feedback content
[1610] Step 9:
[1611] User: Enters voice commands for nearby tourist attractions.
[1612] Specific operation: The user gives a voice command such as "Tell me about nearby tourist spots."
[1613] Input: Voice command
[1614] Output: Audio data
[1615] Step 10:
[1616] Device: Converts voice commands into text and sends location information to the server.
[1617] Specific operation: The voice data is converted into text data using Google Cloud Speech-to-Text and sent to the server along with GPS information.
[1618] Input: Audio data
[1619] Output: Text data and current location information
[1620] Step 11:
[1621] Server: Search for potential tourist destinations and obtain popularity, reviews, photos, and parking information.
[1622] Specific operation: Filters tourist destination candidates based on the current location from the database and obtains information about each location.
[1623] Input: Current location information
[1624] Output: Tourist destination information
[1625] Step 12:
[1626] Server: Sends the acquired tourist spot information to the terminal.
[1627] Specific operation: Send tourist spot information to the terminal.
[1628] Input: Tourist destination information
[1629] Output: Response to terminal
[1630] Step 13:
[1631] Terminal: Announces the received tourist spot information to the user by voice and receives feedback.
[1632] Specific operation: Uses a speech synthesis engine to notify the user of tourist information and receive feedback.
[1633] Input: Tourist destination information
[1634] Output: Feedback speech and user feedback
[1635] Step 14:
[1636] Terminal: Recalculate the route to the selected tourist spot and adjust the estimated arrival time.
[1637] What it does: Uses the route calculation engine to calculate a new route and adjust the estimated arrival time.
[1638] Input: Selected tourist destination
[1639] Output: New route and estimated arrival time
[1640] Step 15:
[1641] Device: The emotion engine recognizes the user's emotions and optimizes the suggestions.
[1642] What it does: It uses an emotion engine to analyze the user's emotions and adjusts suggestions based on the results.
[1643] Input: Emotion data
[1644] Output: Optimized proposals
[1645] (Application example 2)
[1646] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1647] While traveling by car, users need to obtain real-time information on tourist spots, restaurants, hotels, etc., and make reservations. However, conventional car navigation systems do not provide personalized suggestions based on the user's emotions. Another issue is that it is difficult to ensure safety when users set destinations or make reservations while driving. To solve these issues, a new car navigation system that combines voice input and emotion recognition functions is needed.
[1648] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1649] In this invention, the server includes means for activating a voice input engine and receiving voice commands from the user, processing means using an artificial intelligence model to convert voice data into text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing voice feedback on the estimated arrival time to the user, and means including an emotion engine for recognizing emotions from the user's voice data and adjusting the feedback content. This makes it possible to make personalized suggestions based on the user's emotions, providing a safe and comfortable travel experience.
[1650] "Voice input engine" is a general term for devices and software for receiving voice commands from users.
[1651] "Artificial Intelligence Model" means the machine learning algorithms and techniques used to convert voice data into text.
[1652] "Destination and desired arrival time" refers to the final destination set by the user and the desired arrival time at that destination.
[1653] A "server" is a computer system that processes information entered by a user and provides the necessary information and estimated time.
[1654] The "estimated arrival time" is the estimated time required to arrive at the destination specified by the user.
[1655] The "emotion engine" is a technology and algorithm that recognizes emotions from the user's voice data and optimizes the feedback content based on the results.
[1656] "Candidate tourist destinations" are potential travel destinations suggested based on the user's current location and route information.
[1657] "Popularity, word-of-mouth reputation, photos and parking information" is a general term for ratings and reviews of tourist spots and restaurants, as well as images and information about parking.
[1658] "Feedback" refers to the response or guidance provided by the system to the user.
[1659] "Recalculating the route" means recalculating the travel route based on the user's selection and the situation.
[1660] The system of the present invention is a car navigation application system that supports travel in autonomous vehicles, and uses the following main hardware and software:
[1661] 1. Voice Input Engine
[1662] A voice input engine is a device or software that receives voice commands from users, specifically voice recognition engines such as Google Voice Recognition and Apple's Siri, allowing users to input commands using only their voice without using their hands.
[1663] 2. Artificial Intelligence Model
[1664] Artificial intelligence models are used to convert the speech data into text, such as Google Cloud Speech-to-Text API and IBM Watson Speech to Text. This process converts the speech data into text.
[1665] 3. Emotion Engine
[1666] The emotion engine is a technology and algorithm that recognizes emotions from users' voice data and optimizes feedback content. Specifically, it uses Microsoft Azure Emotion API and Affectiva's emotion recognition technology. This enables personalized suggestions to be provided based on the user's emotions.
[1667] 4. Navigation system
[1668] The navigation system uses the Google Maps API, Here Maps API, etc. to calculate the route to the destination, locate the current location, and estimate the arrival time. It also obtains information on nearby tourist attractions, restaurants, hotels, etc.
[1669] 5. Recommendation Services
[1670] The recommendation service provides information on tourist spots, restaurants, and hotels using APIs such as FourSquare and Yelp, and obtains detailed information such as popularity, reviews, photos, and parking information, and makes suggestions to users.
[1671] Specific examples
[1672] Example of user voice input and system response
[1673] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[1674] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[1675] Server: Calculates the estimated arrival time and returns it to the device.
[1676] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[1677] On the device: The emotion engine recognizes emotions from the user's voice data and provides additional advice and information depending on the user's mood.
[1678] Example of tourist information
[1679] User: "What are some nearby tourist spots?"
[1680] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[1681] User: "I'll be there."
[1682] Terminal: "Route set to museum. Arrival time 15 minutes."
[1683] On the device: The emotion engine recognizes the user's emotions and suggests additional tourist destinations that may be of interest.
[1684] This allows users to have a comfortable and personalized travel experience.
[1685] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1686] Step 1:
[1687] Receiving voice commands
[1688] The user issues a voice command.
[1689] The terminal activates a voice input engine and receives voice commands from the user.
[1690] Input: User's voice command
[1691] Output: Audio data
[1692] Step 2:
[1693] Converting audio data to text
[1694] The device converts the voice data into text using an artificial intelligence model (e.g., Google Cloud Speech-to-Text API).
[1695] Input: Audio data
[1696] Output: Text data
[1697] Step 3:
[1698] Extracting destination and desired arrival time
[1699] The device uses natural language processing technology to extract the destination and desired arrival time from the text data.
[1700] Input: Text data
[1701] Output: Extracted information about destination and desired arrival time
[1702] Step 4:
[1703] Destination and desired arrival time sent to server
[1704] The terminal transmits the extracted information to the server and requests it to calculate the estimated arrival time.
[1705] Input: Extract information about destination and desired arrival time
[1706] Output: Request to calculate estimated arrival time
[1707] Step 5:
[1708] Calculating and receiving estimated arrival times
[1709] The server uses a navigation system (e.g., Google Maps API) to calculate the estimated arrival time and sends this information back to the device.
[1710] Input: Extract information about destination and desired arrival time
[1711] Output: Estimated arrival time
[1712] Step 6:
[1713] emotion recognition
[1714] The device uses an emotion engine (e.g., Microsoft Azure Emotion API) to recognize emotions from the user's voice data.
[1715] Input: User's voice data
[1716] Output: Emotion data
[1717] Step 7:
[1718] Personalized Feedback
[1719] The terminal adjusts the feedback content of the estimated arrival time based on the emotion data and provides the feedback to the user by voice.
[1720] Input: Estimated arrival time, emotion data
[1721] Output: Audio feedback
[1722] Step 8:
[1723] Search and submit tourist destination suggestions
[1724] The server searches for nearby tourist spots based on the current location and route information using a recommendation service (e.g., FourSquare API) and sends detailed information to the device.
[1725] Input: Current location and route information
[1726] Output: Tourist destination information
[1727] Step 9:
[1728] To announce tourist destination information and receive feedback
[1729] The terminal notifies the user of tourist spot information by voice and receives feedback from the user.
[1730] Input: Tourist destination information
[1731] Output: User feedback
[1732] Step 10:
[1733] Recalculate route to selected tourist spot
[1734] The server recalculates the route to the tourist spot selected by the user and adjusts the estimated arrival time.
[1735] Input: User feedback
[1736] Output: Recalculated route and estimated arrival time
[1737] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1738] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1739] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1740] [Third embodiment]
[1741] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1742] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1743] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1744] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1745] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1746] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1747] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1748] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1749] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1750] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1751] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1752] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1753] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time.
[1754] Functionality Overview
[1755] 1. Pre-departure settings
[1756] Device:
[1757] It starts the voice input engine and receives the user's voice commands.
[1758] The voice data is sent to an artificial intelligence model and converted into text.
[1759] Information about the final destination and desired arrival time set by the user is sent to the server.
[1760] The destination information returned from the server is checked and audio feedback is given to the user.
[1761] 2. Tourist information while driving
[1762] server:
[1763] Potential tourist spots in the area are selected based on the user's route and current location.
[1764] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[1765] This information is sent to the terminal.
[1766] Device:
[1767] Tourist information from the server is announced to the user by voice.
[1768] It receives user feedback and sends it to the server.
[1769] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[1770] 3. Restaurant directions while driving
[1771] server:
[1772] Select nearby restaurant options based on the user's current location and route.
[1773] Obtain information such as restaurant reviews, budget, and whether parking is available.
[1774] If necessary, the reservation telephone number is transmitted to the terminal.
[1775] Device:
[1776] Restaurant information sent from the server is announced to the user by voice.
[1777] Receive user feedback and send selections to the server.
[1778] Make a reservation if necessary.
[1779] It provides the functionality to receive user feedback, send it to the server, and make reservation calls.
[1780] 4. Hotel reservation information while driving
[1781] server:
[1782] Select nearby hotel options based on the user's current location and destination.
[1783] Obtain hotel reviews, popularity, budget, parking information, etc.
[1784] If necessary, information for reservation procedures is sent to the terminal.
[1785] Device:
[1786] Hotel information sent from the server is announced to the user by voice.
[1787] Receive user feedback and send the selection results to the server.
[1788] Proceed with the booking process as needed.
[1789] Once the booking is completed, the user is given audio feedback.
[1790] Recalculate the route to the user's selected hotel and adjust the estimated arrival time.
[1791] Specific examples
[1792] 1. Pre-departure settings
[1793] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[1794] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[1795] Server: Calculates the estimated arrival time and returns it to the device.
[1796] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[1797] 2. Tourist information
[1798] User: "What are some nearby tourist spots?"
[1799] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[1800] User: "I'll be there."
[1801] Terminal: "Route set to museum. Arrival time 15 minutes."
[1802] 3. Restaurant Information
[1803] User: "I'm hungry, can you find a restaurant nearby?"
[1804] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[1805] User: "Yes, I'd like to make a reservation."
[1806] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[1807] 4. Hotel Reservation Information
[1808] User: "Find a hotel tonight."
[1809] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[1810] User: "Yes, please make a reservation."
[1811] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[1812] In this way, the system of the present invention allows all operations to be performed by voice, allowing you to safely set your destination or change your schedule while driving, providing a comfortable travel experience.
[1813] The processing flow will be explained below.
[1814] 1. Pre-departure settings
[1815] Step 1:
[1816] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[1817] Step 2:
[1818] The device activates the voice input engine and receives the user's voice command.
[1819] Step 3:
[1820] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1821] Step 4:
[1822] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[1823] Step 5:
[1824] The terminal transmits the extracted information to the server.
[1825] Step 6:
[1826] The server calculates the estimated arrival time and returns the result to the terminal.
[1827] Step 7:
[1828] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[1829] 2. Tourist information while driving
[1830] Step 1:
[1831] The user issues a voice command: "Tell me about nearby tourist attractions."
[1832] Step 2:
[1833] The device activates the voice input engine and receives the user's voice command.
[1834] Step 3:
[1835] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1836] Step 4:
[1837] The terminal transmits the converted text data to the server.
[1838] Step 5:
[1839] The server selects nearby tourist spots based on the user's route and current location.
[1840] Step 6:
[1841] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[1842] Step 7:
[1843] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[1844] Step 8:
[1845] The user gives verbal feedback: "I'm going there."
[1846] Step 9:
[1847] The device sends the user's feedback to the server.
[1848] Step 10:
[1849] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[1850] Step 11:
[1851] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[1852] 3. Restaurant directions while driving
[1853] Step 1:
[1854] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[1855] Step 2:
[1856] The device activates the voice input engine and receives the user's voice command.
[1857] Step 3:
[1858] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1859] Step 4:
[1860] The terminal transmits the converted text data to the server.
[1861] Step 5:
[1862] The server searches for nearby restaurant options based on the user's current location and route.
[1863] Step 6:
[1864] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[1865] Step 7:
[1866] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[1867] Step 8:
[1868] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[1869] Step 9:
[1870] The terminal sends the user's feedback to the server and activates the reservation call function.
[1871] Step 10:
[1872] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[1873] Step 11:
[1874] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[1875] 4. Hotel reservation information while driving
[1876] Step 1:
[1877] A user issues a voice command: "Find a hotel for tonight."
[1878] Step 2:
[1879] The device activates the voice input engine and receives the user's voice command.
[1880] Step 3:
[1881] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[1882] Step 4:
[1883] The terminal transmits the converted text data to the server.
[1884] Step 5:
[1885] The server searches for nearby hotel options based on the user's current location and destination.
[1886] Step 6:
[1887] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[1888] Step 7:
[1889] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[1890] Step 8:
[1891] The user gives verbal feedback: "Yes, please book."
[1892] Step 9:
[1893] The device sends the user's feedback to the server and proceeds with the reservation process.
[1894] Step 10:
[1895] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[1896] Step 11:
[1897] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[1898] Example 1
[1899] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1900] Current car navigation systems have the challenge of making it difficult for users to safely and efficiently set destinations and adjust schedules while driving. Furthermore, they lack the functionality to provide comprehensive, real-time information on tourist spots, restaurants, hotels, and other information, preventing users from making appropriate choices quickly. This often results in a loss of convenience and safety while driving.
[1901] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1902] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for suggesting tourist spots, restaurants, and hotels based on the user's current location and destination and acquiring information, means for notifying the user of the acquired information audio-visually and receiving feedback, means for recalculating a route to the selected destination and adjusting the estimated arrival time, and means for making reservations if necessary. This allows destination setting and schedule adjustments to be performed safely and efficiently even while driving, providing a comfortable travel experience.
[1903] A "voice input engine" is a device or software that receives a user's voice commands and converts them into digital data.
[1904] "Voice Data" means human speech information converted into digital form by a speech input engine.
[1905] "Artificial intelligence model" refers to technologies such as machine learning algorithms and neural networks used to convert voice data into text data.
[1906] "Text data" is character string information converted from voice data by an artificial intelligence model.
[1907] "Destination" is information indicating the place or location to which the user wishes to travel.
[1908] The "desired arrival time" is information indicating a specific time at which the user desires to arrive at the destination.
[1909] A "server" is a central system that receives requests on a computer network, processes data, and sends and receives information.
[1910] "Estimated arrival time" is information indicating the estimated arrival time from the current location to the destination.
[1911] "Tourist destination" refers to tourist spots and famous places that users aim to visit.
[1912] "Restaurant" means an eating and drinking establishment selected by a patron for dining.
[1913] "Hotel" means the accommodation facility selected for the Guest's stay.
[1914] "Means of acquisition" refers to the methods and technologies used to collect and receive the required information or data.
[1915] "Means of receiving feedback" refers to the methods and techniques for receiving responses or reactions from users.
[1916] "Means for recalculating a route" refers to a method or technology for recalculating a new route based on specified conditions.
[1917] "Means for completing reservation procedures" refers to the methods and technologies used to make reservations for the service selected by the user (such as restaurant reservations or hotel reservations).
[1918] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, adjusting routes and arrival times in real time.
[1919] The main components of the system are the "terminals" used by users and the "servers" that process data. The specific configuration and operation are explained below.
[1920] The system's devices, which include smartphones and car navigation systems, are equipped with a voice input engine that receives users' voice commands and converts the voice data into a digital format. This digital voice data is then converted into text data using voice recognition software such as Google Cloud Speech-to-Text.
[1921] The converted text data is sent to the server and used to extract the destination and desired arrival time. The server calculates the estimated arrival time based on this information and returns the result to the terminal. The terminal then provides this information as feedback to the user via voice. Specifically, it notifies the user by saying something like, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[1922] In the case of tourist attraction guidance while driving, the device sends the user's current location and route information to the server, and the server searches for nearby tourist attractions. The server obtains the tourist attraction's popularity, reviews, photos, and parking information and sends them to the device. The device then announces this information to the user by voice (e.g., "Yes, there's a famous museum nearby. It has a 4.5 rating and there's parking available. What would you like to do?"). Based on the user's feedback, the server recalculates the route and calculates a new estimated arrival time.
[1923] Regarding restaurant guidance, the system searches for nearby restaurants based on the current location and route, and obtains information such as reviews, budget, and parking information. If necessary, it obtains a reservation phone number and sends it to the terminal. The terminal then announces this information to the user by voice (e.g., "There's a highly rated Chinese restaurant nearby. The reviews are 4.7, and your budget is around 1,500 yen for lunch. Would you like to make a reservation?") and proceeds with the reservation process based on the user's feedback.
[1924] Additionally, the hotel guide searches for nearby hotels based on the current location and destination, and obtains information such as reviews, popularity, budget, and parking information. If a reservation is required, the server sends the necessary information to the terminal, which then proceeds with the reservation process (e.g., "Looking for a hotel to stay at tonight," "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?").
[1925] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set their destination or change their schedule while driving, providing a comfortable travel experience.
[1926] Specific prompt examples
[1927] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[1928] "Tell me about nearby tourist spots."
[1929] I'm hungry, so I'm looking for a nearby restaurant.
[1930] "Find a hotel to stay at tonight."
[1931] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1932] The flow of this system's program processing
[1933] 1. Pre-departure settings
[1934] Step 1:
[1935] Subject: User
[1936] Description: The user speaks a voice command into the device.
[1937] Specific actions: For example, say, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 PM."
[1938] Step 2:
[1939] Subject: Device
[1940] Description: Starts the voice input engine and receives voice commands.
[1941] Input: Voice command from the user
[1942] Output: Digital audio data
[1943] What it does: The voice input engine receives the voice and converts it into a digital format.
[1944] Step 3:
[1945] Subject: Device
[1946] Description: Uses artificial intelligence models to convert voice data into text.
[1947] Input: Digital audio data
[1948] Output: Text data
[1949] What it does: Converts speech to text using a service like Google Cloud Speech-to-Text.
[1950] Step 4:
[1951] Subject: Device
[1952] Description: Sends the converted text data to the server.
[1953] Input: Text data (e.g., "Set a route to Tokyo Station, with a desired arrival time of 3:00 PM.")
[1954] Output: Request to server
[1955] Specific operation: Sends an HTTP request to the server.
[1956] Step 5:
[1957] Subject: Server
[1958] Description: Extracts the destination and desired arrival time and calculates the estimated arrival time.
[1959] Input: Text data
[1960] Output: Estimated arrival time
[1961] Specific operation: Uses natural language processing technology to analyze text and calculate arrival times by referencing traffic information and road conditions.
[1962] Step 6:
[1963] Subject: Server
[1964] Description: Sends the calculation result back to the terminal.
[1965] Input: Estimated arrival time
[1966] Output: Feedback data
[1967] Specific operation: Returns the calculation result to the terminal.
[1968] Step 7:
[1969] Subject: Device
[1970] Description: Provides audio feedback data from the server to the user.
[1971] Input: Feedback data
[1972] Output: Audio notification to the user
[1973] Specific operation: The text data is converted into speech, and a message such as "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM" is spoken to the user.
[1974] 2. Tourist information while driving
[1975] Step 1:
[1976] Subject: User
[1977] Description: Request "Tell me about nearby tourist spots."
[1978] Specific action: Speak a voice command into the device.
[1979] Step 2:
[1980] Subject: Device
[1981] Description: Converts voice commands into text and sends it to the server.
[1982] Input: Voice command
[1983] Output: Text data
[1984] Specific operation: The audio is converted into text using Google Cloud Speech-to-Text or similar and sent to the server.
[1985] Step 3:
[1986] Subject: Server
[1987] Description: Searches for nearby tourist spots based on the user's current location and route.
[1988] Input: Current location and route information
[1989] Output: List of tourist destination candidates
[1990] Specific behavior: Generate a list of tourist destinations by retrieving information from a database or external API.
[1991] Step 4:
[1992] Subject: Server
[1993] Description: Get tourist attraction popularity, reviews, photos, and parking information.
[1994] Input: Tourist destination candidate list
[1995] Output: Detailed information
[1996] Specific actions: Collect and list detailed information about each tourist destination.
[1997] Step 5:
[1998] Subject: Server
[1999] Description: Sends the acquired information to the device.
[2000] Input: More information
[2001] Output: Feedback data
[2002] Specific operation: Sends collected information to the device.
[2003] Step 6:
[2004] Subject: Device
[2005] Description: Presents information to the user audibly.
[2006] Input: Feedback data
[2007] Output: Audio notification to the user
[2008] Action: The audio message is, "There's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[2009] Step 7:
[2010] Subject: User
[2011] Description: Gives feedback saying "There you go."
[2012] Specific actions: Select a tourist spot specified by voice.
[2013] Step 8:
[2014] Subject: Device
[2015] Description: Recalculates route based on user selection.
[2016] Input: User's choice
[2017] Output: Recalculated route
[2018] Specific behavior: Calculate a new route and send a notification such as "Route to the museum has been set. It will take 15 minutes to arrive."
[2019] 3. Restaurant directions while driving
[2020] Step 1:
[2021] Subject: User
[2022] Description: "I'm hungry, find me a nearby restaurant."
[2023] Specific action: Speak a voice command into the device.
[2024] Step 2:
[2025] Subject: Device
[2026] Description: Sends a voice command to the server.
[2027] Input: Voice command
[2028] Output: Text data
[2029] Specific operation: Converts speech into text and sends it to the server.
[2030] Step 3:
[2031] Subject: Server
[2032] Description: Find nearby restaurant suggestions based on your current location and route.
[2033] Input: Current location and route information
[2034] Output: Restaurant candidate list
[2035] Specific operation: Retrieves restaurant information from a database or external API and creates a list.
[2036] Step 4:
[2037] Subject: Server
[2038] Description: Get detailed information like reviews, budget, parking info, etc.
[2039] Input: Restaurant candidate list
[2040] Output: Detailed information
[2041] What it does: Collect and list detailed information about each restaurant.
[2042] Step 5:
[2043] Subject: Server
[2044] Description: Sends the acquired information to the device.
[2045] Input: More information
[2046] Output: Feedback data
[2047] Specific operation: Sends collected information to the device.
[2048] Step 6:
[2049] Subject: Device
[2050] Description: Presents information to the user audibly.
[2051] Input: Feedback data
[2052] Output: Audio notification to the user
[2053] Specific action: Say, "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[2054] Step 7:
[2055] Subject: User
[2056] Instructions: Answer "Yes, I'd like to make a reservation as well."
[2057] Specific action: Indicate your intention to make a reservation by voice.
[2058] Step 8:
[2059] Subject: Device
[2060] Description: Proceed with the booking process.
[2061] Input: User's booking request
[2062] Output: Reservation completion notification
[2063] Specific operation: Make a reservation using the OpenTable API or similar, and notify the customer, "Your reservation has been arranged. You will arrive in 10 minutes."
[2064] 4. Hotel reservation information while driving
[2065] Step 1:
[2066] Subject: User
[2067] Description: "Find me a hotel tonight."
[2068] Specific action: Speak a voice command into the device.
[2069] Step 2:
[2070] Subject: Device
[2071] Description: Sends a voice command to the server.
[2072] Input: Voice command
[2073] Output: Text data
[2074] Specific operation: Converts voice commands into text and sends it to the server.
[2075] Step 3:
[2076] Subject: Server
[2077] Description: Search for nearby hotel suggestions based on your current location and destination.
[2078] Input: Current location and destination information
[2079] Output: Hotel candidate list
[2080] Specific operation: Retrieves hotel information from a database or external API and creates a list.
[2081] Step 4:
[2082] Subject: Server
[2083] Description: Get detailed information like reviews, popularity, budget, parking information, and more.
[2084] Input: Hotel candidate list
[2085] Output: Detailed information
[2086] What to do: Collect and list detailed information about each hotel.
[2087] Step 5:
[2088] Subject: Server
[2089] Description: Sends the acquired information to the device.
[2090] Input: More information
[2091] Output: Feedback data
[2092] Specific operation: Sends collected information to the device.
[2093] Step 6:
[2094] Subject: Device
[2095] Description: Presents information to the user audibly.
[2096] Input: Feedback data
[2097] Output: Audio notification to the user
[2098] Action: The voice will say, "There's a highly rated hotel nearby with a 4.8 rating and parking. Would you like to make a reservation?"
[2099] Step 7:
[2100] Subject: User
[2101] Instructions: "Yes, please make a reservation."
[2102] Specific action: Indicate your intention to make a reservation by voice.
[2103] Step 8:
[2104] Subject: Device
[2105] Description: Proceed with the booking process.
[2106] Input: User's booking request
[2107] Output: Reservation completion notification
[2108] Specific behavior: Make a reservation using Expedia API or Booking.com API and notify the user, "Hotel has been booked. Estimated arrival time is 6:00 PM."
[2109] (Application example 1)
[2110] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2111] When traveling by car, there is a demand for voice control to select and reserve destinations, tourist attractions, restaurants, and hotels along the way, and for integration with autonomous driving systems to provide a safe and comfortable travel experience. However, with conventional systems, it can be difficult to provide sufficient information, complete reservations, and adjust routes using voice control alone. Furthermore, there is a lack of integration with autonomous driving systems, and manual operation by the user is often required. This can make operations while driving cumbersome, potentially compromising safety and efficiency.
[2112] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2113] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for adding control means compatible with the autonomous driving system, and means for automating automatic route adjustment to selected facilities and reservation procedures. This makes it possible for an autonomous vehicle to utilize voice operation and AI technology to select and reserve tourist spots, restaurants, and hotels in real time based on user instructions, and to automatically adjust the route and provide feedback on the arrival time.
[2114] A "voice input engine" is a device or software that receives voice commands from a user.
[2115] "Artificial intelligence model" is a machine learning algorithm used to convert voice data into text.
[2116] "Text data" is voice data converted into text format.
[2117] A "server" is a computer system that provides services to clients over a computer network.
[2118] The "destination" is the destination point set by the user.
[2119] "Desired arrival time" is the time the user desires to arrive at the destination.
[2120] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the server.
[2121] An "autonomous driving system" is a technology that enables vehicles to perform driving operations autonomously.
[2122] "Route adjustment" refers to recalculating the route to a selected destination.
[2123] "Tourist destination candidates" is a list of tourist destinations suggested to the user.
[2124] "Popularity" is an indicator that shows the evaluation of tourist destinations and facilities.
[2125] "Word of mouth" refers to reviews and feedback from users.
[2126] "Photos" are image data that provide visual images of tourist spots and facilities.
[2127] "Parking information" is information about parking spaces at tourist spots and facilities.
[2128] "Restaurant candidates" is a list of restaurants suggested to the user.
[2129] The "budget" is an estimate of the cost for the user to use the service.
[2130] "Reservation phone" refers to a means of making a reservation for a facility by telephone.
[2131] In one embodiment of the present invention, the system first activates a voice input engine to receive a user's voice command. The voice input engine may be, for example, the Google Speech Recognition API or a similar voice recognition engine. The device that receives the voice data converts it into text data using an artificial intelligence model (e.g., a generative AI model).
[2132] Natural language processing (NLP) techniques are used to extract the destination and desired arrival time from the converted text data. The text data is analyzed to extract specific information (in this case, the destination and desired arrival time). This information is sent to a server, which calculates the estimated arrival time. The calculation result is sent back to the device, which then communicates the result to the user as voice feedback.
[2133] By adding control means compatible with the autonomous driving system, it becomes possible to automatically adjust routes to selected facilities and make reservations. Specifically, the terminal provides route information to the autonomous driving system based on information received from the server and uses an API for reservation procedures. This allows users to set destinations and make reservations using only voice commands without manual operation.
[2134] For example, if a user voice-inputs "My destination is Tokyo Station, and I would like to arrive at 3:00 PM," the device converts the voice data into text data and extracts the destination and desired arrival time. This information is sent to the server, which calculates the estimated arrival time and receives feedback. Furthermore, if the user issues the command "Find a nearby restaurant," the server searches for restaurant candidates based on the current location and route information, obtains reviews and budget information, and sends it to the device. The device notifies the user of this by voice, and if the user responds "Make a reservation," the device will automatically complete the reservation procedure.
[2135] An example of a prompt sentence might be:
[2136] Please let us know your destination and desired arrival time.
[2137] Could you tell me about nearby tourist spots?
[2138] Find a restaurant near you.
[2139] Find a hotel to stay in tonight.
[2140] In this way, the present invention combines voice control with an automated driving system to provide users with a safe and convenient travel experience.
[2141] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2142] Step 1:
[2143] The device starts a voice input engine to receive the user's voice command. The voice input engine used here is a speech recognition engine such as the Google Speech Recognition API. The voice input engine captures the user's voice data and processes it as digital voice data.
[2144] Input: User's voice command
[2145] Output: Digital audio data
[2146] Step 2:
[2147] The voice data received by the device is converted into text data using a generative AI model (a speech recognition algorithm). This process involves analyzing the voice signal and generating the corresponding text.
[2148] Input: Digital audio data
[2149] Output: Text data
[2150] Step 3:
[2151] The device extracts the destination and desired arrival time from the generated text data, using natural language processing (NLP) techniques to identify keywords and phrases related to the destination and desired arrival time.
[2152] Input: Text data
[2153] Output: Destination and desired arrival time information
[2154] Step 4:
[2155] The device sends the extracted information to a server, which calculates and returns an estimated arrival time. The server then uses a map database and traffic data to calculate a route from the current location to the destination.
[2156] Input: Destination and desired arrival time information
[2157] Output: Estimated arrival time
[2158] Step 5:
[2159] The terminal receives the estimated arrival time from the server and provides the user with audio feedback. Here, a synthetic speech engine (e.g., pyttsx3) is used to generate speech from text and communicate it to the user.
[2160] Input: Estimated arrival time
[2161] Output: Audio feedback
[2162] Step 6:
[2163] The user inputs an additional voice command (e.g., "Tell me about nearby tourist spots," "Find nearby restaurants," etc.). Based on this prompt, the device queries the server.
[2164] Input: User's additional voice command
[2165] Output: prompt statement
[2166] Step 7:
[2167] The server searches for potential tourist spots and restaurants based on the user's current location and route information, and then retrieves and sends the information to the device. The retrieved information includes popularity, reviews, photos, parking information, budget, etc.
[2168] Input: User's current location and route information
[2169] Output: Information on tourist spots and restaurants
[2170] Step 8:
[2171] The terminal announces the information received from the server to the user by voice and receives user feedback. When the user makes a selection, the selection is sent to the server, which then processes the new route and reservation.
[2172] Input: Information about tourist attractions and restaurants
[2173] Output: Audio feedback and user selection results
[2174] Step 9:
[2175] Based on the user's feedback, the device issues instructions to the autonomous driving system, automatically adjusting the route to the selected facility and making reservations, allowing the user to automatically head to the next destination without manual intervention.
[2176] Input: User selection
[2177] Output: Instructions to the autonomous driving system and reservation procedures
[2178] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2179] The system of this invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice commands, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience.
[2180] Functionality Overview
[2181] 1. Pre-departure settings
[2182] Device:
[2183] It starts the voice input engine and receives the user's voice commands.
[2184] The voice data is sent to an artificial intelligence model and converted into text.
[2185] Information about the final destination and desired arrival time set by the user is sent to the server.
[2186] The destination information returned from the server is checked and audio feedback is given to the user.
[2187] The emotion engine recognizes emotions from the user's voice data and adjusts the feedback content according to the user's emotions.
[2188] 2. Tourist information while driving
[2189] server:
[2190] Potential tourist spots in the area are selected based on the user's route and current location.
[2191] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[2192] This information is sent to the terminal.
[2193] Device:
[2194] Tourist information from the server is announced to the user by voice.
[2195] The emotion engine recognizes the user's emotions and optimizes tourist destination suggestions based on the results.
[2196] It receives user feedback via voice and sends it to the server.
[2197] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[2198] 3. Restaurant directions while driving
[2199] server:
[2200] Select nearby restaurant options based on the user's current location and route.
[2201] Get restaurant reviews, budget and parking information.
[2202] If necessary, the reservation telephone number is transmitted to the terminal.
[2203] Device:
[2204] Restaurant information sent from the server is announced to the user by voice.
[2205] The emotion engine recognizes the user's emotions and optimizes restaurant suggestions based on the results.
[2206] Receive user feedback and send selections to the server.
[2207] Make a reservation if necessary.
[2208] Once the booking is completed, the user is given audio feedback.
[2209] Recalculate your route to the restaurant and adjust your estimated arrival time at your final destination.
[2210] 4. Hotel reservation information while driving
[2211] server:
[2212] Search for nearby hotel suggestions based on the user's current location and destination.
[2213] Get hotel reviews, popularity, budget, and parking information.
[2214] If necessary, information regarding the reservation procedure is sent to the terminal.
[2215] Device:
[2216] Hotel information sent from the server is announced to the user by voice.
[2217] The emotion engine recognizes the user's emotions and optimizes hotel recommendations based on the results.
[2218] Receive user feedback and send the selections to the server.
[2219] Proceed with the booking process as needed.
[2220] Once the booking is completed, the user is given audio feedback.
[2221] Recalculate your route to the hotel and adjust your estimated arrival time at your final destination.
[2222] Specific examples
[2223] 1. Pre-departure settings
[2224] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[2225] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[2226] Server: Calculates the estimated arrival time and returns it to the device.
[2227] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[2228] On the device: The emotion engine recognizes emotions from the user's voice data and provides additional advice and information depending on the user's mood.
[2229] 2. Tourist information
[2230] User: "What are some nearby tourist spots?"
[2231] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[2232] User: "I'll be there."
[2233] Terminal: "Route set to museum. Arrival time 15 minutes."
[2234] On the device: The emotion engine recognizes the user's emotions and suggests additional tourist destinations that may be of interest.
[2235] 3. Restaurant Information
[2236] User: "I'm hungry, can you find a restaurant nearby?"
[2237] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[2238] User: "Yes, I'd like to make a reservation."
[2239] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[2240] On the device: The emotion engine recognizes the user's emotions and provides adaptive responses to suggested restaurant choices.
[2241] 4. Hotel Reservation Information
[2242] User: "Find a hotel tonight."
[2243] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[2244] User: "Yes, please make a reservation."
[2245] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[2246] On the device: The emotion engine takes into account the user's emotions and provides additional information and suggestions to create a relaxing atmosphere.
[2247] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set destinations and change schedules while driving, and also recognizes the user's emotions through an emotion engine, providing a more personalized and comfortable travel experience.
[2248] The processing flow will be explained below.
[2249] 1. Pre-departure settings
[2250] Step 1:
[2251] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[2252] Step 2:
[2253] The device activates the voice input engine and receives the user's voice command.
[2254] Step 3:
[2255] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2256] Step 4:
[2257] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[2258] Step 5:
[2259] The terminal transmits the extracted information to the server.
[2260] Step 6:
[2261] The server calculates the estimated arrival time and returns the result to the terminal.
[2262] Step 7:
[2263] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[2264] Step 8:
[2265] The device uses an emotion engine to recognize emotions from the user's voice data and provides additional feedback based on the user's emotions, such as "You seem to be in a good mood. Enjoy your trip!"
[2266] 2. Tourist information while driving
[2267] Step 1:
[2268] The user issues a voice command: "Tell me about nearby tourist attractions."
[2269] Step 2:
[2270] The device activates the voice input engine and receives the user's voice command.
[2271] Step 3:
[2272] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2273] Step 4:
[2274] The terminal transmits the converted text data to the server.
[2275] Step 5:
[2276] The server selects nearby tourist spots based on the user's route and current location.
[2277] Step 6:
[2278] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[2279] Step 7:
[2280] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[2281] Step 8:
[2282] The device uses an emotion engine to recognize the user's emotions and adjusts the tourist attraction suggestions accordingly: "Since you seem to be in a good mood, we'll also give you more information about this museum."
[2283] Step 9:
[2284] The user gives verbal feedback: "I'm going there."
[2285] Step 10:
[2286] The device sends the user's feedback to the server.
[2287] Step 11:
[2288] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[2289] Step 12:
[2290] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[2291] 3. Restaurant directions while driving
[2292] Step 1:
[2293] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[2294] Step 2:
[2295] The device activates the voice input engine and receives the user's voice command.
[2296] Step 3:
[2297] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2298] Step 4:
[2299] The terminal transmits the converted text data to the server.
[2300] Step 5:
[2301] The server searches for nearby restaurant options based on the user's current location and route.
[2302] Step 6:
[2303] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[2304] Step 7:
[2305] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[2306] Step 8:
[2307] The device uses an emotion engine to recognize the user's emotions and adjusts restaurant suggestions accordingly: "You seem hungry, so I highly recommend this!"
[2308] Step 9:
[2309] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[2310] Step 10:
[2311] The terminal sends the user's feedback to the server and activates the reservation call function.
[2312] Step 11:
[2313] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[2314] Step 12:
[2315] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[2316] 4. Hotel reservation information while driving
[2317] Step 1:
[2318] A user issues a voice command: "Find a hotel for tonight."
[2319] Step 2:
[2320] The device activates the voice input engine and receives the user's voice command.
[2321] Step 3:
[2322] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2323] Step 4:
[2324] The terminal transmits the converted text data to the server.
[2325] Step 5:
[2326] The server searches for nearby hotel options based on the user's current location and destination.
[2327] Step 6:
[2328] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[2329] Step 7:
[2330] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[2331] Step 8:
[2332] The device uses an emotion engine to recognize the user's emotions and adjusts hotel suggestions accordingly: "I was looking for a place with a relaxing atmosphere."
[2333] Step 9:
[2334] The user gives verbal feedback: "Yes, please book."
[2335] Step 10:
[2336] The device sends the user's feedback to the server and proceeds with the reservation process.
[2337] Step 11:
[2338] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[2339] Step 12:
[2340] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[2341] Example 2
[2342] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2343] While traveling by car, it can be difficult for drivers to safely obtain information on tourist spots, restaurants, and hotels, and to plan optimal routes and make reservations. Furthermore, conventional systems have had difficulty providing personalized information based on the driver's emotions and mood. To solve this problem, a smarter car navigation application system that supports emotion recognition is needed.
[2344] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2345] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model for converting voice data into text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing voice feedback on the estimated arrival time to the user, and means for recognizing emotions from the user's voice data and adjusting the feedback content based on the emotions. This allows the user to safely set a destination and desired arrival time and receive personalized feedback according to their emotions.
[2346] A "voice input engine" is software or hardware that converts voice into a digital signal and analyzes it.
[2347] "Voice command" is a method by which a user gives instructions to a system through voice.
[2348] An "artificial intelligence model" is an algorithm or system that analyzes and processes data to automatically perform a specific task.
[2349] "Text data" is voice data converted into a character string format.
[2350] A "destination" is a location that a user intends to reach using the system.
[2351] "Desired arrival time" is the time at which the user wishes to arrive at the destination.
[2352] A "server" is a computer system that provides functions and data to clients over a network.
[2353] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the system.
[2354] "Feedback" is the response or information that a system provides to a user.
[2355] An "emotion engine" is an algorithm or system that analyzes and recognizes a user's emotions and adjusts responses based on the results.
[2356] "Current location" refers to the location where the user is currently located.
[2357] "Route information" is information about the route to the destination.
[2358] "Candidate tourist destinations" are tourist destination options that the system suggests to users.
[2359] "Popularity" is an indicator that shows how highly a tourist destination or facility is rated by many people.
[2360] "Word of mouth reputation" is review information based on user ratings and impressions.
[2361] "Parking information" is information about where you can park at tourist spots and facilities.
[2362] "Restaurant candidates" are restaurant options that the system suggests to users.
[2363] "Budget" is the amount the user plans to pay.
[2364] A "reservation call" is a telephone means of contact for making restaurant or hotel reservations.
[2365] "Facilities" are places and buildings that users visit, such as tourist attractions, restaurants, and hotels.
[2366] The system of the present invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience. A specific embodiment of this system is described below.
[2367] Hardware and Software Use
[2368] Device: Smartphone or dedicated car navigation device (e.g., Android device)
[2369] Server: Cloud-based backend server (e.g., AWS or Google Cloud)
[2370] software:
[2371] Speech recognition engine: Google Cloud Speech-to-Text
[2372] Artificial intelligence model: GPT-4
[2373] Emotion recognition engine: Microsoft Azure Emotion API
[2374] Route Calculation Engine
[2375] Database: tourist attractions, restaurants, and hotel information
[2376] Processing flow
[2377] 1. Voice to Text
[2378] The user enters a voice command.
[2379] The device activates a voice input engine and converts the voice data into text using Google Cloud Speech-to-Text.
[2380] The converted text data is sent to an artificial intelligence model (GPT-4) for analysis.
[2381] 2. Set your destination and desired arrival time
[2382] The terminal transmits information about the final destination and desired arrival time set by the user to the server.
[2383] The server uses a route calculation engine to calculate the estimated arrival time and returns the result to the terminal.
[2384] The terminal provides the user with audio feedback on the estimated arrival time.
[2385] An emotion engine is used to recognize emotions from the user's voice data and adjust the feedback content.
[2386] 3. Tourist information
[2387] The user inputs a voice command such as "Tell me about nearby tourist spots."
[2388] The device converts the voice command into text and sends the current location information to the server.
[2389] The server searches for potential tourist destinations and obtains popularity, reviews, photos, and parking information.
[2390] The terminal notifies the user of the received information by voice and receives feedback.
[2391] Recalculate the route to the tourist spot selected by the user and adjust the estimated arrival time.
[2392] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[2393] 4. Restaurant Information
[2394] The user enters a voice command such as "I'm hungry, find a nearby restaurant."
[2395] The device converts the voice command into text and sends the current location information to the server.
[2396] The server searches for restaurant candidates and retrieves reviews, budget, and parking information.
[2397] The terminal notifies the user of the received information by voice and receives feedback.
[2398] Make a reservation call if necessary.
[2399] Recalculate your route to the selected restaurant and adjust your estimated arrival time.
[2400] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[2401] 5. Hotel Reservation Information
[2402] The user enters the voice command "Find a hotel tonight."
[2403] The device converts the voice command into text and sends the current location information to the server.
[2404] The server searches for hotel options and retrieves reviews, popularity, budget, and parking information.
[2405] The terminal notifies the user of the received information by voice and receives feedback.
[2406] Make a reservation call if necessary.
[2407] Recalculate your route to the selected hotel and adjust your estimated arrival time.
[2408] The emotion engine recognizes the user's emotions and optimizes the suggestions.
[2409] Prompt Sentence Examples
[2410] As a concrete example, the following prompt sentence will be used.
[2411] Pre-departure setup:
[2412] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[2413] Tourist Information:
[2414] "Tell me about nearby tourist spots."
[2415] Restaurant Information:
[2416] I'm hungry, so I'm looking for a nearby restaurant.
[2417] Hotel Reservation Information:
[2418] "Find a hotel to stay at tonight."
[2419] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set destinations and change schedules while driving. It also recognizes the user's emotions through an emotion engine, providing a more personalized and comfortable travel experience.
[2420] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2421] Step 1:
[2422] User: Enters a voice command.
[2423] Specific operation: The user gives voice instructions such as, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 p.m."
[2424] Input: Voice command
[2425] Output: Audio data
[2426] Step 2:
[2427] Device: Activates the voice input engine and converts the voice data into text.
[2428] Specific operation: Converts audio data into text data using Google Cloud Speech-to-Text.
[2429] Input: Audio data
[2430] Output: Text data
[2431] Step 3:
[2432] Terminal: The converted text data is sent to the artificial intelligence model for analysis.
[2433] What it does: It uses GPT-4 to parse text and extract the destination and desired arrival time.
[2434] Input: Text data
[2435] Output: Analysis results (destination and desired arrival time)
[2436] Step 4:
[2437] Terminal: Sends the extracted information to the server.
[2438] Specific operation: Send information to the server about the final destination "Tokyo Station" and the desired arrival time "3:00 PM".
[2439] Input: Analysis results (destination and desired arrival time)
[2440] Output: Request to server
[2441] Step 5:
[2442] Server: Calculates the estimated arrival time using a route calculation engine and returns the result to the terminal.
[2443] Specific operation: Uses a route calculation engine to calculate the optimal route to the destination and derives the estimated arrival time.
[2444] Input: Destination and desired arrival time
[2445] Output: Estimated arrival time
[2446] Step 6:
[2447] Terminal: Provides audio feedback to the user on the estimated arrival time.
[2448] Specific operation: Using a speech synthesis engine, the system notifies the user, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[2449] Input: Estimated arrival time
[2450] Output: Feedback audio
[2451] Step 7:
[2452] Terminal: Activates the emotion engine and recognizes emotions from the user's voice data.
[2453] Specific behavior: Analyzes user emotions using the Microsoft Azure Emotion API.
[2454] Input: Audio data
[2455] Output: Emotion data
[2456] Step 8:
[2457] Device: Adjusts feedback content based on user emotional data.
[2458] Specific behavior: Based on the results of the emotion engine, adjust the feedback content and provide appropriate additional information.
[2459] Input: Emotion data
[2460] Output: Adjusted feedback content
[2461] Step 9:
[2462] User: Enters voice commands for nearby tourist attractions.
[2463] Specific operation: The user gives a voice command such as "Tell me about nearby tourist spots."
[2464] Input: Voice command
[2465] Output: Audio data
[2466] Step 10:
[2467] Device: Converts voice commands into text and sends location information to the server.
[2468] Specific operation: The voice data is converted into text data using Google Cloud Speech-to-Text and sent to the server along with GPS information.
[2469] Input: Audio data
[2470] Output: Text data and current location information
[2471] Step 11:
[2472] Server: Search for potential tourist destinations and obtain popularity, reviews, photos, and parking information.
[2473] Specific operation: Filters tourist destination candidates based on the current location from the database and obtains information about each location.
[2474] Input: Current location information
[2475] Output: Tourist destination information
[2476] Step 12:
[2477] Server: Sends the acquired tourist spot information to the terminal.
[2478] Specific operation: Send tourist spot information to the terminal.
[2479] Input: Tourist destination information
[2480] Output: Response to terminal
[2481] Step 13:
[2482] Terminal: Announces the received tourist spot information to the user by voice and receives feedback.
[2483] Specific operation: Uses a speech synthesis engine to notify the user of tourist information and receive feedback.
[2484] Input: Tourist destination information
[2485] Output: Feedback speech and user feedback
[2486] Step 14:
[2487] Terminal: Recalculate the route to the selected tourist spot and adjust the estimated arrival time.
[2488] What it does: Uses the route calculation engine to calculate a new route and adjust the estimated arrival time.
[2489] Input: Selected tourist destination
[2490] Output: New route and estimated arrival time
[2491] Step 15:
[2492] Device: The emotion engine recognizes the user's emotions and optimizes the suggestions.
[2493] What it does: It uses an emotion engine to analyze the user's emotions and adjusts suggestions based on the results.
[2494] Input: Emotion data
[2495] Output: Optimized proposals
[2496] (Application example 2)
[2497] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2498] While traveling by car, users need to obtain real-time information on tourist spots, restaurants, hotels, etc., and make reservations. However, conventional car navigation systems do not provide personalized suggestions based on the user's emotions. Another issue is that it is difficult to ensure safety when users set destinations or make reservations while driving. To solve these issues, a new car navigation system that combines voice input and emotion recognition functions is needed.
[2499] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2500] In this invention, the server includes means for activating a voice input engine and receiving voice commands from the user, processing means using an artificial intelligence model to convert voice data into text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing voice feedback on the estimated arrival time to the user, and means including an emotion engine for recognizing emotions from the user's voice data and adjusting the feedback content. This makes it possible to make personalized suggestions based on the user's emotions, providing a safe and comfortable travel experience.
[2501] "Voice input engine" is a general term for devices and software for receiving voice commands from users.
[2502] "Artificial Intelligence Model" means the machine learning algorithms and techniques used to convert voice data into text.
[2503] "Destination and desired arrival time" refers to the final destination set by the user and the desired arrival time at that destination.
[2504] A "server" is a computer system that processes information entered by a user and provides the necessary information and estimated time.
[2505] The "estimated arrival time" is the estimated time required to arrive at the destination specified by the user.
[2506] The "emotion engine" is a technology and algorithm that recognizes emotions from the user's voice data and optimizes the feedback content based on the results.
[2507] "Candidate tourist destinations" are potential travel destinations suggested based on the user's current location and route information.
[2508] "Popularity, word-of-mouth reputation, photos and parking information" is a general term for ratings and reviews of tourist spots and restaurants, as well as images and information about parking.
[2509] "Feedback" refers to the response or guidance provided by the system to the user.
[2510] "Recalculating the route" means recalculating the travel route based on the user's selection and the situation.
[2511] The system of the present invention is a car navigation application system that supports travel in autonomous vehicles, and uses the following main hardware and software:
[2512] 1. Voice Input Engine
[2513] A voice input engine is a device or software that receives voice commands from users, specifically voice recognition engines such as Google Voice Recognition and Apple's Siri, allowing users to input commands using only their voice without using their hands.
[2514] 2. Artificial Intelligence Model
[2515] Artificial intelligence models are used to convert the speech data into text, such as Google Cloud Speech-to-Text API and IBM Watson Speech to Text. This process converts the speech data into text.
[2516] 3. Emotion Engine
[2517] The emotion engine is a technology and algorithm that recognizes emotions from users' voice data and optimizes feedback content. Specifically, it uses Microsoft Azure Emotion API and Affectiva's emotion recognition technology. This enables personalized suggestions to be provided based on the user's emotions.
[2518] 4. Navigation system
[2519] The navigation system uses the Google Maps API, Here Maps API, etc. to calculate the route to the destination, locate the current location, and estimate the arrival time. It also obtains information on nearby tourist attractions, restaurants, hotels, etc.
[2520] 5. Recommendation Services
[2521] The recommendation service provides information on tourist spots, restaurants, and hotels using APIs such as FourSquare and Yelp, and obtains detailed information such as popularity, reviews, photos, and parking information, and makes suggestions to users.
[2522] Specific examples
[2523] Example of user voice input and system response
[2524] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[2525] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[2526] Server: Calculates the estimated arrival time and returns it to the device.
[2527] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[2528] On the device: The emotion engine recognizes emotions from the user's voice data and provides additional advice and information depending on the user's mood.
[2529] Example of tourist information
[2530] User: "What are some nearby tourist spots?"
[2531] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[2532] User: "I'll be there."
[2533] Terminal: "Route set to museum. Arrival time 15 minutes."
[2534] On the device: The emotion engine recognizes the user's emotions and suggests additional tourist destinations that may be of interest.
[2535] This allows users to have a comfortable and personalized travel experience.
[2536] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2537] Step 1:
[2538] Receiving voice commands
[2539] The user issues a voice command.
[2540] The terminal activates a voice input engine and receives voice commands from the user.
[2541] Input: User's voice command
[2542] Output: Audio data
[2543] Step 2:
[2544] Converting audio data to text
[2545] The device converts the voice data into text using an artificial intelligence model (e.g., Google Cloud Speech-to-Text API).
[2546] Input: Audio data
[2547] Output: Text data
[2548] Step 3:
[2549] Extracting destination and desired arrival time
[2550] The device uses natural language processing technology to extract the destination and desired arrival time from the text data.
[2551] Input: Text data
[2552] Output: Extracted information about destination and desired arrival time
[2553] Step 4:
[2554] Destination and desired arrival time sent to server
[2555] The terminal transmits the extracted information to the server and requests it to calculate the estimated arrival time.
[2556] Input: Extract information about destination and desired arrival time
[2557] Output: Request to calculate estimated arrival time
[2558] Step 5:
[2559] Calculating and receiving estimated arrival times
[2560] The server uses a navigation system (e.g., Google Maps API) to calculate the estimated arrival time and sends this information back to the device.
[2561] Input: Extract information about destination and desired arrival time
[2562] Output: Estimated arrival time
[2563] Step 6:
[2564] emotion recognition
[2565] The device uses an emotion engine (e.g., Microsoft Azure Emotion API) to recognize emotions from the user's voice data.
[2566] Input: User's voice data
[2567] Output: Emotion data
[2568] Step 7:
[2569] Personalized Feedback
[2570] The terminal adjusts the feedback content of the estimated arrival time based on the emotion data and provides the feedback to the user by voice.
[2571] Input: Estimated arrival time, emotion data
[2572] Output: Audio feedback
[2573] Step 8:
[2574] Search and submit tourist destination suggestions
[2575] The server searches for nearby tourist spots based on the current location and route information using a recommendation service (e.g., FourSquare API) and sends detailed information to the device.
[2576] Input: Current location and route information
[2577] Output: Tourist destination information
[2578] Step 9:
[2579] To announce tourist destination information and receive feedback
[2580] The terminal notifies the user of tourist spot information by voice and receives feedback from the user.
[2581] Input: Tourist destination information
[2582] Output: User feedback
[2583] Step 10:
[2584] Recalculate route to selected tourist spot
[2585] The server recalculates the route to the tourist spot selected by the user and adjusts the estimated arrival time.
[2586] Input: User feedback
[2587] Output: Recalculated route and estimated arrival time
[2588] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2589] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2590] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2591] [Fourth embodiment]
[2592] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2593] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2594] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2595] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2596] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2597] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2598] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2599] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2600] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2601] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2602] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2603] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2604] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2605] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, and adjusts routes and arrival times in real time.
[2606] Functionality Overview
[2607] 1. Pre-departure settings
[2608] Device:
[2609] It starts the voice input engine and receives the user's voice commands.
[2610] The voice data is sent to an artificial intelligence model and converted into text.
[2611] Information about the final destination and desired arrival time set by the user is sent to the server.
[2612] The destination information returned from the server is checked and audio feedback is given to the user.
[2613] 2. Tourist information while driving
[2614] server:
[2615] Potential tourist spots in the area are selected based on the user's route and current location.
[2616] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[2617] This information is sent to the terminal.
[2618] Device:
[2619] Tourist information from the server is announced to the user by voice.
[2620] It receives user feedback and sends it to the server.
[2621] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[2622] 3. Restaurant directions while driving
[2623] server:
[2624] Select nearby restaurant options based on the user's current location and route.
[2625] Obtain information such as restaurant reviews, budget, and whether parking is available.
[2626] If necessary, the reservation telephone number is transmitted to the terminal.
[2627] Device:
[2628] Restaurant information sent from the server is announced to the user by voice.
[2629] Receive user feedback and send selections to the server.
[2630] Make a reservation if necessary.
[2631] It provides the functionality to receive user feedback, send it to the server, and make reservation calls.
[2632] 4. Hotel reservation information while driving
[2633] server:
[2634] Select nearby hotel options based on the user's current location and destination.
[2635] Obtain hotel reviews, popularity, budget, parking information, etc.
[2636] If necessary, information for reservation procedures is sent to the terminal.
[2637] Device:
[2638] Hotel information sent from the server is announced to the user by voice.
[2639] Receive user feedback and send the selection results to the server.
[2640] Proceed with the booking process as needed.
[2641] Once the booking is completed, the user is given audio feedback.
[2642] Recalculate the route to the user's selected hotel and adjust the estimated arrival time.
[2643] Specific examples
[2644] 1. Pre-departure settings
[2645] User: "Car Navigation Assist, set a route to Tokyo Station, desired arrival time 3pm."
[2646] Terminal: "Tokyo Station, right? Desired arrival time is 3:00 PM. I'll confirm."
[2647] Server: Calculates the estimated arrival time and returns it to the device.
[2648] Terminal: "Settings complete. Scheduled to arrive at Tokyo Station at 3:00 PM."
[2649] 2. Tourist information
[2650] User: "What are some nearby tourist spots?"
[2651] Terminal: "Yes, there's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[2652] User: "I'll be there."
[2653] Terminal: "Route set to museum. Arrival time 15 minutes."
[2654] 3. Restaurant Information
[2655] User: "I'm hungry, can you find a restaurant nearby?"
[2656] Terminal: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and my budget is around 1500 yen for lunch. Would you like to make a reservation?"
[2657] User: "Yes, I'd like to make a reservation."
[2658] Terminal: "Reservation has been arranged. Expected to arrive in 10 minutes."
[2659] 4. Hotel Reservation Information
[2660] User: "Find a hotel tonight."
[2661] Terminal: "There's a highly rated hotel nearby with 4.8 reviews and parking. Would you like to make a reservation?"
[2662] User: "Yes, please make a reservation."
[2663] Terminal: "Hotel booked. Estimated arrival time is 6pm."
[2664] In this way, the system of the present invention allows all operations to be performed by voice, allowing you to safely set your destination or change your schedule while driving, providing a comfortable travel experience.
[2665] The processing flow will be explained below.
[2666] 1. Pre-departure settings
[2667] Step 1:
[2668] The user issues a voice command: "Car Navigation Assist, set route to Tokyo Station, desired arrival time 3:00 PM."
[2669] Step 2:
[2670] The device activates the voice input engine and receives the user's voice command.
[2671] Step 3:
[2672] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2673] Step 4:
[2674] The device extracts the destination (Tokyo Station) and desired arrival time (3:00 p.m.) from the converted text data.
[2675] Step 5:
[2676] The terminal transmits the extracted information to the server.
[2677] Step 6:
[2678] The server calculates the estimated arrival time and returns the result to the terminal.
[2679] Step 7:
[2680] The terminal will then provide the user with a voice feedback of the estimated arrival time received from the server: "Settings complete. Expected arrival at Tokyo Station at 3:00 PM."
[2681] 2. Tourist information while driving
[2682] Step 1:
[2683] The user issues a voice command: "Tell me about nearby tourist attractions."
[2684] Step 2:
[2685] The device activates the voice input engine and receives the user's voice command.
[2686] Step 3:
[2687] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2688] Step 4:
[2689] The terminal transmits the converted text data to the server.
[2690] Step 5:
[2691] The server selects nearby tourist spots based on the user's route and current location.
[2692] Step 6:
[2693] The server obtains the popularity of the tourist spot, reviews, photos, and parking information, and sends this information to the terminal.
[2694] Step 7:
[2695] The device receives tourist information from the server and announces it to the user by voice: "There's a famous museum nearby. It has a 4.5 star rating and there's parking available. What would you like to do?"
[2696] Step 8:
[2697] The user gives verbal feedback: "I'm going there."
[2698] Step 9:
[2699] The device sends the user's feedback to the server.
[2700] Step 10:
[2701] The device recalculates the route to the tourist spot and adjusts it to make it in time for the maximum estimated arrival time.
[2702] Step 11:
[2703] The device will then provide audible feedback to the user about the recalculated route: "Route to the museum has been set. It will take 15 minutes to arrive."
[2704] 3. Restaurant directions while driving
[2705] Step 1:
[2706] A user issues a voice command: "I'm hungry, find me a nearby restaurant."
[2707] Step 2:
[2708] The device activates the voice input engine and receives the user's voice command.
[2709] Step 3:
[2710] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2711] Step 4:
[2712] The terminal transmits the converted text data to the server.
[2713] Step 5:
[2714] The server searches for nearby restaurant options based on the user's current location and route.
[2715] Step 6:
[2716] The server retrieves restaurant reviews, budget, and parking information and sends that information to the terminal.
[2717] Step 7:
[2718] The device receives restaurant information from the server and announces it to the user via voice: "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[2719] Step 8:
[2720] The user gives verbal feedback: "Yes, I'd like to make a reservation as well."
[2721] Step 9:
[2722] The terminal sends the user's feedback to the server and activates the reservation call function.
[2723] Step 10:
[2724] The device completes the reservation and provides the user with a voice feedback confirming the reservation: "Your reservation has been arranged. You will arrive in 10 minutes."
[2725] Step 11:
[2726] The device recalculates the route to the restaurant and adjusts the estimated arrival time to the final destination.
[2727] 4. Hotel reservation information while driving
[2728] Step 1:
[2729] A user issues a voice command: "Find a hotel for tonight."
[2730] Step 2:
[2731] The device activates the voice input engine and receives the user's voice command.
[2732] Step 3:
[2733] The device sends the voice data to an artificial intelligence model, which converts it into text data.
[2734] Step 4:
[2735] The terminal transmits the converted text data to the server.
[2736] Step 5:
[2737] The server searches for nearby hotel options based on the user's current location and destination.
[2738] Step 6:
[2739] The server obtains hotel reviews, popularity, budget, and parking information and sends this information to the terminal.
[2740] Step 7:
[2741] The terminal receives hotel information from the server and announces it to the user by voice: "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?"
[2742] Step 8:
[2743] The user gives verbal feedback: "Yes, please book."
[2744] Step 9:
[2745] The device sends the user's feedback to the server and proceeds with the reservation process.
[2746] Step 10:
[2747] The device completes the reservation and provides the user with audio feedback confirming the reservation: "Hotel has been booked. Estimated arrival time is 6:00 PM."
[2748] Step 11:
[2749] The device recalculates the route to the hotel and adjusts the estimated arrival time to the final destination.
[2750] Example 1
[2751] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2752] Current car navigation systems have the challenge of making it difficult for users to safely and efficiently set destinations and adjust schedules while driving. Furthermore, they lack the functionality to provide comprehensive, real-time information on tourist spots, restaurants, hotels, and other information, preventing users from making appropriate choices quickly. This often results in a loss of convenience and safety while driving.
[2753] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2754] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for suggesting tourist spots, restaurants, and hotels based on the user's current location and destination and acquiring information, means for notifying the user of the acquired information audio-visually and receiving feedback, means for recalculating a route to the selected destination and adjusting the estimated arrival time, and means for making reservations if necessary. This allows destination setting and schedule adjustments to be performed safely and efficiently even while driving, providing a comfortable travel experience.
[2755] A "voice input engine" is a device or software that receives a user's voice commands and converts them into digital data.
[2756] "Voice Data" means human speech information converted into digital form by a speech input engine.
[2757] "Artificial intelligence model" refers to technologies such as machine learning algorithms and neural networks used to convert voice data into text data.
[2758] "Text data" is character string information converted from voice data by an artificial intelligence model.
[2759] "Destination" is information indicating the place or location to which the user wishes to travel.
[2760] The "desired arrival time" is information indicating a specific time at which the user desires to arrive at the destination.
[2761] A "server" is a central system that receives requests on a computer network, processes data, and sends and receives information.
[2762] "Estimated arrival time" is information indicating the estimated arrival time from the current location to the destination.
[2763] "Tourist destination" refers to tourist spots and famous places that users aim to visit.
[2764] "Restaurant" means an eating and drinking establishment selected by a patron for dining.
[2765] "Hotel" means the accommodation facility selected for the Guest's stay.
[2766] "Means of acquisition" refers to the methods and technologies used to collect and receive the required information or data.
[2767] "Means of receiving feedback" refers to the methods and techniques for receiving responses or reactions from users.
[2768] "Means for recalculating a route" refers to a method or technology for recalculating a new route based on specified conditions.
[2769] "Means for completing reservation procedures" refers to the methods and technologies used to make reservations for the service selected by the user (such as restaurant reservations or hotel reservations).
[2770] The system of the present invention is a car navigation application system that supports car travel, and suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice operations, adjusting routes and arrival times in real time.
[2771] The main components of the system are the "terminals" used by users and the "servers" that process data. The specific configuration and operation are explained below.
[2772] The system's devices, which include smartphones and car navigation systems, are equipped with a voice input engine that receives users' voice commands and converts the voice data into a digital format. This digital voice data is then converted into text data using voice recognition software such as Google Cloud Speech-to-Text.
[2773] The converted text data is sent to the server and used to extract the destination and desired arrival time. The server calculates the estimated arrival time based on this information and returns the result to the terminal. The terminal then provides this information as feedback to the user via voice. Specifically, it notifies the user by saying something like, "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM."
[2774] In the case of tourist attraction guidance while driving, the device sends the user's current location and route information to the server, and the server searches for nearby tourist attractions. The server obtains the tourist attraction's popularity, reviews, photos, and parking information and sends them to the device. The device then announces this information to the user by voice (e.g., "Yes, there's a famous museum nearby. It has a 4.5 rating and there's parking available. What would you like to do?"). Based on the user's feedback, the server recalculates the route and calculates a new estimated arrival time.
[2775] Regarding restaurant guidance, the system searches for nearby restaurants based on the current location and route, and obtains information such as reviews, budget, and parking information. If necessary, it obtains a reservation phone number and sends it to the terminal. The terminal then announces this information to the user by voice (e.g., "There's a highly rated Chinese restaurant nearby. The reviews are 4.7, and your budget is around 1,500 yen for lunch. Would you like to make a reservation?") and proceeds with the reservation process based on the user's feedback.
[2776] Additionally, the hotel guide searches for nearby hotels based on the current location and destination, and obtains information such as reviews, popularity, budget, and parking information. If a reservation is required, the server sends the necessary information to the terminal, which then proceeds with the reservation process (e.g., "Looking for a hotel to stay at tonight," "There's a highly rated hotel nearby. It has a 4.8 rating and parking. Would you like to make a reservation?").
[2777] In this way, the system of the present invention allows all operations to be performed by voice, allowing users to safely set their destination or change their schedule while driving, providing a comfortable travel experience.
[2778] Specific prompt examples
[2779] "Car navigation assist, set route to Tokyo Station, desired arrival time is 3pm."
[2780] "Tell me about nearby tourist spots."
[2781] I'm hungry, so I'm looking for a nearby restaurant.
[2782] "Find a hotel to stay at tonight."
[2783] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2784] The flow of this system's program processing
[2785] 1. Pre-departure settings
[2786] Step 1:
[2787] Subject: User
[2788] Description: The user speaks a voice command into the device.
[2789] Specific actions: For example, say, "Car navigation assist, set route to Tokyo Station, desired arrival time is 3:00 PM."
[2790] Step 2:
[2791] Subject: Device
[2792] Description: Starts the voice input engine and receives voice commands.
[2793] Input: Voice command from the user
[2794] Output: Digital audio data
[2795] What it does: The voice input engine receives the voice and converts it into a digital format.
[2796] Step 3:
[2797] Subject: Device
[2798] Description: Uses artificial intelligence models to convert voice data into text.
[2799] Input: Digital audio data
[2800] Output: Text data
[2801] What it does: Converts speech to text using a service like Google Cloud Speech-to-Text.
[2802] Step 4:
[2803] Subject: Device
[2804] Description: Sends the converted text data to the server.
[2805] Input: Text data (e.g., "Set a route to Tokyo Station, with a desired arrival time of 3:00 PM.")
[2806] Output: Request to server
[2807] Specific operation: Sends an HTTP request to the server.
[2808] Step 5:
[2809] Subject: Server
[2810] Description: Extracts the destination and desired arrival time and calculates the estimated arrival time.
[2811] Input: Text data
[2812] Output: Estimated arrival time
[2813] Specific operation: Uses natural language processing technology to analyze text and calculate arrival times by referencing traffic information and road conditions.
[2814] Step 6:
[2815] Subject: Server
[2816] Description: Sends the calculation result back to the terminal.
[2817] Input: Estimated arrival time
[2818] Output: Feedback data
[2819] Specific operation: Returns the calculation result to the terminal.
[2820] Step 7:
[2821] Subject: Device
[2822] Description: Provides audio feedback data from the server to the user.
[2823] Input: Feedback data
[2824] Output: Audio notification to the user
[2825] Specific operation: The text data is converted into speech, and a message such as "Settings are complete. You are scheduled to arrive at Tokyo Station at 3:00 PM" is spoken to the user.
[2826] 2. Tourist information while driving
[2827] Step 1:
[2828] Subject: User
[2829] Description: Request "Tell me about nearby tourist spots."
[2830] Specific action: Speak a voice command into the device.
[2831] Step 2:
[2832] Subject: Device
[2833] Description: Converts voice commands into text and sends it to the server.
[2834] Input: Voice command
[2835] Output: Text data
[2836] Specific operation: The audio is converted into text using Google Cloud Speech-to-Text or similar and sent to the server.
[2837] Step 3:
[2838] Subject: Server
[2839] Description: Searches for nearby tourist spots based on the user's current location and route.
[2840] Input: Current location and route information
[2841] Output: List of tourist destination candidates
[2842] Specific behavior: Generate a list of tourist destinations by retrieving information from a database or external API.
[2843] Step 4:
[2844] Subject: Server
[2845] Description: Get tourist attraction popularity, reviews, photos, and parking information.
[2846] Input: Tourist destination candidate list
[2847] Output: Detailed information
[2848] Specific actions: Collect and list detailed information about each tourist destination.
[2849] Step 5:
[2850] Subject: Server
[2851] Description: Sends the acquired information to the device.
[2852] Input: More information
[2853] Output: Feedback data
[2854] Specific operation: Sends collected information to the device.
[2855] Step 6:
[2856] Subject: Device
[2857] Description: Presents information to the user audibly.
[2858] Input: Feedback data
[2859] Output: Audio notification to the user
[2860] Action: The audio message is, "There's a famous museum nearby. It has 4.5 star reviews and parking. What would you like to do?"
[2861] Step 7:
[2862] Subject: User
[2863] Description: Gives feedback saying "There you go."
[2864] Specific actions: Select a tourist spot specified by voice.
[2865] Step 8:
[2866] Subject: Device
[2867] Description: Recalculates route based on user selection.
[2868] Input: User's choice
[2869] Output: Recalculated route
[2870] Specific behavior: Calculate a new route and send a notification such as "Route to the museum has been set. It will take 15 minutes to arrive."
[2871] 3. Restaurant directions while driving
[2872] Step 1:
[2873] Subject: User
[2874] Description: "I'm hungry, find me a nearby restaurant."
[2875] Specific action: Speak a voice command into the device.
[2876] Step 2:
[2877] Subject: Device
[2878] Description: Sends a voice command to the server.
[2879] Input: Voice command
[2880] Output: Text data
[2881] Specific operation: Converts speech into text and sends it to the server.
[2882] Step 3:
[2883] Subject: Server
[2884] Description: Find nearby restaurant suggestions based on your current location and route.
[2885] Input: Current location and route information
[2886] Output: Restaurant candidate list
[2887] Specific operation: Retrieves restaurant information from a database or external API and creates a list.
[2888] Step 4:
[2889] Subject: Server
[2890] Description: Get detailed information like reviews, budget, parking info, etc.
[2891] Input: Restaurant candidate list
[2892] Output: Detailed information
[2893] What it does: Collect and list detailed information about each restaurant.
[2894] Step 5:
[2895] Subject: Server
[2896] Description: Sends the acquired information to the device.
[2897] Input: More information
[2898] Output: Feedback data
[2899] Specific operation: Sends collected information to the device.
[2900] Step 6:
[2901] Subject: Device
[2902] Description: Presents information to the user audibly.
[2903] Input: Feedback data
[2904] Output: Audio notification to the user
[2905] Specific action: Say, "There's a highly rated Chinese restaurant nearby. It has a 4.7 rating and your budget is around 1,500 yen for lunch. Would you like to make a reservation?"
[2906] Step 7:
[2907] Subject: User
[2908] Instructions: Answer "Yes, I'd like to make a reservation as well."
[2909] Specific action: Indicate your intention to make a reservation by voice.
[2910] Step 8:
[2911] Subject: Device
[2912] Description: Proceed with the booking process.
[2913] Input: User's booking request
[2914] Output: Reservation completion notification
[2915] Specific operation: Make a reservation using the OpenTable API or similar, and notify the customer, "Your reservation has been arranged. You will arrive in 10 minutes."
[2916] 4. Hotel reservation information while driving
[2917] Step 1:
[2918] Subject: User
[2919] Description: "Find me a hotel tonight."
[2920] Specific action: Speak a voice command into the device.
[2921] Step 2:
[2922] Subject: Device
[2923] Description: Sends a voice command to the server.
[2924] Input: Voice command
[2925] Output: Text data
[2926] Specific operation: Converts voice commands into text and sends it to the server.
[2927] Step 3:
[2928] Subject: Server
[2929] Description: Search for nearby hotel suggestions based on your current location and destination.
[2930] Input: Current location and destination information
[2931] Output: Hotel candidate list
[2932] Specific operation: Retrieves hotel information from a database or external API and creates a list.
[2933] Step 4:
[2934] Subject: Server
[2935] Description: Get detailed information like reviews, popularity, budget, parking information, and more.
[2936] Input: Hotel candidate list
[2937] Output: Detailed information
[2938] What to do: Collect and list detailed information about each hotel.
[2939] Step 5:
[2940] Subject: Server
[2941] Description: Sends the acquired information to the device.
[2942] Input: More information
[2943] Output: Feedback data
[2944] Specific operation: Sends collected information to the device.
[2945] Step 6:
[2946] Subject: Device
[2947] Description: Presents information to the user audibly.
[2948] Input: Feedback data
[2949] Output: Audio notification to the user
[2950] Action: The voice will say, "There's a highly rated hotel nearby with a 4.8 rating and parking. Would you like to make a reservation?"
[2951] Step 7:
[2952] Subject: User
[2953] Instructions: "Yes, please make a reservation."
[2954] Specific action: Indicate your intention to make a reservation by voice.
[2955] Step 8:
[2956] Subject: Device
[2957] Description: Proceed with the booking process.
[2958] Input: User's booking request
[2959] Output: Reservation completion notification
[2960] Specific behavior: Make a reservation using Expedia API or Booking.com API and notify the user, "Hotel has been booked. Estimated arrival time is 6:00 PM."
[2961] (Application example 1)
[2962] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2963] When traveling by car, there is a demand for voice control to select and reserve destinations, tourist attractions, restaurants, and hotels along the way, and for integration with autonomous driving systems to provide a safe and comfortable travel experience. However, with conventional systems, it can be difficult to provide sufficient information, complete reservations, and adjust routes using voice control alone. Furthermore, there is a lack of integration with autonomous driving systems, and manual operation by the user is often required. This can make operations while driving cumbersome, potentially compromising safety and efficiency.
[2964] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2965] In this invention, the server includes means for activating a voice input engine and receiving voice commands from a user, processing means using an artificial intelligence model to convert voice data to text, means for extracting a destination and desired arrival time from the converted text data, means for transmitting the extracted information to the server and receiving an estimated arrival time, means for providing audio feedback on the estimated arrival time to the user, means for adding control means compatible with the autonomous driving system, and means for automating automatic route adjustment to selected facilities and reservation procedures. This makes it possible for an autonomous vehicle to utilize voice operation and AI technology to select and reserve tourist spots, restaurants, and hotels in real time based on user instructions, and to automatically adjust the route and provide feedback on the arrival time.
[2966] A "voice input engine" is a device or software that receives voice commands from a user.
[2967] "Artificial intelligence model" is a machine learning algorithm used to convert voice data into text.
[2968] "Text data" is voice data converted into text format.
[2969] A "server" is a computer system that provides services to clients over a computer network.
[2970] The "destination" is the destination point set by the user.
[2971] "Desired arrival time" is the time the user desires to arrive at the destination.
[2972] "Estimated arrival time" is the estimated time of arrival at the destination calculated by the server.
[2973] An "autonomous driving system" is a technology that enables vehicles to perform driving operations autonomously.
[2974] "Route adjustment" refers to recalculating the route to a selected destination.
[2975] "Tourist destination candidates" is a list of tourist destinations suggested to the user.
[2976] "Popularity" is an indicator that shows the evaluation of tourist destinations and facilities.
[2977] "Word of mouth" refers to reviews and feedback from users.
[2978] "Photos" are image data that provide visual images of tourist spots and facilities.
[2979] "Parking information" is information about parking spaces at tourist spots and facilities.
[2980] "Restaurant candidates" is a list of restaurants suggested to the user.
[2981] The "budget" is an estimate of the cost for the user to use the service.
[2982] "Reservation phone" refers to a means of making a reservation for a facility by telephone.
[2983] In one embodiment of the present invention, the system first activates a voice input engine to receive a user's voice command. The voice input engine may be, for example, the Google Speech Recognition API or a similar voice recognition engine. The device that receives the voice data converts it into text data using an artificial intelligence model (e.g., a generative AI model).
[2984] Natural language processing (NLP) techniques are used to extract the destination and desired arrival time from the converted text data. The text data is analyzed to extract specific information (in this case, the destination and desired arrival time). This information is sent to a server, which calculates the estimated arrival time. The calculation result is sent back to the device, which then communicates the result to the user as voice feedback.
[2985] By adding control means compatible with the autonomous driving system, it becomes possible to automatically adjust routes to selected facilities and make reservations. Specifically, the terminal provides route information to the autonomous driving system based on information received from the server and uses an API for reservation procedures. This allows users to set destinations and make reservations using only voice commands without manual operation.
[2986] For example, if a user voice-inputs "My destination is Tokyo Station, and I would like to arrive at 3:00 PM," the device converts the voice data into text data and extracts the destination and desired arrival time. This information is sent to the server, which calculates the estimated arrival time and receives feedback. Furthermore, if the user issues the command "Find a nearby restaurant," the server searches for restaurant candidates based on the current location and route information, obtains reviews and budget information, and sends it to the device. The device notifies the user of this by voice, and if the user responds "Make a reservation," the device will automatically complete the reservation procedure.
[2987] An example of a prompt sentence might be:
[2988] Please let us know your destination and desired arrival time.
[2989] Could you tell me about nearby tourist spots?
[2990] Find a restaurant near you.
[2991] Find a hotel to stay in tonight.
[2992] In this way, the present invention combines voice control with an automated driving system to provide users with a safe and convenient travel experience.
[2993] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2994] Step 1:
[2995] The device starts a voice input engine to receive the user's voice command. The voice input engine used here is a speech recognition engine such as the Google Speech Recognition API. The voice input engine captures the user's voice data and processes it as digital voice data.
[2996] Input: User's voice command
[2997] Output: Digital audio data
[2998] Step 2:
[2999] The voice data received by the device is converted into text data using a generative AI model (a speech recognition algorithm). This process involves analyzing the voice signal and generating the corresponding text.
[3000] Input: Digital audio data
[3001] Output: Text data
[3002] Step 3:
[3003] The device extracts the destination and desired arrival time from the generated text data, using natural language processing (NLP) techniques to identify keywords and phrases related to the destination and desired arrival time.
[3004] Input: Text data
[3005] Output: Destination and desired arrival time information
[3006] Step 4:
[3007] The device sends the extracted information to a server, which calculates and returns an estimated arrival time. The server then uses a map database and traffic data to calculate a route from the current location to the destination.
[3008] Input: Destination and desired arrival time information
[3009] Output: Estimated arrival time
[3010] Step 5:
[3011] The terminal receives the estimated arrival time from the server and provides the user with audio feedback. Here, a synthetic speech engine (e.g., pyttsx3) is used to generate speech from text and communicate it to the user.
[3012] Input: Estimated arrival time
[3013] Output: Audio feedback
[3014] Step 6:
[3015] The user inputs an additional voice command (e.g., "Tell me about nearby tourist spots," "Find nearby restaurants," etc.). Based on this prompt, the device queries the server.
[3016] Input: User's additional voice command
[3017] Output: prompt statement
[3018] Step 7:
[3019] The server searches for potential tourist spots and restaurants based on the user's current location and route information, and then retrieves and sends the information to the device. The retrieved information includes popularity, reviews, photos, parking information, budget, etc.
[3020] Input: User's current location and route information
[3021] Output: Information on tourist spots and restaurants
[3022] Step 8:
[3023] The terminal announces the information received from the server to the user by voice and receives user feedback. When the user makes a selection, the selection is sent to the server, which then processes the new route and reservation.
[3024] Input: Information about tourist attractions and restaurants
[3025] Output: Audio feedback and user selection results
[3026] Step 9:
[3027] Based on the user's feedback, the device issues instructions to the autonomous driving system, automatically adjusting the route to the selected facility and making reservations, allowing the user to automatically head to the next destination without manual intervention.
[3028] Input: User selection
[3029] Output: Instructions to the autonomous driving system and reservation procedures
[3030] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[3031] The system of this invention is a car navigation application system that supports car travel. It suggests and reserves tourist spots, rest areas, restaurants, and hotels based on the user's voice commands, and adjusts routes and arrival times in real time. It also incorporates an emotion engine that recognizes the user's emotions, providing a more personalized experience.
[3032] Functionality Overview
[3033] 1. Pre-departure settings
[3034] Device:
[3035] It starts the voice input engine and receives the user's voice commands.
[3036] The voice data is sent to an artificial intelligence model and converted into text.
[3037] Information about the final destination and desired arrival time set by the user is sent to the server.
[3038] The destination information returned from the server is checked and audio feedback is given to the user.
[3039] The emotion engine recognizes emotions from the user's voice data and adjusts the feedback content according to the user's emotions.
[3040] 2. Tourist information while driving
[3041] server:
[3042] Potential tourist spots in the area are selected based on the user's route and current location.
[3043] Obtain information on the popularity of tourist destinations, reviews, photos, and parking information.
[3044] This information is sent to the terminal.
[3045] Device:
[3046] Tourist information from the server is announced to the user by voice.
[3047] The emotion engine recognizes the user's emotions and optimizes tourist destination suggestions based on the results.
[3048] It receives user feedback via voice and sends it to the server.
[3049] Recalculate your route to the tourist spot you selected to meet the maximum estimated arrival time.
[3050] 3. Restaurant directions while driving
[3051] server:
[3052] Select nearby restaurant options based on the user's current location and route.
[3053] Get restaurant reviews, budget and parking information.
[3054] If necessary, the reservation telephone number is transmitted to the terminal.
[3055] Device:
[3056] Restaurant information sent from the server is announced to the user by voice.
[3057] The emotion engine recognizes the user's emotions and optimizes restaurant suggestions based on the results.
[3058] Receive user feedback and send selections to the server.
[3059] Make a reservation if necessary.
[3060] Once the booking is completed, the user is given audio feedback.
[3061] Recalculate your route to the restaurant and adjust your estimated arrival time at your final destination.
[3062] 4. Hotel reservation information while driving
[3063] server:
[3064] Search for nearby hotel suggestions based on the user's current location and destination.
[3065] Get hotel reviews, popularity, budget, and parking information.
[3066] If necessary, information regarding the reservation procedure is sent to the terminal.
[3067] Device:
[3068] Hotel information sent from the server ...
Claims
1. means for activating a voice input engine and receiving voice commands from a user; a processing means using an artificial intelligence model to convert voice data into text; means for extracting a destination and a desired arrival time from the converted text data; means for transmitting the extracted information to a server and receiving an estimated arrival time; and means for providing audio feedback to the user about the estimated arrival time.
2. A means for searching for potential tourist spots in the vicinity based on current location and route information, and acquiring and transmitting popularity, reviews, photos, and parking information; a means for announcing the received information to the user by voice and receiving feedback; and means for recalculating the route to the selected tourist attraction and adjusting the estimated time of arrival.
3. A way to search for nearby restaurant options based on your current location and route, and obtain and submit reviews, budget, and parking information. a means for announcing the received information to the user by voice and receiving feedback; A way to make a reservation call if necessary, and and means for recalculating a route to the selected restaurant.
4. A way to search for nearby accommodation options based on your current location and destination, and obtain and submit reviews, popularity, budget, and parking information; a means for announcing the received information to the user by voice and receiving feedback; A means to proceed with the booking process if necessary; and means for recalculating the route to the selected accommodation and adjusting the estimated time of arrival.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A