system

The system addresses the inefficiencies of conventional navigation systems by allowing voice or text input for destination requests, providing integrated navigation and tourist guidance through a terminal and server, enhancing user convenience and efficiency.

JP2026064692APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Conventional car navigation systems require significant user effort for inputting destinations and separate searches for detailed information, making it difficult to provide intuitive and seamless guidance to destinations and tourist spots.

Method used

A system that allows users to request destinations by voice or text, analyzes the data, extracts relevant information, and provides navigation and tourist guidance through a terminal and server using generative artificial intelligence, enabling real-time information retrieval and presentation.

Benefits of technology

Enables users to receive destination guidance and sightseeing information with simple operations, improving convenience and efficiency by integrating voice or text input with real-time navigation and tourist information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064692000001_ABST
    Figure 2026064692000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for the user to request a destination by voice or text, A means of analyzing the requested data and extracting destination information, A means for sending the extracted destination information to the server, A means for searching for details of the relevant destination and generating candidate information based on destination information received by the server, A means for sending the generated candidate information to the terminal, A means for the terminal to present the received candidate information to the user via voice or display and initiate navigation, During navigation, a means of receiving tourist information requests from users and providing further information, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, with the development of autonomous driving technology, there has been an increasing demand for users to provide guidance to destinations more smoothly and intuitively. However, in conventional car navigation systems, it takes a great deal of effort for users to input a destination, and separate searches are required to obtain detailed information such as tourist spots. In such a situation, in order to improve user convenience, there is a need for a system that can easily request a destination by voice or text and then provide subsequent navigation and tourist guidance collectively.

Means for Solving the Problems

[0005] To solve this problem, the present invention provides the following means: a means for the user to request a destination by voice or text; a means for analyzing the requested data and extracting destination information; a means for transmitting the extracted destination information to a server; a means for the server to search for details of the relevant destination based on the destination information received and generate candidate information; a means for transmitting the generated candidate information to a terminal; a means for the terminal to present the received candidate information to the user by voice or display and start navigation; and further, a means for receiving requests for sightseeing information from the user during navigation and providing additional information. As a result, the user can receive both destination guidance and sightseeing information with simple operation, greatly improving convenience.

[0006] A "user" refers to a person who uses this system and is the entity that makes a destination request via voice or text.

[0007] "Means of making requests by voice or text" refers to an interface that allows users to provide information such as destinations to the system via voice input or text input.

[0008] "Means for analyzing requested data and extracting destination information" refers to an algorithm or program that analyzes voice or text data received from a user and identifies information related to the destination.

[0009] A "server" is a central computer system that communicates with terminals via a network and performs tasks such as searching for destination information and processing generative artificial intelligence.

[0010] "Means of sending to the server" refers to a communication module used to send data from a terminal to a server.

[0011] "Means for searching for details of a relevant destination and generating candidate information" refers to an algorithm or program that uses a database or external API to retrieve and organize detailed destination information based on destination information received by the server.

[0012] "Means for sending candidate information to the terminal" refers to a communication module for sending candidate information generated from the server to the terminal.

[0013] A "terminal" is a device that functions as a user interface and provides information to the user through voice or display.

[0014] "Means for initiating navigation" refers to an algorithm or program that activates the car navigation system based on received destination information and guides the user to the destination.

[0015] "Means for receiving tourist information requests and providing further information" refers to an algorithm or program that processes additional information requests from users during navigation and provides detailed tourist information using generative artificial intelligence. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6]It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the language used in the following description will be explained.

[0019] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0020] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as the user interface and a server that includes data processing and generative artificial intelligence.

[0038] System Overview

[0039] User interface (terminal)

[0040] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0041] server

[0042] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0043] Program Processing Overview

[0044] Request reception and analysis

[0045] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination.

[0046] Requests to the server and searches

[0047] The terminal sends the extracted keywords to the server. Based on the received keywords, the server retrieves detailed destination information from its database or external APIs. Based on the retrieved information, it generates the most suitable candidate information and sends it back to the terminal.

[0048] Presentation and navigation to the user

[0049] The terminal presents the received candidate information to the user via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and starts navigation. During navigation, the terminal guides the user while updating route information in real time.

[0050] Providing tourist information

[0051] During navigation, if the user requests additional information about a tourist destination, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist destination and sends it to the device. The device then provides this information to the user.

[0052] Specific example

[0053] For example, if a user requests by voice, "I want to go to Tokyo Tower," the device converts the voice to text and extracts the keyword "Tokyo Tower." This keyword is sent to the server, which retrieves location information for "Tokyo Tower," information on the nearest parking lot, etc., and sends it to the device. The device then presents the retrieved information to the user and starts navigation. During navigation, if the user requests, "Tell me about the history of Tokyo Tower," the server generates detailed information about its history and provides it to the user through the device. In this way, users can receive destination guidance and sightseeing information with simple voice commands.

[0054] The following describes the processing flow.

[0055] Step 1:

[0056] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[0057] Step 2:

[0058] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[0059] Step 3:

[0060] The terminal analyzes the converted text "I want to go to Tokyo Tower." The natural language processing module extracts the keyword "Tokyo Tower" from the text data.

[0061] Step 4:

[0062] The terminal generates request data containing the extracted keywords and sends it to the server using the communication module.

[0063] Step 5:

[0064] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data from databases and external APIs.

[0065] Step 6:

[0066] Based on the data acquired by the server, destination information (e.g., location information for Tokyo Tower, information on the nearest parking lot, etc.) is generated. The generated information is optimized and sent to the terminal.

[0067] Step 7:

[0068] The terminal outputs destination information received from the server to the display device. The speech synthesis module then informs the user, "Your destination is Tokyo Tower. Let's depart."

[0069] Step 8:

[0070] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[0071] Step 9:

[0072] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[0073] Step 10:

[0074] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection."

[0075] Step 11:

[0076] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[0077] Step 12:

[0078] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[0079] Step 13:

[0080] The server analyzes the additional requests it receives. Generative artificial intelligence retrieves information about the "history of Tokyo Tower" from the database.

[0081] Step 14:

[0082] The server generates tourist information and sends it to the terminal. The terminal then uses a speech synthesis module to guide the user through this information.

[0083] Step 15:

[0084] The device continues navigation, guiding the user to their destination while updating route information in real time.

[0085] (Example 1)

[0086] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0087] Conventional car navigation systems have struggled to efficiently retrieve and provide detailed destination information and tourist information when users request a destination via voice or text. Furthermore, while there is a demand for real-time tourist information in addition to route guidance, the means to achieve this have been insufficient.

[0088] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0089] In this invention, the server includes means for converting speech to text using a speech recognition module, means for extracting keywords using a natural language processing algorithm, and means for retrieving and generating detailed information using a server located in a cloud computing environment. This makes it possible to obtain and provide detailed destination and tourist information in real time when a user requests a destination by voice or text.

[0090] A "user" refers to a person who uses the system to obtain directions to their destination or tourist information.

[0091] A "speech recognition module" refers to a software or hardware component that converts voice input from a user into text data.

[0092] A "natural language processing algorithm" refers to a programmatic method for extracting and analyzing specific keywords from text data.

[0093] A "cloud computing environment" refers to an infrastructure that provides computing resources and data storage via the internet.

[0094] A "server" refers to a computer system located in a central cloud computing environment that processes data and provides information based on user requests.

[0095] An "external API" refers to an interface used to access external databases and services and retrieve information.

[0096] A "terminal" refers to a device that, as part of an in-vehicle system, receives voice or text input from the user and communicates with a server.

[0097] A "navigation system" refers to a system that provides route guidance to a user-specified destination and can update route information in real time.

[0098] "Generative artificial intelligence" refers to advanced AI models used to generate detailed information about destinations and tourist attractions.

[0099] A "prompt" refers to a text-based question or instruction that is input to a generative artificial intelligence system.

[0100] This invention relates to a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as a user interface and a server including data processing and generative artificial intelligence.

[0101] User interface (terminal)

[0102] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0103] server

[0104] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. As a specific example of the software, a high-performance speech recognition module (e.g., Google® Cloud Speech-to-Text API) is used for speech recognition, NLP algorithms are used for natural language processing, and external APIs (e.g., Google Places API) are used for searching for detailed information.

[0105] Specifically, if a user requests by voice, "I want to go to Tokyo Tower," the terminal converts this voice into text and extracts the keyword "Tokyo Tower." The extracted keyword is sent to the server as a data packet. The server retrieves information based on the received keyword through a database or external API, organizes the retrieved information, and sends it to the terminal. The terminal displays the retrieved information on its screen or presents it to the user using a voice output module. After the user confirms the destination, the terminal configures the car navigation system and starts navigation. The navigation system guides the user to the destination while updating route information in real time.

[0106] If the user requests again during navigation to "Tell me about the history of Tokyo Tower," the device will send another request to the server. The server will use generative artificial intelligence (for example, a model from OpenAI®) to generate detailed historical information and send it to the device. The device will then provide this information to the user in voice or text.

[0107] Example of a prompt

[0108] Examples of prompt statements include the following:

[0109] "Please describe a program that provides detailed information about the location specified by the user as their destination and then performs navigation."

[0110] "Please explain, with specific examples, the processing steps of a system that provides tourist information based on user requests."

[0111] As described above, the present invention efficiently processes voice or text requests from users and provides detailed navigation and tourist information in real time.

[0112] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0113] Step 1:

[0114] A user submits a voice request.

[0115] The user sends a voice request from inside the car saying, "I want to go to Tokyo Tower."

[0116] Input: User voice.

[0117] Output: Acquisition of audio data by the terminal.

[0118] Step 2:

[0119] The device converts speech to text.

[0120] The device receives audio using its built-in microphone and converts the audio into text data using a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API).

[0121] Input: Audio data.

[0122] Output: Text data.

[0123] Specific operation: The speech recognition module analyzes the speech pattern and outputs it as a string.

[0124] Step 3:

[0125] The device extracts keywords from the text.

[0126] The device uses a natural language processing algorithm (e.g., NLP) to extract keywords like "Tokyo Tower" from text data.

[0127] Input: Text data.

[0128] Output: Keywords.

[0129] Specific operation: An NLP algorithm analyzes the context and structure of the text to identify the destination.

[0130] Step 4:

[0131] The device sends the keyword to the server.

[0132] The terminal compiles the extracted keywords into a data packet and sends it to the server via the communication module.

[0133] Input: Keyword.

[0134] Output: Data packet containing the keyword.

[0135] Specific operation: The communication module uploads the keyword to the server.

[0136] Step 5:

[0137] The server searches for information based on keywords.

[0138] Based on the received keywords, the server searches for detailed destination information using its database and external APIs (e.g., Google Places API).

[0139] Input: Data packet containing keywords.

[0140] Output: Detailed destination information (location, parking information, etc.).

[0141] Specific operations: Database search and data retrieval from external APIs.

[0142] Step 6:

[0143] The server sends the search results to the device.

[0144] The server organizes the search results, combines location information and nearest parking information into a single data packet, and sends it to the terminal.

[0145] Input: Destination details.

[0146] Output: Data packet containing search results.

[0147] Specific operation: Configures a data packet and sends it to the terminal.

[0148] Step 7:

[0149] The device displays search results to the user.

[0150] The terminal displays the received information on its screen or notifies the user via voice through an audio output module.

[0151] Input: Data packet containing search results.

[0152] Output: Information presented to the user.

[0153] Specific actions: Display the results on the screen or communicate them via voice.

[0154] Step 8:

[0155] The user confirms the destination.

[0156] The user reviews the information provided and selects "Tokyo Tower" as their destination.

[0157] Input: The search results presented.

[0158] Output: Action to confirm destination.

[0159] Specific actions: The user operates the display or audio confirmation button.

[0160] Step 9:

[0161] The device starts navigation.

[0162] The device sets Tokyo Tower as the destination in the car's navigation system and begins navigation. The navigation system (e.g., Garmin or TomTom) provides the latest route information in real time and guides the user to the destination.

[0163] Input: Action to confirm destination.

[0164] Output: Navigation started.

[0165] Specific operation: The navigation system calculates and displays the route from the current location to the destination.

[0166] Step 10:

[0167] The user requests additional information.

[0168] During navigation, the user requests additional information via voice, saying, "Tell me about the history of Tokyo Tower."

[0169] Input: Voice request for additional information.

[0170] Output: Acquisition of audio data by the terminal.

[0171] Specific operation: The device receives audio via the microphone and recognizes the audio pattern.

[0172] Step 11:

[0173] The device sends the request to the server.

[0174] The terminal converts the user's request into text data and sends it back to the server.

[0175] Input: Audio data of the request for additional information.

[0176] Output: Data packet containing the request.

[0177] Specific operation: The speech recognition module converts speech to text and sends it to the server.

[0178] Step 12:

[0179] The server generates additional information.

[0180] The server uses generative artificial intelligence (e.g., an AI model) to generate detailed information about the "history of Tokyo Tower" as requested by the user.

[0181] Input: Data packet containing the request.

[0182] Output: Generated detailed information.

[0183] Specific operation: The artificial intelligence model generates information and outputs it as data.

[0184] Step 13:

[0185] The device provides additional information to the user.

[0186] The terminal provides the user with the received detailed information via an audio output module or display.

[0187] Input: Data packet containing detailed information.

[0188] Output: Information presented to the user.

[0189] Specific actions: Display information on the screen or provide audio notifications.

[0190] (Application Example 1)

[0191] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0192] Conventional navigation systems make it difficult for users to obtain destination information and tourist information while operating the vehicle, and often fail to provide appropriate guidance, especially during long-distance drives or when visiting unfamiliar places. Furthermore, current car navigation systems rely heavily on manual operation, which poses a problem in terms of the efficiency of information provision to users in autonomous vehicles.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0194] In this invention, the server includes means for the user to request a destination by voice or text, means for analyzing the requested data and extracting destination information, means for transmitting the extracted destination information to the server, means for searching for details of the relevant destination based on the destination information received by the server and generating candidate information, means for transmitting the generated candidate information to a terminal, means for the terminal to present the received candidate information to the user by voice or display and start navigation, means for receiving requests for sightseeing information from the user during navigation and providing further information, means installed in an autonomous vehicle, which processes voice input using a terminal as a user interface, and means for communicating with a cloud server to generate destination information and sightseeing information and provide it to the user. As a result, the user can efficiently and safely receive destination guidance and sightseeing information while riding in an autonomous vehicle.

[0195] "Voice input" is a technology that acquires voice data and converts it into a digital format.

[0196] "Text input" refers to the technology of entering text information using input devices such as keyboards and touch panels.

[0197] "Destination information" refers to detailed data about a location specified by the user, including location information and related tourist information.

[0198] A "server" is a computer system located in a cloud computing environment that provides information using data processing and generative artificial intelligence.

[0199] A "terminal" is a device that receives input from a user and displays information, and is often installed as part of an in-vehicle system.

[0200] "Navigation" is a function that calculates and guides the user along the optimal route to reach their destination.

[0201] "Tourist information" refers to detailed data about tourist attractions, history, and facilities in a destination or its surrounding area.

[0202] An "autonomous vehicle" is a vehicle equipped with autonomous driving technology that can drive autonomously without user intervention.

[0203] "Generative artificial intelligence" refers to AI technology that analyzes and generates data, automatically creating and providing destination information and tourist guides.

[0204] A "cloud server" is a server used remotely via the internet, and it is a computing environment for processing and storing large amounts of data.

[0205] The system of this invention has a configuration for realizing destination guidance and sightseeing information within an autonomous vehicle. A specific embodiment thereof is described below.

[0206] System Configuration

[0207] User interface (terminal)

[0208] The terminal is installed inside the autonomous vehicle and functions as a user input interface. The terminal includes the following hardware:

[0209] Microphone: A device used to acquire user voice input.

[0210] Display: A monitor used to display destination information and tourist information.

[0211] Built-in computer: A processing unit for speech recognition and data analysis.

[0212] server

[0213] The server is located in a cloud environment and operates using the following software and APIs.

[0214] Speech recognition module: Software that converts speech input into text.

[0215] Generative artificial intelligence (generative AI): AI models that perform data analysis and information generation.

[0216] Database: A storage system for storing destination information and tourist guides.

[0217] Communication module (such as the requests library): A module for sending and receiving data between a terminal and a cloud server.

[0218] Processing flow

[0219] 1. Voice input

[0220] The user makes a voice request inside the autonomous vehicle, saying "I want to go to XX." The microphone captures the voice, and the built-in computer converts it into text using a speech recognition module.

[0221] 2. Data transmission

[0222] The converted text data is sent to a cloud server. The server uses generative artificial intelligence to extract destination information from the submitted keywords and generate relevant tourist information.

[0223] 3. Information presentation

[0224] The generated destination information and tourist information are sent to the terminal. The terminal displays the information on its screen and provides guidance to the user via voice or text.

[0225] 4. Start Navigation

[0226] The autonomous vehicle calculates a route based on the destination information it has acquired and then begins autonomous driving. During navigation, the user can request additional sightseeing information, and the corresponding information will be generated and provided.

[0227] Specific example

[0228] For example, if a user requests by voice, "I want to go to Shibuya Station," the microphone collects the voice, and the built-in computing system converts the voice into text. The converted keyword "Shibuya Station" is sent to a cloud server, which generates location information for "Shibuya Station" and nearby tourist information. This information is returned to the device, displayed on the screen, and navigation begins.

[0229] Example of a prompt

[0230] User input: "I would like directions to Shibuya Station and some sightseeing information."

[0231] System response:

[0232] "Calculating the route to Shibuya Station. Performing voice recognition to continue..."

[0233] "This is the route to Shibuya Station. Near Shibuya Station, you'll find tourist attractions such as the Shibuya Scramble Crossing, Center Gai, and Shibuya Hikarie."

[0234] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0235] Step 1:

[0236] The user enters a voice request into the microphone inside the autonomous vehicle, saying "I want to go to XX." The terminal acquires the voice data, and the built-in computer's voice recognition module converts it into text data. The input is voice data, and the output is text data.

[0237] Step 2:

[0238] The terminal's built-in computer analyzes the text data output from the speech recognition module and extracts keywords related to the destination. At this stage, the input is the text data obtained in step 1, and the output is the extracted keywords.

[0239] Step 3:

[0240] The terminal sends the extracted keywords to the cloud server. A communication module (e.g., the requests library) is used to send and receive data with the cloud server. The input for this step is the extracted keywords, and the output is the result of sending a request to the server.

[0241] Step 4:

[0242] The server analyzes the received keywords using generative artificial intelligence (generative AI model) and retrieves detailed information about the corresponding destination from a database or external API. After retrieving the information, the server generates optimal candidate information using the generative AI. Here, the input is the received keywords, and the output is the generated destination candidate information.

[0243] Step 5:

[0244] The server sends the generated destination candidate information to the terminal. Data is exchanged in real time using a communication module. The input is the generated candidate information, and the output is the result of the information transmission to the terminal.

[0245] Step 6:

[0246] The terminal presents the received candidate information to the user via voice or display. The user confirms the destination by checking the display or voice guidance. The input is candidate information from the server, and the output is destination candidate information presented to the user.

[0247] Step 7:

[0248] Once the user confirms the destination, the terminal sets the destination information in the autonomous vehicle's navigation system and starts navigation. The input is the destination information confirmed by the user, and the output is the autonomous vehicle's route calculation and driving instructions.

[0249] Step 8:

[0250] During navigation, if the user requests additional sightseeing information, the device sends another request to the server. The server uses a generative AI model to generate sightseeing information and sends it to the device. The input for this step is the user's sightseeing information request, and the output is the generated sightseeing information.

[0251] Step 9:

[0252] The terminal provides the user with received tourist information. It presents information to the user using a display and voice guidance, providing guidance in real time. The input is tourist information from the server, and the output is detailed tourist information presented to the user.

[0253] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0254] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[0255] System Overview

[0256] User interface (terminal)

[0257] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0258] server

[0259] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0260] Emotional Engine

[0261] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[0262] Program Processing Overview

[0263] Request reception and analysis

[0264] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination. In addition, an emotion engine analyzes the user's emotions from the voice or text.

[0265] Requests to the server and searches

[0266] The device sends request data to the server, which includes keywords extracted by the device and analyzed sentiment data. Based on the received keywords, the server retrieves destination details from its database or external APIs and generates candidate information that takes sentiment data into consideration. The generated information is then sent back to the device.

[0267] Presentation and navigation to the user

[0268] The terminal presents the user with received candidate information via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and begins navigation. During navigation, the terminal guides the user while updating route information in real time. It also adjusts the tone and content of the guidance according to the user's emotional state.

[0269] Providing tourist information

[0270] During navigation, if the user requests additional information about a tourist spot, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist spot and sends it to the device, taking sentiment data into consideration. The device then provides this information to the user and adjusts the tone and content of the guidance according to their emotions.

[0271] Specific example

[0272] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[0273] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[0274] The following describes the processing flow.

[0275] Step 1:

[0276] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[0277] Step 2:

[0278] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[0279] Step 3:

[0280] The device analyzes the converted text, "I want to go to Tokyo Tower." A natural language processing module extracts the keyword "Tokyo Tower" from the text data. Additionally, an emotion engine analyzes the user's emotions (e.g., excitement, fatigue) from the audio and text data.

[0281] Step 4:

[0282] The terminal generates request data containing extracted keywords and analyzed sentiment data, and sends it to the server using a communication module.

[0283] Step 5:

[0284] The server analyzes the received request. The generative artificial intelligence recognizes "Tokyo Tower" and obtains related data (location information, nearby facilities, parking lot information, etc.) from the database or external APIs. It also generates guidance information considering the sentiment data.

[0285] Step 6:

[0286] Based on the data obtained by the server, candidate information corresponding to the destination information and sentiment is generated and sent to the terminal.

[0287] Step 7:

[0288] The terminal outputs the destination information and candidate information received from the server to the display device. The speech synthesis module guides the user with "The destination is Tokyo Tower. Let's go." The tone and content of the guidance are adjusted according to the sentiment data.

[0289] Step 8:

[0290] The user confirms the guidance and gives approval with a voice input such as "Yes". The terminal receives this voice input and converts it into text.

[0291] Step 9:

[0292] The terminal confirms the user's approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[0293] Step 10:

[0294] The terminal starts the navigation. It gives instructions to the user such as "Turn left at the next intersection" with voice guidance. During navigation, the emotion engine continuously analyzes the user's emotion and adjusts the tone and content of the guidance.

[0295] Step 11:

[0296] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[0297] Step 12:

[0298] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[0299] Step 13:

[0300] The server analyzes the additional requests it receives, and a generative artificial intelligence retrieves tourist information (for example, the history of Tokyo Tower) from databases and external APIs. Based on the user's emotions analyzed by the emotion engine, the tone and content of the generated guide information are adjusted.

[0301] Step 14:

[0302] The server generates tourist information and sends corresponding guidance information to the terminal. The terminal then uses a speech synthesis module to provide this information to the user.

[0303] Step 15:

[0304] The device continues navigation, guiding the user to their destination while updating route information in real time. The tone and content of the guidance continue to be adjusted based on the user's emotions.

[0305] (Example 2)

[0306] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0307] Conventional car navigation systems simply provided route guidance to the destination and lacked the ability to provide personalized guidance information according to the user's emotional state. For this reason, user satisfaction was low, and it was not particularly suitable for tourism guidance intended for relaxation or excitement. Also, if the requested tourism information did not match the user's emotional state, the entire travel experience could become unpleasant.

[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following respective means.

[0309] In this invention, the server includes means for analyzing the requested data and extracting destination information, means for transmitting the extracted destination information and the analyzed user emotion data to the server, and means for searching for details of the corresponding destination based on the received destination information and emotion data and generating candidate information according to the emotion using a generated artificial intelligence. Thereby, personalized tourism guidance and navigation adapted to the user's emotional state become possible.

[0310] The "user" refers to a user of the system who requests a destination or tourism information using voice or text.

[0311] The "terminal" refers to a device installed as an in-vehicle system that receives voice input and text input from the user and processes request data.

[0312] [[ID=第十九]] The "voice recognition module" refers to software or hardware for converting the user's voice input into text data.

[0313] The "emotion engine" refers to a technology for analyzing the emotional state from the user's voice or text input and outputting the analysis result.

[0314] The "server" refers to a device arranged in a central computer environment that processes the requested data and generates detailed information about the destination and tourism guidance.

[0315] A "generative AI model" refers to artificial intelligence technology that generates personalized tourist information and navigation information based on the user's emotional state.

[0316] A "database" refers to a data management system that stores detailed information and tourist information about a destination, and allows users to search for and retrieve that information as needed.

[0317] An "external API" refers to an application program interface used to retrieve information by interacting with other services or databases.

[0318] A "navigation system" refers to a device and software that provides route guidance to a destination and updates route information in real time.

[0319] "Personalized guidance" refers to guidance information whose content and tone are adjusted based on the user's current emotional state.

[0320] This invention relates to a system that allows users to request destinations via voice or text, provides car navigation and tourist information based on those requests, analyzes the user's emotions, and provides personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[0321] User interface (terminal)

[0322] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and has built-in speech recognition and text conversion capabilities. It can also exchange data with a server in real time via a communication module. Specifically, the terminal uses the Google Speech-to-Text API to convert speech to text and further uses an emotion engine to analyze the user's emotions from the speech and text.

[0323] server

[0324] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. Specifically, destination data is obtained using the Google Maps API and MongoDB, and tourist information is generated using generative AI models such as OpenAI's GPT-4 (registered trademark).

[0325] Emotional Engine

[0326] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[0327] Specific example

[0328] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[0329] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[0330] Prompt example

[0331] The prompt message when a user requests to go to Tokyo Tower.

[0332] "The user requested to visit Tokyo Tower, and the emotion engine detected excitement. Please generate an energetic sightseeing guide suitable for the excited user."

[0333] The prompt text when a user requests, "Tell me about the history of Tokyo Tower."

[0334] "The user requested information about the history of Tokyo Tower, and the emotion engine detected fatigue. Please generate a gentle, user-friendly tourist guide."

[0335] In this way, the present invention allows users to receive more satisfying, personalized navigation and tourist information.

[0336] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0337] Step 1:

[0338] The user requests the destination via voice or text.

[0339] Input: User voice or text request (e.g., "I want to go to Tokyo Tower").

[0340] Specific action: The user speaks to the in-car system, saying, "I want to go to Tokyo Tower."

[0341] Output: The device receives the audio data.

[0342] Step 2:

[0343] The device converts voice input into text using a speech recognition module.

[0344] Input: User voice data received by the terminal.

[0345] Specific operation: The speech recognition module installed in the device (e.g., Google Speech-to-Text API) converts speech to text.

[0346] Output: Text data (e.g., "I want to go to Tokyo Tower").

[0347] Step 3:

[0348] The terminal analyzes the request content and extracts destination information.

[0349] Input: Converted text data.

[0350] Specific operation: The terminal uses a text analysis module to extract keywords related to the destination (e.g., "Tokyo Tower").

[0351] Output: Extracted keywords (e.g., "Tokyo Tower").

[0352] Step 4:

[0353] The device sends the converted text and audio data to the emotion engine for sentiment analysis.

[0354] Input: Converted text data and audio data.

[0355] Specific operation: The emotion engine analyzes the voice tone and text content to determine the user's emotional state (e.g., "excited").

[0356] Output: Sentiment data (e.g., "excited").

[0357] Step 5:

[0358] The device sends the extracted keywords and analyzed sentiment data to the server.

[0359] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[0360] Specific operation: The terminal structures keyword and sentiment data, generates data packets, and sends them to the server.

[0361] Output: The server receives the request data.

[0362] Step 6:

[0363] The server retrieves destination information from external APIs and databases based on the keywords it receives, and uses a generative AI model to generate candidate information that reflects emotions.

[0364] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[0365] Specific operation: The server uses the Google Maps API or similar to retrieve detailed information about "Tokyo Tower," and then uses a generative AI model like OpenAI's GPT-4 to generate a guide text appropriate for the emotion "excitement."

[0366] Output: Generated candidate information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[0367] Step 7:

[0368] The server sends the candidate information it generates to the terminal.

[0369] Input: Generated candidate information (Example: Guide text "Tokyo Tower is a very popular tourist spot, and the night view is especially wonderful!").

[0370] Specific operation: The server structures the candidate information as a data packet and sends it to the terminal.

[0371] Output: The terminal receives candidate information.

[0372] Step 8:

[0373] The terminal presents the received candidate information to the user via voice or display.

[0374] Input: Received suggested information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[0375] Specific operation: The device uses a speech synthesis module and display function to present instructions to the user.

[0376] Output: The user reviews the suggested information.

[0377] Step 9:

[0378] The user confirms the destination, the device sets the destination in the car's navigation system, and starts navigation.

[0379] Input: User confirmation instructions (e.g., "Set Tokyo Tower as destination").

[0380] Specific operation: The device receives user instructions, sets the destination "Tokyo Tower" in the car navigation system, and starts navigation.

[0381] Output: The car navigation system begins route guidance.

[0382] Step 10:

[0383] During navigation, the device guides the user while updating route information in real time.

[0384] Input: Real-time data from the car navigation system.

[0385] Specific operation: The terminal receives real-time data from the navigation system and guides the user while updating route information.

[0386] Output: Guidance based on the latest route information.

[0387] Step 11:

[0388] During navigation, the user requests additional information about tourist attractions.

[0389] Input: User voice or text request (e.g., "Tell me about the history of Tokyo Tower").

[0390] Specific operation: The user requests additional information from the in-vehicle system via voice or text.

[0391] Output: The device receives voice or text data.

[0392] Step 12:

[0393] The device receives the request, converts it to text using a speech recognition module, and performs sentiment analysis using an emotion engine.

[0394] Input: User voice data (e.g., "Tell me about the history of Tokyo Tower").

[0395] Specific operation: The device converts the voice data into text using a speech recognition module, and then performs sentiment analysis using an emotion engine (e.g., "fatigue").

[0396] Output: Text data and sentiment data (e.g., "fatigue").

[0397] Step 13:

[0398] The terminal sends the request data to the server again along with the analysis results.

[0399] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[0400] Specific operation: The device structures text data and sentiment data and sends them to the server as data packets.

[0401] Output: The server receives the request data.

[0402] Step 14:

[0403] The server uses an AI model to generate tourist information that responds to emotions.

[0404] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[0405] Specific operation: The server retrieves information from an internal database or external API, and a generative AI model generates a gentle, user-friendly guidance message suitable for "fatigue."

[0406] Output: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[0407] Step 15:

[0408] The server generates a message and sends it to the terminal.

[0409] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[0410] Specific operation: The server structures the generated guidance message into a data packet and sends it to the terminal.

[0411] Output: The terminal receives the generated guidance message.

[0412] Step 16:

[0413] The terminal provides the user with the received notification message, and adjusts the tone and content of the message according to the user's emotions.

[0414] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[0415] Specific operation: The device uses a speech synthesis module and display function to present guidance text to the user in an appropriate tone.

[0416] Output: Users understand and become interested in tourist information.

[0417] (Application Example 2)

[0418] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0419] Conventional car navigation and tourist information systems only provide destinations and tourist information based on user requests, and are unable to provide personalized guidance that takes into account the user's emotional state. This results in a lack of improved user experience. Furthermore, methods for integrating and providing multiple pieces of information are insufficient, making it difficult to provide appropriate information based on the user's emotions.

[0420] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing emotions from the user's voice or text input, means for adjusting the tone and content of the guidance based on the emotion analysis results, and means for generating tourist information using generative artificial intelligence on the server and providing the user with personalized guidance based on emotion data. This makes it possible to provide personalized guidance according to the user's emotional state.

[0421] A "user" refers to a person who uses the system to request destinations or receive tourist information.

[0422] "Voice or text" refers to a means by which a user inputs instructions to the system, and includes both voice input and text input.

[0423] A "request" refers to an instruction or request that a user makes to a system.

[0424] "Analyzing data" refers to the process of understanding user requests and extracting their content.

[0425] "Destination information" refers to detailed data related to the destination specified by the user through the system.

[0426] A "server" is a central computer system that uses data processing and generation artificial intelligence to retrieve and generate information based on user requests.

[0427] "Extraction" refers to the act of selecting destination information from request data.

[0428] "Searching" refers to the process by which a server retrieves relevant information from external databases or APIs based on destination information it has received.

[0429] "Candidate information" refers to multiple suggestions or options generated by the server to respond to a user's request.

[0430] A "terminal" refers to a device that provides a user interface and allows users to access a system.

[0431] "Navigation" refers to a guide function that directs the user to a specified destination.

[0432] "Tourism information" refers to additional information about tourist attractions and facilities related to the destination.

[0433] "Emotion" refers to the emotional state analyzed from the user's voice and text.

[0434] "Emotion analysis" refers to the process of identifying a user's emotional state from their input data.

[0435] "Generative artificial intelligence" refers to artificial intelligence technology that generates tourist information based on user requests and sentiment data.

[0436] "Personalization" refers to adjusting the content of information and guidance provided according to each user's emotional state and preferences.

[0437] "The tone and content of the guidance" refers to the tone and specific content of the information provided to the user, and is adjusted based on their emotional state.

[0438] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[0439] User interface (terminal)

[0440] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0441] server

[0442] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0443] Emotional Engine

[0444] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[0445] Specific examples

[0446] Processing voice requests

[0447] If a user requests to go to Senso-ji Temple by voice while inside the vehicle, the voice is input through the device's microphone and converted into text by a voice recognition module. Next, this text data is sent to an emotion engine for emotion analysis. Let's say "joy" is detected at this point.

[0448] Data processing on the server

[0449] The analyzed sentiment data and text data are integrated and sent to the server. The server identifies the destination "Senso-ji Temple" from the received text data and generates personalized tourist information based on the sentiment data. For example, it might generate information such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple."

[0450] Provide feedback

[0451] The generated personalized guidance information is sent back to the device, which then presents the guidance to the user via voice or display. At this time, the tone and content of the guidance are adjusted based on emotional data.

[0452] As described above, this system allows users to receive not only navigation to their destination, but also optimal sightseeing information tailored to their mood at the time.

[0453] Example of a prompt

[0454] "Please analyze the following voice request and generate appropriate, emotion-sensitive tourist information:

[0455] Voice request: "I want to go to Senso-ji Temple."

[0456] Emotion analysis result: "Joy"

[0457] In this way, by providing guidance tailored to the user's emotions, it becomes possible to offer a more personalized navigation experience.

[0458] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0459] Step 1:

[0460] The user requests a destination by voice or text. When the user speaks into the in-car microphone, saying "I want to go to Senso-ji Temple," the voice is picked up by the device's microphone.

[0461] Input: Voice Request

[0462] Output: Audio data

[0463] Step 2:

[0464] The device's speech recognition module converts the audio data into text data. The speech recognition module (for example, the Vosk library) analyzes the audio data and generates the text "I want to go to Senso-ji Temple."

[0465] Input: Audio data

[0466] Output: Text data

[0467] Step 3:

[0468] The emotion engine analyzes the user's emotions based on text data. The emotion engine (for example, EmotionRecognizer) analyzes text data and voice tone information to identify the emotion of "joy" in this case.

[0469] Input: Text data, voice tone

[0470] Output: Sentiment data

[0471] Step 4:

[0472] The terminal sends generated text data and sentiment data to the server. A communication module is used to send request data to the server.

[0473] Input: Text data, sentiment data

[0474] Output: Request data

[0475] Step 5:

[0476] The server processes the request data and generates tourist information based on destination and sentiment data. The server's data processing module analyzes the request data, retrieves tourist information for "Senso-ji Temple" from an external API, and then a generative artificial intelligence adjusts the personalized guidance.

[0477] Input: Request data

[0478] Output: Personalized tourist information

[0479] Step 6:

[0480] The server sends the generated personalized tourist information to the terminal. The server sends the tourist information to the terminal via a communication module.

[0481] Input: Personalized tourist information

[0482] Output: Transmitted data

[0483] Step 7:

[0484] The terminal presents tourist information received by the user via voice or display. A speech synthesis engine (e.g., pyttsx3) generates a guidance message, such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple," which is presented aloud.

[0485] Input: Data to send

[0486] Output: Audio or display

[0487] Step 8:

[0488] The system checks the information received by the user and inputs instructions to start navigation. When the user gives a voice command such as "Start navigation," that command is entered into the terminal.

[0489] Input: User instructions

[0490] Output: Navigation start instruction

[0491] Step 9:

[0492] The terminal activates the navigation system and provides directions to the destination. The navigation module calculates the optimal route based on the current location and destination, and provides real-time guidance.

[0493] Input: Navigation start command

[0494] Output: Navigation Guide

[0495] In this way, the system analyzes the user's emotions based on their request, provides optimal sightseeing information, and initiates navigation.

[0496] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0497] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0498] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0499] [Second Embodiment]

[0500] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0501] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0502] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0503] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0504] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0505] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0506] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0507] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0508] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0509] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0510] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0511] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0512] This invention is a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as the user interface and a server that includes data processing and generative artificial intelligence.

[0513] System Overview

[0514] User interface (terminal)

[0515] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0516] server

[0517] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0518] Program Processing Overview

[0519] Request reception and analysis

[0520] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination.

[0521] Requests to the server and searches

[0522] The terminal sends the extracted keywords to the server. Based on the received keywords, the server retrieves detailed destination information from its database or external APIs. Based on the retrieved information, it generates the most suitable candidate information and sends it back to the terminal.

[0523] Presentation and navigation to the user

[0524] The terminal presents the received candidate information to the user via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and starts navigation. During navigation, the terminal guides the user while updating route information in real time.

[0525] Providing tourist information

[0526] During navigation, if the user requests additional information about a tourist destination, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist destination and sends it to the device. The device then provides this information to the user.

[0527] Specific example

[0528] For example, if a user requests by voice, "I want to go to Tokyo Tower," the device converts the voice to text and extracts the keyword "Tokyo Tower." This keyword is sent to the server, which retrieves location information for "Tokyo Tower," information on the nearest parking lot, etc., and sends it to the device. The device then presents the retrieved information to the user and starts navigation. During navigation, if the user requests, "Tell me about the history of Tokyo Tower," the server generates detailed information about its history and provides it to the user through the device. In this way, users can receive destination guidance and sightseeing information with simple voice commands.

[0529] The following describes the processing flow.

[0530] Step 1:

[0531] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[0532] Step 2:

[0533] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[0534] Step 3:

[0535] The terminal analyzes the converted text "I want to go to Tokyo Tower." The natural language processing module extracts the keyword "Tokyo Tower" from the text data.

[0536] Step 4:

[0537] The terminal generates request data containing the extracted keywords and sends it to the server using the communication module.

[0538] Step 5:

[0539] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data from databases and external APIs.

[0540] Step 6:

[0541] Based on the data acquired by the server, destination information (e.g., location information for Tokyo Tower, information on the nearest parking lot, etc.) is generated. The generated information is optimized and sent to the terminal.

[0542] Step 7:

[0543] The terminal outputs destination information received from the server to the display device. The speech synthesis module then informs the user, "Your destination is Tokyo Tower. Let's depart."

[0544] Step 8:

[0545] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[0546] Step 9:

[0547] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[0548] Step 10:

[0549] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection."

[0550] Step 11:

[0551] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[0552] Step 12:

[0553] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[0554] Step 13:

[0555] The server analyzes the additional requests it receives. Generative artificial intelligence retrieves information about the "history of Tokyo Tower" from the database.

[0556] Step 14:

[0557] The server generates tourist information and sends it to the terminal. The terminal then uses a speech synthesis module to guide the user through this information.

[0558] Step 15:

[0559] The device continues navigation, guiding the user to their destination while updating route information in real time.

[0560] (Example 1)

[0561] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0562] Conventional car navigation systems have struggled to efficiently retrieve and provide detailed destination information and tourist information when users request a destination via voice or text. Furthermore, while there is a demand for real-time tourist information in addition to route guidance, the means to achieve this have been insufficient.

[0563] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0564] In this invention, the server includes means for converting speech to text using a speech recognition module, means for extracting keywords using a natural language processing algorithm, and means for retrieving and generating detailed information using a server located in a cloud computing environment. This makes it possible to obtain and provide detailed destination and tourist information in real time when a user requests a destination by voice or text.

[0565] A "user" refers to a person who uses the system to obtain directions to their destination or tourist information.

[0566] A "speech recognition module" refers to a software or hardware component that converts voice input from a user into text data.

[0567] A "natural language processing algorithm" refers to a programmatic method for extracting and analyzing specific keywords from text data.

[0568] A "cloud computing environment" refers to an infrastructure that provides computing resources and data storage via the internet.

[0569] A "server" refers to a computer system located in a central cloud computing environment that processes data and provides information based on user requests.

[0570] An "external API" refers to an interface used to access external databases and services and retrieve information.

[0571] A "terminal" refers to a device that, as part of an in-vehicle system, receives voice or text input from the user and communicates with a server.

[0572] A "navigation system" refers to a system that provides route guidance to a user-specified destination and can update route information in real time.

[0573] "Generative artificial intelligence" refers to advanced AI models used to generate detailed information about destinations and tourist attractions.

[0574] A "prompt" refers to a text-based question or instruction that is input to a generative artificial intelligence system.

[0575] This invention relates to a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as a user interface and a server including data processing and generative artificial intelligence.

[0576] User interface (terminal)

[0577] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0578] server

[0579] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating suggested information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. As a specific example of the software, a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API) is used for speech recognition, NLP algorithms are used for natural language processing, and external APIs (e.g., Google Places API) are used for searching for detailed information.

[0580] Specifically, if a user requests by voice, "I want to go to Tokyo Tower," the terminal converts this voice into text and extracts the keyword "Tokyo Tower." The extracted keyword is sent to the server as a data packet. The server retrieves information based on the received keyword through a database or external API, organizes the retrieved information, and sends it to the terminal. The terminal displays the retrieved information on its screen or presents it to the user using a voice output module. After the user confirms the destination, the terminal configures the car navigation system and starts navigation. The navigation system guides the user to the destination while updating route information in real time.

[0581] If the user requests again during navigation, "Tell me about the history of Tokyo Tower," the device will send another request to the server. The server will use generative artificial intelligence (for example, an OpenAI model) to generate detailed historical information and send it to the device. The device will then provide this information to the user in voice or text.

[0582] Example of a prompt

[0583] Examples of prompt statements include the following:

[0584] "Please describe a program that provides detailed information about the location specified by the user as their destination and then performs navigation."

[0585] "Please explain, with specific examples, the processing steps of a system that provides tourist information based on user requests."

[0586] As described above, the present invention efficiently processes voice or text requests from users and provides detailed navigation and tourist information in real time.

[0587] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0588] Step 1:

[0589] A user submits a voice request.

[0590] The user sends a voice request from inside the car saying, "I want to go to Tokyo Tower."

[0591] Input: User voice.

[0592] Output: Acquisition of audio data by the terminal.

[0593] Step 2:

[0594] The device converts speech to text.

[0595] The device receives audio using its built-in microphone and converts the audio into text data using a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API).

[0596] Input: Audio data.

[0597] Output: Text data.

[0598] Specific operation: The speech recognition module analyzes the speech pattern and outputs it as a string.

[0599] Step 3:

[0600] The device extracts keywords from the text.

[0601] The device uses a natural language processing algorithm (e.g., NLP) to extract keywords like "Tokyo Tower" from text data.

[0602] Input: Text data.

[0603] Output: Keywords.

[0604] Specific operation: An NLP algorithm analyzes the context and structure of the text to identify the destination.

[0605] Step 4:

[0606] The device sends the keyword to the server.

[0607] The terminal compiles the extracted keywords into a data packet and sends it to the server via the communication module.

[0608] Input: Keyword.

[0609] Output: Data packet containing the keyword.

[0610] Specific operation: The communication module uploads the keyword to the server.

[0611] Step 5:

[0612] The server searches for information based on keywords.

[0613] Based on the received keywords, the server searches for detailed destination information using its database and external APIs (e.g., Google Places API).

[0614] Input: Data packet containing keywords.

[0615] Output: Detailed destination information (location, parking information, etc.).

[0616] Specific operations: Database search and data retrieval from external APIs.

[0617] Step 6:

[0618] The server sends the search results to the device.

[0619] The server organizes the search results, combines location information and nearest parking information into a single data packet, and sends it to the terminal.

[0620] Input: Destination details.

[0621] Output: Data packet containing search results.

[0622] Specific operation: Configures a data packet and sends it to the terminal.

[0623] Step 7:

[0624] The device displays search results to the user.

[0625] The terminal displays the received information on its screen or notifies the user via voice through an audio output module.

[0626] Input: Data packet containing search results.

[0627] Output: Information presented to the user.

[0628] Specific actions: Display the results on the screen or communicate them via voice.

[0629] Step 8:

[0630] The user confirms the destination.

[0631] The user reviews the information provided and selects "Tokyo Tower" as their destination.

[0632] Input: The search results presented.

[0633] Output: Action to confirm destination.

[0634] Specific actions: The user operates the display or audio confirmation button.

[0635] Step 9:

[0636] The device starts navigation.

[0637] The device sets Tokyo Tower as the destination in the car's navigation system and begins navigation. The navigation system (e.g., Garmin or TomTom) provides the latest route information in real time and guides the user to the destination.

[0638] Input: Action to confirm destination.

[0639] Output: Navigation started.

[0640] Specific operation: The navigation system calculates and displays the route from the current location to the destination.

[0641] Step 10:

[0642] The user requests additional information.

[0643] During navigation, the user requests additional information via voice, saying, "Tell me about the history of Tokyo Tower."

[0644] Input: Voice request for additional information.

[0645] Output: Acquisition of audio data by the terminal.

[0646] Specific operation: The device receives audio via the microphone and recognizes the audio pattern.

[0647] Step 11:

[0648] The device sends the request to the server.

[0649] The terminal converts the user's request into text data and sends it back to the server.

[0650] Input: Audio data of the request for additional information.

[0651] Output: Data packet containing the request.

[0652] Specific operation: The speech recognition module converts speech to text and sends it to the server.

[0653] Step 12:

[0654] The server generates additional information.

[0655] The server uses generative artificial intelligence (e.g., an AI model) to generate detailed information about the "history of Tokyo Tower" as requested by the user.

[0656] Input: Data packet containing the request.

[0657] Output: Generated detailed information.

[0658] Specific operation: The artificial intelligence model generates information and outputs it as data.

[0659] Step 13:

[0660] The device provides additional information to the user.

[0661] The terminal provides the user with the received detailed information via an audio output module or display.

[0662] Input: Data packet containing detailed information.

[0663] Output: Information presented to the user.

[0664] Specific actions: Display information on the screen or provide audio notifications.

[0665] (Application Example 1)

[0666] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0667] Conventional navigation systems make it difficult for users to obtain destination information and tourist information while operating the vehicle, and often fail to provide appropriate guidance, especially during long-distance drives or when visiting unfamiliar places. Furthermore, current car navigation systems rely heavily on manual operation, which poses a problem in terms of the efficiency of information provision to users in autonomous vehicles.

[0668] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0669] In this invention, the server includes means for the user to request a destination by voice or text, means for analyzing the requested data and extracting destination information, means for transmitting the extracted destination information to the server, means for searching for details of the relevant destination based on the destination information received by the server and generating candidate information, means for transmitting the generated candidate information to a terminal, means for the terminal to present the received candidate information to the user by voice or display and start navigation, means for receiving requests for sightseeing information from the user during navigation and providing further information, means installed in an autonomous vehicle, which processes voice input using a terminal as a user interface, and means for communicating with a cloud server to generate destination information and sightseeing information and provide it to the user. As a result, the user can efficiently and safely receive destination guidance and sightseeing information while riding in an autonomous vehicle.

[0670] "Voice input" is a technology that acquires voice data and converts it into a digital format.

[0671] "Text input" refers to the technology of entering text information using input devices such as keyboards and touch panels.

[0672] "Destination information" refers to detailed data about a location specified by the user, including location information and related tourist information.

[0673] A "server" is a computer system located in a cloud computing environment that provides information using data processing and generative artificial intelligence.

[0674] A "terminal" is a device that receives input from a user and displays information, and is often installed as part of an in-vehicle system.

[0675] "Navigation" is a function that calculates and guides the user along the optimal route to reach their destination.

[0676] "Tourist information" refers to detailed data about tourist attractions, history, and facilities in a destination or its surrounding area.

[0677] An "autonomous vehicle" is a vehicle equipped with autonomous driving technology that can drive autonomously without user intervention.

[0678] "Generative artificial intelligence" refers to AI technology that analyzes and generates data, automatically creating and providing destination information and tourist guides.

[0679] A "cloud server" is a server used remotely via the internet, and it is a computing environment for processing and storing large amounts of data.

[0680] The system of this invention has a configuration for realizing destination guidance and sightseeing information within an autonomous vehicle. A specific embodiment thereof is described below.

[0681] System Configuration

[0682] User interface (terminal)

[0683] The terminal is installed inside the autonomous vehicle and functions as a user input interface. The terminal includes the following hardware:

[0684] Microphone: A device used to acquire user voice input.

[0685] Display: A monitor used to display destination information and tourist information.

[0686] Built-in computer: A processing unit for speech recognition and data analysis.

[0687] server

[0688] The server is located in a cloud environment and operates using the following software and APIs.

[0689] Speech recognition module: Software that converts speech input into text.

[0690] Generative artificial intelligence (generative AI): AI models that perform data analysis and information generation.

[0691] Database: A storage system for storing destination information and tourist guides.

[0692] Communication module (such as the requests library): A module for sending and receiving data between a terminal and a cloud server.

[0693] Processing flow

[0694] 1. Voice input

[0695] The user makes a voice request inside the autonomous vehicle, saying "I want to go to XX." The microphone captures the voice, and the built-in computer converts it into text using a speech recognition module.

[0696] 2. Data transmission

[0697] The converted text data is sent to a cloud server. The server uses generative artificial intelligence to extract destination information from the submitted keywords and generate relevant tourist information.

[0698] 3. Information presentation

[0699] The generated destination information and tourist information are sent to the terminal. The terminal displays the information on its screen and provides guidance to the user via voice or text.

[0700] 4. Start Navigation

[0701] The autonomous vehicle calculates a route based on the destination information it has acquired and then begins autonomous driving. During navigation, the user can request additional sightseeing information, and the corresponding information will be generated and provided.

[0702] Specific example

[0703] For example, if a user requests by voice, "I want to go to Shibuya Station," the microphone collects the voice, and the built-in computing system converts the voice into text. The converted keyword "Shibuya Station" is sent to a cloud server, which generates location information for "Shibuya Station" and nearby tourist information. This information is returned to the device, displayed on the screen, and navigation begins.

[0704] Example of a prompt

[0705] User input: "I would like directions to Shibuya Station and some sightseeing information."

[0706] System response:

[0707] "Calculating the route to Shibuya Station. Performing voice recognition to continue..."

[0708] "This is the route to Shibuya Station. Near Shibuya Station, you'll find tourist attractions such as the Shibuya Scramble Crossing, Center Gai, and Shibuya Hikarie."

[0709] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0710] Step 1:

[0711] The user enters a voice request into the microphone inside the autonomous vehicle, saying "I want to go to XX." The terminal acquires the voice data, and the built-in computer's voice recognition module converts it into text data. The input is voice data, and the output is text data.

[0712] Step 2:

[0713] The terminal's built-in computer analyzes the text data output from the speech recognition module and extracts keywords related to the destination. At this stage, the input is the text data obtained in step 1, and the output is the extracted keywords.

[0714] Step 3:

[0715] The terminal sends the extracted keywords to the cloud server. A communication module (e.g., the requests library) is used to send and receive data with the cloud server. The input for this step is the extracted keywords, and the output is the result of sending a request to the server.

[0716] Step 4:

[0717] The server analyzes the received keywords using generative artificial intelligence (generative AI model) and retrieves detailed information about the corresponding destination from a database or external API. After retrieving the information, the server generates optimal candidate information using the generative AI. Here, the input is the received keywords, and the output is the generated destination candidate information.

[0718] Step 5:

[0719] The server sends the generated destination candidate information to the terminal. Data is exchanged in real time using a communication module. The input is the generated candidate information, and the output is the result of the information transmission to the terminal.

[0720] Step 6:

[0721] The terminal presents the received candidate information to the user via voice or display. The user confirms the destination by checking the display or voice guidance. The input is candidate information from the server, and the output is destination candidate information presented to the user.

[0722] Step 7:

[0723] Once the user confirms the destination, the terminal sets the destination information in the autonomous vehicle's navigation system and starts navigation. The input is the destination information confirmed by the user, and the output is the autonomous vehicle's route calculation and driving instructions.

[0724] Step 8:

[0725] During navigation, if the user requests additional sightseeing information, the device sends another request to the server. The server uses a generative AI model to generate sightseeing information and sends it to the device. The input for this step is the user's sightseeing information request, and the output is the generated sightseeing information.

[0726] Step 9:

[0727] The terminal provides the user with received tourist information. It presents information to the user using a display and voice guidance, providing guidance in real time. The input is tourist information from the server, and the output is detailed tourist information presented to the user.

[0728] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0729] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[0730] System Overview

[0731] User interface (terminal)

[0732] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0733] server

[0734] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0735] Emotional Engine

[0736] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[0737] Program Processing Overview

[0738] Request reception and analysis

[0739] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination. In addition, an emotion engine analyzes the user's emotions from the voice or text.

[0740] Requests to the server and searches

[0741] The device sends request data to the server, which includes keywords extracted by the device and analyzed sentiment data. Based on the received keywords, the server retrieves destination details from its database or external APIs and generates candidate information that takes sentiment data into consideration. The generated information is then sent back to the device.

[0742] Presentation and navigation to the user

[0743] The terminal presents the user with received candidate information via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and begins navigation. During navigation, the terminal guides the user while updating route information in real time. It also adjusts the tone and content of the guidance according to the user's emotional state.

[0744] Providing tourist information

[0745] During navigation, if the user requests additional information about a tourist spot, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist spot and sends it to the device, taking sentiment data into consideration. The device then provides this information to the user and adjusts the tone and content of the guidance according to their emotions.

[0746] Specific example

[0747] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[0748] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[0749] The following describes the processing flow.

[0750] Step 1:

[0751] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[0752] Step 2:

[0753] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[0754] Step 3:

[0755] The device analyzes the converted text, "I want to go to Tokyo Tower." A natural language processing module extracts the keyword "Tokyo Tower" from the text data. Additionally, an emotion engine analyzes the user's emotions (e.g., excitement, fatigue) from the audio and text data.

[0756] Step 4:

[0757] The terminal generates request data containing extracted keywords and analyzed sentiment data, and sends it to the server using a communication module.

[0758] Step 5:

[0759] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data (location information, nearby facilities, parking information, etc.) from databases and external APIs. It also generates guidance information that takes emotional data into consideration.

[0760] Step 6:

[0761] Based on the data acquired by the server, it generates destination information and candidate information corresponding to emotions, and sends this to the terminal.

[0762] Step 7:

[0763] The terminal outputs destination information and suggested destinations received from the server to the display device. The speech synthesis module guides the user, saying, "Your destination is Tokyo Tower. Let's go." The tone and content of the guidance are adjusted according to the user's emotional data.

[0764] Step 8:

[0765] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[0766] Step 9:

[0767] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[0768] Step 10:

[0769] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection." During navigation, the emotion engine continuously analyzes the user's emotions and adjusts the tone and content of the guidance accordingly.

[0770] Step 11:

[0771] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[0772] Step 12:

[0773] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[0774] Step 13:

[0775] The server analyzes the additional requests it receives, and a generative artificial intelligence retrieves tourist information (for example, the history of Tokyo Tower) from databases and external APIs. Based on the user's emotions analyzed by the emotion engine, the tone and content of the generated guide information are adjusted.

[0776] Step 14:

[0777] The server generates tourist information and sends corresponding guidance information to the terminal. The terminal then uses a speech synthesis module to provide this information to the user.

[0778] Step 15:

[0779] The device continues navigation, guiding the user to their destination while updating route information in real time. The tone and content of the guidance continue to be adjusted based on the user's emotions.

[0780] (Example 2)

[0781] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0782] Traditional car navigation systems simply provide route guidance to a destination and lack the ability to provide personalized guidance information tailored to the user's emotional state. This resulted in low user satisfaction and made them unsuitable for sightseeing trips intended for relaxation or excitement. Furthermore, if the requested sightseeing information did not align with the user's emotional state, the entire travel experience could become unpleasant.

[0783] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0784] In this invention, the server includes means for analyzing requested data and extracting destination information, means for transmitting the extracted destination information and analyzed user emotion data to the server, and means for searching for details of the relevant destination based on the received destination information and emotion data, and generating candidate information corresponding to the emotion using generative artificial intelligence. This enables personalized sightseeing guidance and navigation that is tailored to the user's emotional state.

[0785] A "user" refers to a person who uses a system to request destinations or tourist information using voice or text.

[0786] A "terminal" refers to a device installed as part of a vehicle system that receives voice and text input from the user and processes the requested data.

[0787] A "speech recognition module" refers to software or hardware that converts a user's voice input into text data.

[0788] An "emotion engine" refers to a technology that analyzes a user's emotional state from their voice or text input and outputs the results of that analysis.

[0789] A "server" refers to a device located in a central computing environment that processes requested data and generates detailed destination information and tourist guides.

[0790] A "generative AI model" refers to artificial intelligence technology that generates personalized tourist information and navigation information based on the user's emotional state.

[0791] A "database" refers to a data management system that stores detailed information and tourist information about a destination, and allows users to search for and retrieve that information as needed.

[0792] An "external API" refers to an application program interface used to retrieve information by interacting with other services or databases.

[0793] A "navigation system" refers to a device and software that provides route guidance to a destination and updates route information in real time.

[0794] "Personalized guidance" refers to guidance information whose content and tone are adjusted based on the user's current emotional state.

[0795] This invention relates to a system that allows users to request destinations via voice or text, provides car navigation and tourist information based on those requests, analyzes the user's emotions, and provides personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[0796] User interface (terminal)

[0797] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and has built-in speech recognition and text conversion capabilities. It can also exchange data with a server in real time via a communication module. Specifically, the terminal uses the Google Speech-to-Text API to convert speech to text and further uses an emotion engine to analyze the user's emotions from the speech and text.

[0798] server

[0799] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. Specifically, destination data is obtained using the Google Maps API and MongoDB, and tourist information is generated using generative AI models such as OpenAI's GPT-4.

[0800] Emotional Engine

[0801] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[0802] Specific example

[0803] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[0804] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[0805] Prompt example

[0806] The prompt message when a user requests to go to Tokyo Tower.

[0807] "The user requested to visit Tokyo Tower, and the emotion engine detected excitement. Please generate an energetic sightseeing guide suitable for the excited user."

[0808] The prompt text when a user requests, "Tell me about the history of Tokyo Tower."

[0809] "The user requested information about the history of Tokyo Tower, and the emotion engine detected fatigue. Please generate a gentle, user-friendly tourist guide."

[0810] In this way, the present invention allows users to receive more satisfying, personalized navigation and tourist information.

[0811] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0812] Step 1:

[0813] The user requests the destination via voice or text.

[0814] Input: User voice or text request (e.g., "I want to go to Tokyo Tower").

[0815] Specific action: The user speaks to the in-car system, saying, "I want to go to Tokyo Tower."

[0816] Output: The device receives the audio data.

[0817] Step 2:

[0818] The device converts voice input into text using a speech recognition module.

[0819] Input: User voice data received by the terminal.

[0820] Specific operation: The speech recognition module installed in the device (e.g., Google Speech-to-Text API) converts speech to text.

[0821] Output: Text data (e.g., "I want to go to Tokyo Tower").

[0822] Step 3:

[0823] The terminal analyzes the request content and extracts destination information.

[0824] Input: Converted text data.

[0825] Specific operation: The terminal uses a text analysis module to extract keywords related to the destination (e.g., "Tokyo Tower").

[0826] Output: Extracted keywords (e.g., "Tokyo Tower").

[0827] Step 4:

[0828] The device sends the converted text and audio data to the emotion engine for sentiment analysis.

[0829] Input: Converted text data and audio data.

[0830] Specific operation: The emotion engine analyzes the voice tone and text content to determine the user's emotional state (e.g., "excited").

[0831] Output: Sentiment data (e.g., "excited").

[0832] Step 5:

[0833] The device sends the extracted keywords and analyzed sentiment data to the server.

[0834] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[0835] Specific operation: The terminal structures keyword and sentiment data, generates data packets, and sends them to the server.

[0836] Output: The server receives the request data.

[0837] Step 6:

[0838] The server retrieves destination information from external APIs and databases based on the keywords it receives, and uses a generative AI model to generate candidate information that reflects emotions.

[0839] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[0840] Specific operation: The server uses the Google Maps API or similar to retrieve detailed information about "Tokyo Tower," and then uses a generative AI model like OpenAI's GPT-4 to generate a guide text appropriate for the emotion "excitement."

[0841] Output: Generated candidate information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[0842] Step 7:

[0843] The server sends the candidate information it generates to the terminal.

[0844] Input: Generated candidate information (Example: Guide text "Tokyo Tower is a very popular tourist spot, and the night view is especially wonderful!").

[0845] Specific operation: The server structures the candidate information as a data packet and sends it to the terminal.

[0846] Output: The terminal receives candidate information.

[0847] Step 8:

[0848] The terminal presents the received candidate information to the user via voice or display.

[0849] Input: Received suggested information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[0850] Specific operation: The device uses a speech synthesis module and display function to present instructions to the user.

[0851] Output: The user reviews the suggested information.

[0852] Step 9:

[0853] The user confirms the destination, the device sets the destination in the car's navigation system, and starts navigation.

[0854] Input: User confirmation instructions (e.g., "Set Tokyo Tower as destination").

[0855] Specific operation: The device receives user instructions, sets the destination "Tokyo Tower" in the car navigation system, and starts navigation.

[0856] Output: The car navigation system begins route guidance.

[0857] Step 10:

[0858] During navigation, the device guides the user while updating route information in real time.

[0859] Input: Real-time data from the car navigation system.

[0860] Specific operation: The terminal receives real-time data from the navigation system and guides the user while updating route information.

[0861] Output: Guidance based on the latest route information.

[0862] Step 11:

[0863] During navigation, the user requests additional information about tourist attractions.

[0864] Input: User voice or text request (e.g., "Tell me about the history of Tokyo Tower").

[0865] Specific operation: The user requests additional information from the in-vehicle system via voice or text.

[0866] Output: The device receives voice or text data.

[0867] Step 12:

[0868] The device receives the request, converts it to text using a speech recognition module, and performs sentiment analysis using an emotion engine.

[0869] Input: User voice data (e.g., "Tell me about the history of Tokyo Tower").

[0870] Specific operation: The device converts the voice data into text using a speech recognition module, and then performs sentiment analysis using an emotion engine (e.g., "fatigue").

[0871] Output: Text data and sentiment data (e.g., "fatigue").

[0872] Step 13:

[0873] The terminal sends the request data to the server again along with the analysis results.

[0874] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[0875] Specific operation: The device structures text data and sentiment data and sends them to the server as data packets.

[0876] Output: The server receives the request data.

[0877] Step 14:

[0878] The server uses an AI model to generate tourist information that responds to emotions.

[0879] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[0880] Specific operation: The server retrieves information from an internal database or external API, and a generative AI model generates a gentle, user-friendly guidance message suitable for "fatigue."

[0881] Output: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[0882] Step 15:

[0883] The server generates a message and sends it to the terminal.

[0884] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[0885] Specific operation: The server structures the generated guidance message into a data packet and sends it to the terminal.

[0886] Output: The terminal receives the generated guidance message.

[0887] Step 16:

[0888] The terminal provides the user with the received notification message, and adjusts the tone and content of the message according to the user's emotions.

[0889] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[0890] Specific operation: The device uses a speech synthesis module and display function to present guidance text to the user in an appropriate tone.

[0891] Output: Users understand and become interested in tourist information.

[0892] (Application Example 2)

[0893] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0894] Conventional car navigation and tourist information systems only provide destinations and tourist information based on user requests, and are unable to provide personalized guidance that takes into account the user's emotional state. This results in a lack of improved user experience. Furthermore, methods for integrating and providing multiple pieces of information are insufficient, making it difficult to provide appropriate information based on the user's emotions.

[0895] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing emotions from the user's voice or text input, means for adjusting the tone and content of the guidance based on the emotion analysis results, and means for generating tourist information using generative artificial intelligence on the server and providing the user with personalized guidance based on emotion data. This makes it possible to provide personalized guidance according to the user's emotional state.

[0896] A "user" refers to a person who uses the system to request destinations or receive tourist information.

[0897] "Voice or text" refers to a means by which a user inputs instructions to the system, and includes both voice input and text input.

[0898] A "request" refers to an instruction or request that a user makes to a system.

[0899] "Analyzing data" refers to the process of understanding user requests and extracting their content.

[0900] "Destination information" refers to detailed data related to the destination specified by the user through the system.

[0901] A "server" is a central computer system that uses data processing and generation artificial intelligence to retrieve and generate information based on user requests.

[0902] "Extraction" refers to the act of selecting destination information from request data.

[0903] "Searching" refers to the process by which a server retrieves relevant information from external databases or APIs based on destination information it has received.

[0904] "Candidate information" refers to multiple suggestions or options generated by the server to respond to a user's request.

[0905] A "terminal" refers to a device that provides a user interface and allows users to access a system.

[0906] "Navigation" refers to a guide function that directs the user to a specified destination.

[0907] "Tourism information" refers to additional information about tourist attractions and facilities related to the destination.

[0908] "Emotion" refers to the emotional state analyzed from the user's voice and text.

[0909] "Emotion analysis" refers to the process of identifying a user's emotional state from their input data.

[0910] "Generative artificial intelligence" refers to artificial intelligence technology that generates tourist information based on user requests and sentiment data.

[0911] "Personalization" refers to adjusting the content of information and guidance provided according to each user's emotional state and preferences.

[0912] "The tone and content of the guidance" refers to the tone and specific content of the information provided to the user, and is adjusted based on their emotional state.

[0913] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[0914] User interface (terminal)

[0915] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0916] server

[0917] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0918] Emotional Engine

[0919] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[0920] Specific examples

[0921] Processing voice requests

[0922] If a user requests to go to Senso-ji Temple by voice while inside the vehicle, the voice is input through the device's microphone and converted into text by a voice recognition module. Next, this text data is sent to an emotion engine for emotion analysis. Let's say "joy" is detected at this point.

[0923] Data processing on the server

[0924] The analyzed sentiment data and text data are integrated and sent to the server. The server identifies the destination "Senso-ji Temple" from the received text data and generates personalized tourist information based on the sentiment data. For example, it might generate information such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple."

[0925] Provide feedback

[0926] The generated personalized guidance information is sent back to the device, which then presents the guidance to the user via voice or display. At this time, the tone and content of the guidance are adjusted based on emotional data.

[0927] As described above, this system allows users to receive not only navigation to their destination, but also optimal sightseeing information tailored to their mood at the time.

[0928] Example of a prompt

[0929] "Please analyze the following voice request and generate appropriate, emotion-sensitive tourist information:

[0930] Voice request: "I want to go to Senso-ji Temple."

[0931] Emotion analysis result: "Joy"

[0932] In this way, by providing guidance tailored to the user's emotions, it becomes possible to offer a more personalized navigation experience.

[0933] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0934] Step 1:

[0935] The user requests a destination by voice or text. When the user speaks into the in-car microphone, saying "I want to go to Senso-ji Temple," the voice is picked up by the device's microphone.

[0936] Input: Voice Request

[0937] Output: Audio data

[0938] Step 2:

[0939] The device's speech recognition module converts the audio data into text data. The speech recognition module (for example, the Vosk library) analyzes the audio data and generates the text "I want to go to Senso-ji Temple."

[0940] Input: Audio data

[0941] Output: Text data

[0942] Step 3:

[0943] The emotion engine analyzes the user's emotions based on text data. The emotion engine (for example, EmotionRecognizer) analyzes text data and voice tone information to identify the emotion of "joy" in this case.

[0944] Input: Text data, voice tone

[0945] Output: Sentiment data

[0946] Step 4:

[0947] The terminal sends generated text data and sentiment data to the server. A communication module is used to send request data to the server.

[0948] Input: Text data, sentiment data

[0949] Output: Request data

[0950] Step 5:

[0951] The server processes the request data and generates tourist information based on destination and sentiment data. The server's data processing module analyzes the request data, retrieves tourist information for "Senso-ji Temple" from an external API, and then a generative artificial intelligence adjusts the personalized guidance.

[0952] Input: Request data

[0953] Output: Personalized tourist information

[0954] Step 6:

[0955] The server sends the generated personalized tourist information to the terminal. The server sends the tourist information to the terminal via a communication module.

[0956] Input: Personalized tourist information

[0957] Output: Transmitted data

[0958] Step 7:

[0959] The terminal presents tourist information received by the user via voice or display. A speech synthesis engine (e.g., pyttsx3) generates a guidance message, such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple," which is presented aloud.

[0960] Input: Data to send

[0961] Output: Audio or display

[0962] Step 8:

[0963] The system checks the information received by the user and inputs instructions to start navigation. When the user gives a voice command such as "Start navigation," that command is entered into the terminal.

[0964] Input: User instructions

[0965] Output: Navigation start instruction

[0966] Step 9:

[0967] The terminal activates the navigation system and provides directions to the destination. The navigation module calculates the optimal route based on the current location and destination, and provides real-time guidance.

[0968] Input: Navigation start command

[0969] Output: Navigation Guide

[0970] In this way, the system analyzes the user's emotions based on their request, provides optimal sightseeing information, and initiates navigation.

[0971] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0972] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0973] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0974] [Third Embodiment]

[0975] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0976] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0977] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0978] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0979] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0980] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0981] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0982] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0983] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0984] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0985] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0986] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0987] This invention is a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as the user interface and a server that includes data processing and generative artificial intelligence.

[0988] System Overview

[0989] User interface (terminal)

[0990] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[0991] server

[0992] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[0993] Program Processing Overview

[0994] Request reception and analysis

[0995] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination.

[0996] Requests to the server and searches

[0997] The terminal sends the extracted keywords to the server. Based on the received keywords, the server retrieves detailed destination information from its database or external APIs. Based on the retrieved information, it generates the most suitable candidate information and sends it back to the terminal.

[0998] Presentation and navigation to the user

[0999] The terminal presents the received candidate information to the user via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and starts navigation. During navigation, the terminal guides the user while updating route information in real time.

[1000] Providing tourist information

[1001] During navigation, if the user requests additional information about a tourist destination, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist destination and sends it to the device. The device then provides this information to the user.

[1002] Specific example

[1003] For example, if a user requests by voice, "I want to go to Tokyo Tower," the device converts the voice to text and extracts the keyword "Tokyo Tower." This keyword is sent to the server, which retrieves location information for "Tokyo Tower," information on the nearest parking lot, etc., and sends it to the device. The device then presents the retrieved information to the user and starts navigation. During navigation, if the user requests, "Tell me about the history of Tokyo Tower," the server generates detailed information about its history and provides it to the user through the device. In this way, users can receive destination guidance and sightseeing information with simple voice commands.

[1004] The following describes the processing flow.

[1005] Step 1:

[1006] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[1007] Step 2:

[1008] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[1009] Step 3:

[1010] The terminal analyzes the converted text "I want to go to Tokyo Tower." The natural language processing module extracts the keyword "Tokyo Tower" from the text data.

[1011] Step 4:

[1012] The terminal generates request data containing the extracted keywords and sends it to the server using the communication module.

[1013] Step 5:

[1014] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data from databases and external APIs.

[1015] Step 6:

[1016] Based on the data acquired by the server, destination information (e.g., location information for Tokyo Tower, information on the nearest parking lot, etc.) is generated. The generated information is optimized and sent to the terminal.

[1017] Step 7:

[1018] The terminal outputs destination information received from the server to the display device. The speech synthesis module then informs the user, "Your destination is Tokyo Tower. Let's depart."

[1019] Step 8:

[1020] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[1021] Step 9:

[1022] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[1023] Step 10:

[1024] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection."

[1025] Step 11:

[1026] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[1027] Step 12:

[1028] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[1029] Step 13:

[1030] The server analyzes the additional requests it receives. Generative artificial intelligence retrieves information about the "history of Tokyo Tower" from the database.

[1031] Step 14:

[1032] The server generates tourist information and sends it to the terminal. The terminal then uses a speech synthesis module to guide the user through this information.

[1033] Step 15:

[1034] The device continues navigation, guiding the user to their destination while updating route information in real time.

[1035] (Example 1)

[1036] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1037] Conventional car navigation systems have struggled to efficiently retrieve and provide detailed destination information and tourist information when users request a destination via voice or text. Furthermore, while there is a demand for real-time tourist information in addition to route guidance, the means to achieve this have been insufficient.

[1038] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1039] In this invention, the server includes means for converting speech to text using a speech recognition module, means for extracting keywords using a natural language processing algorithm, and means for retrieving and generating detailed information using a server located in a cloud computing environment. This makes it possible to obtain and provide detailed destination and tourist information in real time when a user requests a destination by voice or text.

[1040] A "user" refers to a person who uses the system to obtain directions to their destination or tourist information.

[1041] A "speech recognition module" refers to a software or hardware component that converts voice input from a user into text data.

[1042] A "natural language processing algorithm" refers to a programmatic method for extracting and analyzing specific keywords from text data.

[1043] A "cloud computing environment" refers to an infrastructure that provides computing resources and data storage via the internet.

[1044] A "server" refers to a computer system located in a central cloud computing environment that processes data and provides information based on user requests.

[1045] An "external API" refers to an interface used to access external databases and services and retrieve information.

[1046] A "terminal" refers to a device that, as part of an in-vehicle system, receives voice or text input from the user and communicates with a server.

[1047] A "navigation system" refers to a system that provides route guidance to a user-specified destination and can update route information in real time.

[1048] "Generative artificial intelligence" refers to advanced AI models used to generate detailed information about destinations and tourist attractions.

[1049] A "prompt" refers to a text-based question or instruction that is input to a generative artificial intelligence system.

[1050] This invention relates to a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as a user interface and a server including data processing and generative artificial intelligence.

[1051] User interface (terminal)

[1052] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1053] server

[1054] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating suggested information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. As a specific example of the software, a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API) is used for speech recognition, NLP algorithms are used for natural language processing, and external APIs (e.g., Google Places API) are used for searching for detailed information.

[1055] Specifically, if a user requests by voice, "I want to go to Tokyo Tower," the terminal converts this voice into text and extracts the keyword "Tokyo Tower." The extracted keyword is sent to the server as a data packet. The server retrieves information based on the received keyword through a database or external API, organizes the retrieved information, and sends it to the terminal. The terminal displays the retrieved information on its screen or presents it to the user using a voice output module. After the user confirms the destination, the terminal configures the car navigation system and starts navigation. The navigation system guides the user to the destination while updating route information in real time.

[1056] If the user requests again during navigation, "Tell me about the history of Tokyo Tower," the device will send another request to the server. The server will use generative artificial intelligence (for example, an OpenAI model) to generate detailed historical information and send it to the device. The device will then provide this information to the user in voice or text.

[1057] Example of a prompt

[1058] Examples of prompt statements include the following:

[1059] "Please describe a program that provides detailed information about the location specified by the user as their destination and then performs navigation."

[1060] "Please explain, with specific examples, the processing steps of a system that provides tourist information based on user requests."

[1061] As described above, the present invention efficiently processes voice or text requests from users and provides detailed navigation and tourist information in real time.

[1062] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1063] Step 1:

[1064] A user submits a voice request.

[1065] The user sends a voice request from inside the car saying, "I want to go to Tokyo Tower."

[1066] Input: User voice.

[1067] Output: Acquisition of audio data by the terminal.

[1068] Step 2:

[1069] The device converts speech to text.

[1070] The device receives audio using its built-in microphone and converts the audio into text data using a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API).

[1071] Input: Audio data.

[1072] Output: Text data.

[1073] Specific operation: The speech recognition module analyzes the speech pattern and outputs it as a string.

[1074] Step 3:

[1075] The device extracts keywords from the text.

[1076] The device uses a natural language processing algorithm (e.g., NLP) to extract keywords like "Tokyo Tower" from text data.

[1077] Input: Text data.

[1078] Output: Keywords.

[1079] Specific operation: An NLP algorithm analyzes the context and structure of the text to identify the destination.

[1080] Step 4:

[1081] The device sends the keyword to the server.

[1082] The terminal compiles the extracted keywords into a data packet and sends it to the server via the communication module.

[1083] Input: Keyword.

[1084] Output: Data packet containing the keyword.

[1085] Specific operation: The communication module uploads the keyword to the server.

[1086] Step 5:

[1087] The server searches for information based on keywords.

[1088] Based on the received keywords, the server searches for detailed destination information using its database and external APIs (e.g., Google Places API).

[1089] Input: Data packet containing keywords.

[1090] Output: Detailed destination information (location, parking information, etc.).

[1091] Specific operations: Database search and data retrieval from external APIs.

[1092] Step 6:

[1093] The server sends the search results to the device.

[1094] The server organizes the search results, combines location information and nearest parking information into a single data packet, and sends it to the terminal.

[1095] Input: Destination details.

[1096] Output: Data packet containing search results.

[1097] Specific operation: Configures a data packet and sends it to the terminal.

[1098] Step 7:

[1099] The device displays search results to the user.

[1100] The terminal displays the received information on its screen or notifies the user via voice through an audio output module.

[1101] Input: Data packet containing search results.

[1102] Output: Information presented to the user.

[1103] Specific actions: Display the results on the screen or communicate them via voice.

[1104] Step 8:

[1105] The user confirms the destination.

[1106] The user reviews the information provided and selects "Tokyo Tower" as their destination.

[1107] Input: The search results presented.

[1108] Output: Action to confirm destination.

[1109] Specific actions: The user operates the display or audio confirmation button.

[1110] Step 9:

[1111] The device starts navigation.

[1112] The device sets Tokyo Tower as the destination in the car's navigation system and begins navigation. The navigation system (e.g., Garmin or TomTom) provides the latest route information in real time and guides the user to the destination.

[1113] Input: Action to confirm destination.

[1114] Output: Navigation started.

[1115] Specific operation: The navigation system calculates and displays the route from the current location to the destination.

[1116] Step 10:

[1117] The user requests additional information.

[1118] During navigation, the user requests additional information via voice, saying, "Tell me about the history of Tokyo Tower."

[1119] Input: Voice request for additional information.

[1120] Output: Acquisition of audio data by the terminal.

[1121] Specific operation: The device receives audio via the microphone and recognizes the audio pattern.

[1122] Step 11:

[1123] The device sends the request to the server.

[1124] The terminal converts the user's request into text data and sends it back to the server.

[1125] Input: Audio data of the request for additional information.

[1126] Output: Data packet containing the request.

[1127] Specific operation: The speech recognition module converts speech to text and sends it to the server.

[1128] Step 12:

[1129] The server generates additional information.

[1130] The server uses generative artificial intelligence (e.g., an AI model) to generate detailed information about the "history of Tokyo Tower" as requested by the user.

[1131] Input: Data packet containing the request.

[1132] Output: Generated detailed information.

[1133] Specific operation: The artificial intelligence model generates information and outputs it as data.

[1134] Step 13:

[1135] The device provides additional information to the user.

[1136] The terminal provides the user with the received detailed information via an audio output module or display.

[1137] Input: Data packet containing detailed information.

[1138] Output: Information presented to the user.

[1139] Specific actions: Display information on the screen or provide audio notifications.

[1140] (Application Example 1)

[1141] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1142] Conventional navigation systems make it difficult for users to obtain destination information and tourist information while operating the vehicle, and often fail to provide appropriate guidance, especially during long-distance drives or when visiting unfamiliar places. Furthermore, current car navigation systems rely heavily on manual operation, which poses a problem in terms of the efficiency of information provision to users in autonomous vehicles.

[1143] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1144] In this invention, the server includes means for the user to request a destination by voice or text, means for analyzing the requested data and extracting destination information, means for transmitting the extracted destination information to the server, means for searching for details of the relevant destination based on the destination information received by the server and generating candidate information, means for transmitting the generated candidate information to a terminal, means for the terminal to present the received candidate information to the user by voice or display and start navigation, means for receiving requests for sightseeing information from the user during navigation and providing further information, means installed in an autonomous vehicle, which processes voice input using a terminal as a user interface, and means for communicating with a cloud server to generate destination information and sightseeing information and provide it to the user. As a result, the user can efficiently and safely receive destination guidance and sightseeing information while riding in an autonomous vehicle.

[1145] "Voice input" is a technology that acquires voice data and converts it into a digital format.

[1146] "Text input" refers to the technology of entering text information using input devices such as keyboards and touch panels.

[1147] "Destination information" refers to detailed data about a location specified by the user, including location information and related tourist information.

[1148] A "server" is a computer system located in a cloud computing environment that provides information using data processing and generative artificial intelligence.

[1149] A "terminal" is a device that receives input from a user and displays information, and is often installed as part of an in-vehicle system.

[1150] "Navigation" is a function that calculates and guides the user along the optimal route to reach their destination.

[1151] "Tourist information" refers to detailed data about tourist attractions, history, and facilities in a destination or its surrounding area.

[1152] An "autonomous vehicle" is a vehicle equipped with autonomous driving technology that can drive autonomously without user intervention.

[1153] "Generative artificial intelligence" refers to AI technology that analyzes and generates data, automatically creating and providing destination information and tourist guides.

[1154] A "cloud server" is a server used remotely via the internet, and it is a computing environment for processing and storing large amounts of data.

[1155] The system of this invention has a configuration for realizing destination guidance and sightseeing information within an autonomous vehicle. A specific embodiment thereof is described below.

[1156] System Configuration

[1157] User interface (terminal)

[1158] The terminal is installed inside the autonomous vehicle and functions as a user input interface. The terminal includes the following hardware:

[1159] Microphone: A device used to acquire user voice input.

[1160] Display: A monitor used to display destination information and tourist information.

[1161] Built-in computer: A processing unit for speech recognition and data analysis.

[1162] server

[1163] The server is located in a cloud environment and operates using the following software and APIs.

[1164] Speech recognition module: Software that converts speech input into text.

[1165] Generative artificial intelligence (generative AI): AI models that perform data analysis and information generation.

[1166] Database: A storage system for storing destination information and tourist guides.

[1167] Communication module (such as the requests library): A module for sending and receiving data between a terminal and a cloud server.

[1168] Processing flow

[1169] 1. Voice input

[1170] The user makes a voice request inside the autonomous vehicle, saying "I want to go to XX." The microphone captures the voice, and the built-in computer converts it into text using a speech recognition module.

[1171] 2. Data transmission

[1172] The converted text data is sent to a cloud server. The server uses generative artificial intelligence to extract destination information from the submitted keywords and generate relevant tourist information.

[1173] 3. Information presentation

[1174] The generated destination information and tourist information are sent to the terminal. The terminal displays the information on its screen and provides guidance to the user via voice or text.

[1175] 4. Start Navigation

[1176] The autonomous vehicle calculates a route based on the destination information it has acquired and then begins autonomous driving. During navigation, the user can request additional sightseeing information, and the corresponding information will be generated and provided.

[1177] Specific example

[1178] For example, if a user requests by voice, "I want to go to Shibuya Station," the microphone collects the voice, and the built-in computing system converts the voice into text. The converted keyword "Shibuya Station" is sent to a cloud server, which generates location information for "Shibuya Station" and nearby tourist information. This information is returned to the device, displayed on the screen, and navigation begins.

[1179] Example of a prompt

[1180] User input: "I would like directions to Shibuya Station and some sightseeing information."

[1181] System response:

[1182] "Calculating the route to Shibuya Station. Performing voice recognition to continue..."

[1183] "This is the route to Shibuya Station. Near Shibuya Station, you'll find tourist attractions such as the Shibuya Scramble Crossing, Center Gai, and Shibuya Hikarie."

[1184] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1185] Step 1:

[1186] The user enters a voice request into the microphone inside the autonomous vehicle, saying "I want to go to XX." The terminal acquires the voice data, and the built-in computer's voice recognition module converts it into text data. The input is voice data, and the output is text data.

[1187] Step 2:

[1188] The terminal's built-in computer analyzes the text data output from the speech recognition module and extracts keywords related to the destination. At this stage, the input is the text data obtained in step 1, and the output is the extracted keywords.

[1189] Step 3:

[1190] The terminal sends the extracted keywords to the cloud server. A communication module (e.g., the requests library) is used to send and receive data with the cloud server. The input for this step is the extracted keywords, and the output is the result of sending a request to the server.

[1191] Step 4:

[1192] The server analyzes the received keywords using generative artificial intelligence (generative AI model) and retrieves detailed information about the corresponding destination from a database or external API. After retrieving the information, the server generates optimal candidate information using the generative AI. Here, the input is the received keywords, and the output is the generated destination candidate information.

[1193] Step 5:

[1194] The server sends the generated destination candidate information to the terminal. Data is exchanged in real time using a communication module. The input is the generated candidate information, and the output is the result of the information transmission to the terminal.

[1195] Step 6:

[1196] The terminal presents the received candidate information to the user via voice or display. The user confirms the destination by checking the display or voice guidance. The input is candidate information from the server, and the output is destination candidate information presented to the user.

[1197] Step 7:

[1198] Once the user confirms the destination, the terminal sets the destination information in the autonomous vehicle's navigation system and starts navigation. The input is the destination information confirmed by the user, and the output is the autonomous vehicle's route calculation and driving instructions.

[1199] Step 8:

[1200] During navigation, if the user requests additional sightseeing information, the device sends another request to the server. The server uses a generative AI model to generate sightseeing information and sends it to the device. The input for this step is the user's sightseeing information request, and the output is the generated sightseeing information.

[1201] Step 9:

[1202] The terminal provides the user with received tourist information. It presents information to the user using a display and voice guidance, providing guidance in real time. The input is tourist information from the server, and the output is detailed tourist information presented to the user.

[1203] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1204] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[1205] System Overview

[1206] User interface (terminal)

[1207] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1208] server

[1209] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[1210] Emotional Engine

[1211] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[1212] Program Processing Overview

[1213] Request reception and analysis

[1214] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination. In addition, an emotion engine analyzes the user's emotions from the voice or text.

[1215] Requests to the server and searches

[1216] The device sends request data to the server, which includes keywords extracted by the device and analyzed sentiment data. Based on the received keywords, the server retrieves destination details from its database or external APIs and generates candidate information that takes sentiment data into consideration. The generated information is then sent back to the device.

[1217] Presentation and navigation to the user

[1218] The terminal presents the user with received candidate information via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and begins navigation. During navigation, the terminal guides the user while updating route information in real time. It also adjusts the tone and content of the guidance according to the user's emotional state.

[1219] Providing tourist information

[1220] During navigation, if the user requests additional information about a tourist spot, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist spot and sends it to the device, taking sentiment data into consideration. The device then provides this information to the user and adjusts the tone and content of the guidance according to their emotions.

[1221] Specific example

[1222] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[1223] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[1224] The following describes the processing flow.

[1225] Step 1:

[1226] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[1227] Step 2:

[1228] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[1229] Step 3:

[1230] The device analyzes the converted text, "I want to go to Tokyo Tower." A natural language processing module extracts the keyword "Tokyo Tower" from the text data. Additionally, an emotion engine analyzes the user's emotions (e.g., excitement, fatigue) from the audio and text data.

[1231] Step 4:

[1232] The terminal generates request data containing extracted keywords and analyzed sentiment data, and sends it to the server using a communication module.

[1233] Step 5:

[1234] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data (location information, nearby facilities, parking information, etc.) from databases and external APIs. It also generates guidance information that takes emotional data into consideration.

[1235] Step 6:

[1236] Based on the data acquired by the server, it generates destination information and candidate information corresponding to emotions, and sends this to the terminal.

[1237] Step 7:

[1238] The terminal outputs destination information and suggested destinations received from the server to the display device. The speech synthesis module guides the user, saying, "Your destination is Tokyo Tower. Let's go." The tone and content of the guidance are adjusted according to the user's emotional data.

[1239] Step 8:

[1240] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[1241] Step 9:

[1242] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[1243] Step 10:

[1244] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection." During navigation, the emotion engine continuously analyzes the user's emotions and adjusts the tone and content of the guidance accordingly.

[1245] Step 11:

[1246] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[1247] Step 12:

[1248] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[1249] Step 13:

[1250] The server analyzes the additional requests it receives, and a generative artificial intelligence retrieves tourist information (for example, the history of Tokyo Tower) from databases and external APIs. Based on the user's emotions analyzed by the emotion engine, the tone and content of the generated guide information are adjusted.

[1251] Step 14:

[1252] The server generates tourist information and sends corresponding guidance information to the terminal. The terminal then uses a speech synthesis module to provide this information to the user.

[1253] Step 15:

[1254] The device continues navigation, guiding the user to their destination while updating route information in real time. The tone and content of the guidance continue to be adjusted based on the user's emotions.

[1255] (Example 2)

[1256] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1257] Traditional car navigation systems simply provide route guidance to a destination and lack the ability to provide personalized guidance information tailored to the user's emotional state. This resulted in low user satisfaction and made them unsuitable for sightseeing trips intended for relaxation or excitement. Furthermore, if the requested sightseeing information did not align with the user's emotional state, the entire travel experience could become unpleasant.

[1258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1259] In this invention, the server includes means for analyzing requested data and extracting destination information, means for transmitting the extracted destination information and analyzed user emotion data to the server, and means for searching for details of the relevant destination based on the received destination information and emotion data, and generating candidate information corresponding to the emotion using generative artificial intelligence. This enables personalized sightseeing guidance and navigation that is tailored to the user's emotional state.

[1260] A "user" refers to a person who uses a system to request destinations or tourist information using voice or text.

[1261] A "terminal" refers to a device installed as part of a vehicle system that receives voice and text input from the user and processes the requested data.

[1262] A "speech recognition module" refers to software or hardware that converts a user's voice input into text data.

[1263] An "emotion engine" refers to a technology that analyzes a user's emotional state from their voice or text input and outputs the results of that analysis.

[1264] A "server" refers to a device located in a central computing environment that processes requested data and generates detailed destination information and tourist guides.

[1265] A "generative AI model" refers to artificial intelligence technology that generates personalized tourist information and navigation information based on the user's emotional state.

[1266] A "database" refers to a data management system that stores detailed information and tourist information about a destination, and allows users to search for and retrieve that information as needed.

[1267] An "external API" refers to an application program interface used to retrieve information by interacting with other services or databases.

[1268] A "navigation system" refers to a device and software that provides route guidance to a destination and updates route information in real time.

[1269] "Personalized guidance" refers to guidance information whose content and tone are adjusted based on the user's current emotional state.

[1270] This invention relates to a system that allows users to request destinations via voice or text, provides car navigation and tourist information based on those requests, analyzes the user's emotions, and provides personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[1271] User interface (terminal)

[1272] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and has built-in speech recognition and text conversion capabilities. It can also exchange data with a server in real time via a communication module. Specifically, the terminal uses the Google Speech-to-Text API to convert speech to text and further uses an emotion engine to analyze the user's emotions from the speech and text.

[1273] server

[1274] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. Specifically, destination data is obtained using the Google Maps API and MongoDB, and tourist information is generated using generative AI models such as OpenAI's GPT-4.

[1275] Emotional Engine

[1276] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[1277] Specific example

[1278] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[1279] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[1280] Prompt example

[1281] The prompt message when a user requests to go to Tokyo Tower.

[1282] "The user requested to visit Tokyo Tower, and the emotion engine detected excitement. Please generate an energetic sightseeing guide suitable for the excited user."

[1283] The prompt text when a user requests, "Tell me about the history of Tokyo Tower."

[1284] "The user requested information about the history of Tokyo Tower, and the emotion engine detected fatigue. Please generate a gentle, user-friendly tourist guide."

[1285] In this way, the present invention allows users to receive more satisfying, personalized navigation and tourist information.

[1286] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1287] Step 1:

[1288] The user requests the destination via voice or text.

[1289] Input: User voice or text request (e.g., "I want to go to Tokyo Tower").

[1290] Specific action: The user speaks to the in-car system, saying, "I want to go to Tokyo Tower."

[1291] Output: The device receives the audio data.

[1292] Step 2:

[1293] The device converts voice input into text using a speech recognition module.

[1294] Input: User voice data received by the terminal.

[1295] Specific operation: The speech recognition module installed in the device (e.g., Google Speech-to-Text API) converts speech to text.

[1296] Output: Text data (e.g., "I want to go to Tokyo Tower").

[1297] Step 3:

[1298] The terminal analyzes the request content and extracts destination information.

[1299] Input: Converted text data.

[1300] Specific operation: The terminal uses a text analysis module to extract keywords related to the destination (e.g., "Tokyo Tower").

[1301] Output: Extracted keywords (e.g., "Tokyo Tower").

[1302] Step 4:

[1303] The device sends the converted text and audio data to the emotion engine for sentiment analysis.

[1304] Input: Converted text data and audio data.

[1305] Specific operation: The emotion engine analyzes the voice tone and text content to determine the user's emotional state (e.g., "excited").

[1306] Output: Sentiment data (e.g., "excited").

[1307] Step 5:

[1308] The device sends the extracted keywords and analyzed sentiment data to the server.

[1309] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[1310] Specific operation: The terminal structures keyword and sentiment data, generates data packets, and sends them to the server.

[1311] Output: The server receives the request data.

[1312] Step 6:

[1313] The server retrieves destination information from external APIs and databases based on the keywords it receives, and uses a generative AI model to generate candidate information that reflects emotions.

[1314] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[1315] Specific operation: The server uses the Google Maps API or similar to retrieve detailed information about "Tokyo Tower," and then uses a generative AI model like OpenAI's GPT-4 to generate a guide text appropriate for the emotion "excitement."

[1316] Output: Generated candidate information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[1317] Step 7:

[1318] The server sends the candidate information it generates to the terminal.

[1319] Input: Generated candidate information (Example: Guide text "Tokyo Tower is a very popular tourist spot, and the night view is especially wonderful!").

[1320] Specific operation: The server structures the candidate information as a data packet and sends it to the terminal.

[1321] Output: The terminal receives candidate information.

[1322] Step 8:

[1323] The terminal presents the received candidate information to the user via voice or display.

[1324] Input: Received suggested information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[1325] Specific operation: The device uses a speech synthesis module and display function to present instructions to the user.

[1326] Output: The user reviews the suggested information.

[1327] Step 9:

[1328] The user confirms the destination, the device sets the destination in the car's navigation system, and starts navigation.

[1329] Input: User confirmation instructions (e.g., "Set Tokyo Tower as destination").

[1330] Specific operation: The device receives user instructions, sets the destination "Tokyo Tower" in the car navigation system, and starts navigation.

[1331] Output: The car navigation system begins route guidance.

[1332] Step 10:

[1333] During navigation, the device guides the user while updating route information in real time.

[1334] Input: Real-time data from the car navigation system.

[1335] Specific operation: The terminal receives real-time data from the navigation system and guides the user while updating route information.

[1336] Output: Guidance based on the latest route information.

[1337] Step 11:

[1338] During navigation, the user requests additional information about tourist attractions.

[1339] Input: User voice or text request (e.g., "Tell me about the history of Tokyo Tower").

[1340] Specific operation: The user requests additional information from the in-vehicle system via voice or text.

[1341] Output: The device receives voice or text data.

[1342] Step 12:

[1343] The device receives the request, converts it to text using a speech recognition module, and performs sentiment analysis using an emotion engine.

[1344] Input: User voice data (e.g., "Tell me about the history of Tokyo Tower").

[1345] Specific operation: The device converts the voice data into text using a speech recognition module, and then performs sentiment analysis using an emotion engine (e.g., "fatigue").

[1346] Output: Text data and sentiment data (e.g., "fatigue").

[1347] Step 13:

[1348] The terminal sends the request data to the server again along with the analysis results.

[1349] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[1350] Specific operation: The device structures text data and sentiment data and sends them to the server as data packets.

[1351] Output: The server receives the request data.

[1352] Step 14:

[1353] The server uses an AI model to generate tourist information that responds to emotions.

[1354] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[1355] Specific operation: The server retrieves information from an internal database or external API, and a generative AI model generates a gentle, user-friendly guidance message suitable for "fatigue."

[1356] Output: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[1357] Step 15:

[1358] The server generates a message and sends it to the terminal.

[1359] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[1360] Specific operation: The server structures the generated guidance message into a data packet and sends it to the terminal.

[1361] Output: The terminal receives the generated guidance message.

[1362] Step 16:

[1363] The terminal provides the user with the received notification message, and adjusts the tone and content of the message according to the user's emotions.

[1364] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[1365] Specific operation: The device uses a speech synthesis module and display function to present guidance text to the user in an appropriate tone.

[1366] Output: Users understand and become interested in tourist information.

[1367] (Application Example 2)

[1368] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1369] Conventional car navigation and tourist information systems only provide destinations and tourist information based on user requests, and are unable to provide personalized guidance that takes into account the user's emotional state. This results in a lack of improved user experience. Furthermore, methods for integrating and providing multiple pieces of information are insufficient, making it difficult to provide appropriate information based on the user's emotions.

[1370] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing emotions from the user's voice or text input, means for adjusting the tone and content of the guidance based on the emotion analysis results, and means for generating tourist information using generative artificial intelligence on the server and providing the user with personalized guidance based on emotion data. This makes it possible to provide personalized guidance according to the user's emotional state.

[1371] A "user" refers to a person who uses the system to request destinations or receive tourist information.

[1372] "Voice or text" refers to a means by which a user inputs instructions to the system, and includes both voice input and text input.

[1373] A "request" refers to an instruction or request that a user makes to a system.

[1374] "Analyzing data" refers to the process of understanding user requests and extracting their content.

[1375] "Destination information" refers to detailed data related to the destination specified by the user through the system.

[1376] A "server" is a central computer system that uses data processing and generation artificial intelligence to retrieve and generate information based on user requests.

[1377] "Extraction" refers to the act of selecting destination information from request data.

[1378] "Searching" refers to the process by which a server retrieves relevant information from external databases or APIs based on destination information it has received.

[1379] "Candidate information" refers to multiple suggestions or options generated by the server to respond to a user's request.

[1380] A "terminal" refers to a device that provides a user interface and allows users to access a system.

[1381] "Navigation" refers to a guide function that directs the user to a specified destination.

[1382] "Tourism information" refers to additional information about tourist attractions and facilities related to the destination.

[1383] "Emotion" refers to the emotional state analyzed from the user's voice and text.

[1384] "Emotion analysis" refers to the process of identifying a user's emotional state from their input data.

[1385] "Generative artificial intelligence" refers to artificial intelligence technology that generates tourist information based on user requests and sentiment data.

[1386] "Personalization" refers to adjusting the content of information and guidance provided according to each user's emotional state and preferences.

[1387] "The tone and content of the guidance" refers to the tone and specific content of the information provided to the user, and is adjusted based on their emotional state.

[1388] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[1389] User interface (terminal)

[1390] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1391] server

[1392] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[1393] Emotional Engine

[1394] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[1395] Specific examples

[1396] Processing voice requests

[1397] If a user requests to go to Senso-ji Temple by voice while inside the vehicle, the voice is input through the device's microphone and converted into text by a voice recognition module. Next, this text data is sent to an emotion engine for emotion analysis. Let's say "joy" is detected at this point.

[1398] Data processing on the server

[1399] The analyzed sentiment data and text data are integrated and sent to the server. The server identifies the destination "Senso-ji Temple" from the received text data and generates personalized tourist information based on the sentiment data. For example, it might generate information such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple."

[1400] Provide feedback

[1401] The generated personalized guidance information is sent back to the device, which then presents the guidance to the user via voice or display. At this time, the tone and content of the guidance are adjusted based on emotional data.

[1402] As described above, this system allows users to receive not only navigation to their destination, but also optimal sightseeing information tailored to their mood at the time.

[1403] Example of a prompt

[1404] "Please analyze the following voice request and generate appropriate, emotion-sensitive tourist information:

[1405] Voice request: "I want to go to Senso-ji Temple."

[1406] Emotion analysis result: "Joy"

[1407] In this way, by providing guidance tailored to the user's emotions, it becomes possible to offer a more personalized navigation experience.

[1408] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1409] Step 1:

[1410] The user requests a destination by voice or text. When the user speaks into the in-car microphone, saying "I want to go to Senso-ji Temple," the voice is picked up by the device's microphone.

[1411] Input: Voice Request

[1412] Output: Audio data

[1413] Step 2:

[1414] The device's speech recognition module converts the audio data into text data. The speech recognition module (for example, the Vosk library) analyzes the audio data and generates the text "I want to go to Senso-ji Temple."

[1415] Input: Audio data

[1416] Output: Text data

[1417] Step 3:

[1418] The emotion engine analyzes the user's emotions based on text data. The emotion engine (for example, EmotionRecognizer) analyzes text data and voice tone information to identify the emotion of "joy" in this case.

[1419] Input: Text data, voice tone

[1420] Output: Sentiment data

[1421] Step 4:

[1422] The terminal sends generated text data and sentiment data to the server. A communication module is used to send request data to the server.

[1423] Input: Text data, sentiment data

[1424] Output: Request data

[1425] Step 5:

[1426] The server processes the request data and generates tourist information based on destination and sentiment data. The server's data processing module analyzes the request data, retrieves tourist information for "Senso-ji Temple" from an external API, and then a generative artificial intelligence adjusts the personalized guidance.

[1427] Input: Request data

[1428] Output: Personalized tourist information

[1429] Step 6:

[1430] The server sends the generated personalized tourist information to the terminal. The server sends the tourist information to the terminal via a communication module.

[1431] Input: Personalized tourist information

[1432] Output: Transmitted data

[1433] Step 7:

[1434] The terminal presents tourist information received by the user via voice or display. A speech synthesis engine (e.g., pyttsx3) generates a guidance message, such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple," which is presented aloud.

[1435] Input: Data to send

[1436] Output: Audio or display

[1437] Step 8:

[1438] The system checks the information received by the user and inputs instructions to start navigation. When the user gives a voice command such as "Start navigation," that command is entered into the terminal.

[1439] Input: User instructions

[1440] Output: Navigation start instruction

[1441] Step 9:

[1442] The terminal activates the navigation system and provides directions to the destination. The navigation module calculates the optimal route based on the current location and destination, and provides real-time guidance.

[1443] Input: Navigation start command

[1444] Output: Navigation Guide

[1445] In this way, the system analyzes the user's emotions based on their request, provides optimal sightseeing information, and initiates navigation.

[1446] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1447] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1448] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1449] [Fourth Embodiment]

[1450] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1451] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1452] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1453] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1454] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1456] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1457] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1458] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1459] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1460] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1461] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1462] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1463] This invention is a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as the user interface and a server that includes data processing and generative artificial intelligence.

[1464] System Overview

[1465] User interface (terminal)

[1466] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1467] server

[1468] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[1469] Program Processing Overview

[1470] Request reception and analysis

[1471] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination.

[1472] Requests to the server and searches

[1473] The terminal sends the extracted keywords to the server. Based on the received keywords, the server retrieves detailed destination information from its database or external APIs. Based on the retrieved information, it generates the most suitable candidate information and sends it back to the terminal.

[1474] Presentation and navigation to the user

[1475] The terminal presents the received candidate information to the user via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and starts navigation. During navigation, the terminal guides the user while updating route information in real time.

[1476] Providing tourist information

[1477] During navigation, if the user requests additional information about a tourist destination, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist destination and sends it to the device. The device then provides this information to the user.

[1478] Specific example

[1479] For example, if a user requests by voice, "I want to go to Tokyo Tower," the device converts the voice to text and extracts the keyword "Tokyo Tower." This keyword is sent to the server, which retrieves location information for "Tokyo Tower," information on the nearest parking lot, etc., and sends it to the device. The device then presents the retrieved information to the user and starts navigation. During navigation, if the user requests, "Tell me about the history of Tokyo Tower," the server generates detailed information about its history and provides it to the user through the device. In this way, users can receive destination guidance and sightseeing information with simple voice commands.

[1480] The following describes the processing flow.

[1481] Step 1:

[1482] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[1483] Step 2:

[1484] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[1485] Step 3:

[1486] The terminal analyzes the converted text "I want to go to Tokyo Tower." The natural language processing module extracts the keyword "Tokyo Tower" from the text data.

[1487] Step 4:

[1488] The terminal generates request data containing the extracted keywords and sends it to the server using the communication module.

[1489] Step 5:

[1490] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data from databases and external APIs.

[1491] Step 6:

[1492] Based on the data acquired by the server, destination information (e.g., location information for Tokyo Tower, information on the nearest parking lot, etc.) is generated. The generated information is optimized and sent to the terminal.

[1493] Step 7:

[1494] The terminal outputs destination information received from the server to the display device. The speech synthesis module then informs the user, "Your destination is Tokyo Tower. Let's depart."

[1495] Step 8:

[1496] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[1497] Step 9:

[1498] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[1499] Step 10:

[1500] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection."

[1501] Step 11:

[1502] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[1503] Step 12:

[1504] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[1505] Step 13:

[1506] The server analyzes the additional requests it receives. Generative artificial intelligence retrieves information about the "history of Tokyo Tower" from the database.

[1507] Step 14:

[1508] The server generates tourist information and sends it to the terminal. The terminal then uses a speech synthesis module to guide the user through this information.

[1509] Step 15:

[1510] The device continues navigation, guiding the user to their destination while updating route information in real time.

[1511] (Example 1)

[1512] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1513] Conventional car navigation systems have struggled to efficiently retrieve and provide detailed destination information and tourist information when users request a destination via voice or text. Furthermore, while there is a demand for real-time tourist information in addition to route guidance, the means to achieve this have been insufficient.

[1514] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1515] In this invention, the server includes means for converting speech to text using a speech recognition module, means for extracting keywords using a natural language processing algorithm, and means for retrieving and generating detailed information using a server located in a cloud computing environment. This makes it possible to obtain and provide detailed destination and tourist information in real time when a user requests a destination by voice or text.

[1516] A "user" refers to a person who uses the system to obtain directions to their destination or tourist information.

[1517] A "speech recognition module" refers to a software or hardware component that converts voice input from a user into text data.

[1518] A "natural language processing algorithm" refers to a programmatic method for extracting and analyzing specific keywords from text data.

[1519] A "cloud computing environment" refers to an infrastructure that provides computing resources and data storage via the internet.

[1520] A "server" refers to a computer system located in a central cloud computing environment that processes data and provides information based on user requests.

[1521] An "external API" refers to an interface used to access external databases and services and retrieve information.

[1522] A "terminal" refers to a device that, as part of an in-vehicle system, receives voice or text input from the user and communicates with a server.

[1523] A "navigation system" refers to a system that provides route guidance to a user-specified destination and can update route information in real time.

[1524] "Generative artificial intelligence" refers to advanced AI models used to generate detailed information about destinations and tourist attractions.

[1525] A "prompt" refers to a text-based question or instruction that is input to a generative artificial intelligence system.

[1526] This invention relates to a system in which a user requests a destination by voice or text, and the system provides car navigation and tourist information based on that request. The system consists of a terminal as a user interface and a server including data processing and generative artificial intelligence.

[1527] User interface (terminal)

[1528] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1529] server

[1530] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating suggested information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. As a specific example of the software, a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API) is used for speech recognition, NLP algorithms are used for natural language processing, and external APIs (e.g., Google Places API) are used for searching for detailed information.

[1531] Specifically, if a user requests by voice, "I want to go to Tokyo Tower," the terminal converts this voice into text and extracts the keyword "Tokyo Tower." The extracted keyword is sent to the server as a data packet. The server retrieves information based on the received keyword through a database or external API, organizes the retrieved information, and sends it to the terminal. The terminal displays the retrieved information on its screen or presents it to the user using a voice output module. After the user confirms the destination, the terminal configures the car navigation system and starts navigation. The navigation system guides the user to the destination while updating route information in real time.

[1532] If the user requests again during navigation, "Tell me about the history of Tokyo Tower," the device will send another request to the server. The server will use generative artificial intelligence (for example, an OpenAI model) to generate detailed historical information and send it to the device. The device will then provide this information to the user in voice or text.

[1533] Example of a prompt

[1534] Examples of prompt statements include the following:

[1535] "Please describe a program that provides detailed information about the location specified by the user as their destination and then performs navigation."

[1536] "Please explain, with specific examples, the processing steps of a system that provides tourist information based on user requests."

[1537] As described above, the present invention efficiently processes voice or text requests from users and provides detailed navigation and tourist information in real time.

[1538] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1539] Step 1:

[1540] A user submits a voice request.

[1541] The user sends a voice request from inside the car saying, "I want to go to Tokyo Tower."

[1542] Input: User voice.

[1543] Output: Acquisition of audio data by the terminal.

[1544] Step 2:

[1545] The device converts speech to text.

[1546] The device receives audio using its built-in microphone and converts the audio into text data using a high-performance speech recognition module (e.g., Google Cloud Speech-to-Text API).

[1547] Input: Audio data.

[1548] Output: Text data.

[1549] Specific operation: The speech recognition module analyzes the speech pattern and outputs it as a string.

[1550] Step 3:

[1551] The device extracts keywords from the text.

[1552] The device uses a natural language processing algorithm (e.g., NLP) to extract keywords like "Tokyo Tower" from text data.

[1553] Input: Text data.

[1554] Output: Keywords.

[1555] Specific operation: An NLP algorithm analyzes the context and structure of the text to identify the destination.

[1556] Step 4:

[1557] The device sends the keyword to the server.

[1558] The terminal compiles the extracted keywords into a data packet and sends it to the server via the communication module.

[1559] Input: Keyword.

[1560] Output: Data packet containing the keyword.

[1561] Specific operation: The communication module uploads the keyword to the server.

[1562] Step 5:

[1563] The server searches for information based on keywords.

[1564] Based on the received keywords, the server searches for detailed destination information using its database and external APIs (e.g., Google Places API).

[1565] Input: Data packet containing keywords.

[1566] Output: Detailed destination information (location, parking information, etc.).

[1567] Specific operations: Database search and data retrieval from external APIs.

[1568] Step 6:

[1569] The server sends the search results to the device.

[1570] The server organizes the search results, combines location information and nearest parking information into a single data packet, and sends it to the terminal.

[1571] Input: Destination details.

[1572] Output: Data packet containing search results.

[1573] Specific operation: Configures a data packet and sends it to the terminal.

[1574] Step 7:

[1575] The device displays search results to the user.

[1576] The terminal displays the received information on its screen or notifies the user via voice through an audio output module.

[1577] Input: Data packet containing search results.

[1578] Output: Information presented to the user.

[1579] Specific actions: Display the results on the screen or communicate them via voice.

[1580] Step 8:

[1581] The user confirms the destination.

[1582] The user reviews the information provided and selects "Tokyo Tower" as their destination.

[1583] Input: The search results presented.

[1584] Output: Action to confirm destination.

[1585] Specific actions: The user operates the display or audio confirmation button.

[1586] Step 9:

[1587] The device starts navigation.

[1588] The device sets Tokyo Tower as the destination in the car's navigation system and begins navigation. The navigation system (e.g., Garmin or TomTom) provides the latest route information in real time and guides the user to the destination.

[1589] Input: Action to confirm destination.

[1590] Output: Navigation started.

[1591] Specific operation: The navigation system calculates and displays the route from the current location to the destination.

[1592] Step 10:

[1593] The user requests additional information.

[1594] During navigation, the user requests additional information via voice, saying, "Tell me about the history of Tokyo Tower."

[1595] Input: Voice request for additional information.

[1596] Output: Acquisition of audio data by the terminal.

[1597] Specific operation: The device receives audio via the microphone and recognizes the audio pattern.

[1598] Step 11:

[1599] The device sends the request to the server.

[1600] The terminal converts the user's request into text data and sends it back to the server.

[1601] Input: Audio data of the request for additional information.

[1602] Output: Data packet containing the request.

[1603] Specific operation: The speech recognition module converts speech to text and sends it to the server.

[1604] Step 12:

[1605] The server generates additional information.

[1606] The server uses generative artificial intelligence (e.g., an AI model) to generate detailed information about the "history of Tokyo Tower" as requested by the user.

[1607] Input: Data packet containing the request.

[1608] Output: Generated detailed information.

[1609] Specific operation: The artificial intelligence model generates information and outputs it as data.

[1610] Step 13:

[1611] The device provides additional information to the user.

[1612] The terminal provides the user with the received detailed information via an audio output module or display.

[1613] Input: Data packet containing detailed information.

[1614] Output: Information presented to the user.

[1615] Specific actions: Display information on the screen or provide audio notifications.

[1616] (Application Example 1)

[1617] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1618] Conventional navigation systems make it difficult for users to obtain destination information and tourist information while operating the vehicle, and often fail to provide appropriate guidance, especially during long-distance drives or when visiting unfamiliar places. Furthermore, current car navigation systems rely heavily on manual operation, which poses a problem in terms of the efficiency of information provision to users in autonomous vehicles.

[1619] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1620] In this invention, the server includes means for the user to request a destination by voice or text, means for analyzing the requested data and extracting destination information, means for transmitting the extracted destination information to the server, means for searching for details of the relevant destination based on the destination information received by the server and generating candidate information, means for transmitting the generated candidate information to a terminal, means for the terminal to present the received candidate information to the user by voice or display and start navigation, means for receiving requests for sightseeing information from the user during navigation and providing further information, means installed in an autonomous vehicle, which processes voice input using a terminal as a user interface, and means for communicating with a cloud server to generate destination information and sightseeing information and provide it to the user. As a result, the user can efficiently and safely receive destination guidance and sightseeing information while riding in an autonomous vehicle.

[1621] "Voice input" is a technology that acquires voice data and converts it into a digital format.

[1622] "Text input" refers to the technology of entering text information using input devices such as keyboards and touch panels.

[1623] "Destination information" refers to detailed data about a location specified by the user, including location information and related tourist information.

[1624] A "server" is a computer system located in a cloud computing environment that provides information using data processing and generative artificial intelligence.

[1625] A "terminal" is a device that receives input from a user and displays information, and is often installed as part of an in-vehicle system.

[1626] "Navigation" is a function that calculates and guides the user along the optimal route to reach their destination.

[1627] "Tourist information" refers to detailed data about tourist attractions, history, and facilities in a destination or its surrounding area.

[1628] An "autonomous vehicle" is a vehicle equipped with autonomous driving technology that can drive autonomously without user intervention.

[1629] "Generative artificial intelligence" refers to AI technology that analyzes and generates data, automatically creating and providing destination information and tourist guides.

[1630] A "cloud server" is a server used remotely via the internet, and it is a computing environment for processing and storing large amounts of data.

[1631] The system of this invention has a configuration for realizing destination guidance and sightseeing information within an autonomous vehicle. A specific embodiment thereof is described below.

[1632] System Configuration

[1633] User interface (terminal)

[1634] The terminal is installed inside the autonomous vehicle and functions as a user input interface. The terminal includes the following hardware:

[1635] Microphone: A device used to acquire user voice input.

[1636] Display: A monitor used to display destination information and tourist information.

[1637] Built-in computer: A processing unit for speech recognition and data analysis.

[1638] server

[1639] The server is located in a cloud environment and operates using the following software and APIs.

[1640] Speech recognition module: Software that converts speech input into text.

[1641] Generative artificial intelligence (generative AI): AI models that perform data analysis and information generation.

[1642] Database: A storage system for storing destination information and tourist guides.

[1643] Communication module (such as the requests library): A module for sending and receiving data between a terminal and a cloud server.

[1644] Processing flow

[1645] 1. Voice input

[1646] The user makes a voice request inside the autonomous vehicle, saying "I want to go to XX." The microphone captures the voice, and the built-in computer converts it into text using a speech recognition module.

[1647] 2. Data transmission

[1648] The converted text data is sent to a cloud server. The server uses generative artificial intelligence to extract destination information from the submitted keywords and generate relevant tourist information.

[1649] 3. Information presentation

[1650] The generated destination information and tourist information are sent to the terminal. The terminal displays the information on its screen and provides guidance to the user via voice or text.

[1651] 4. Start Navigation

[1652] The autonomous vehicle calculates a route based on the destination information it has acquired and then begins autonomous driving. During navigation, the user can request additional sightseeing information, and the corresponding information will be generated and provided.

[1653] Specific example

[1654] For example, if a user requests by voice, "I want to go to Shibuya Station," the microphone collects the voice, and the built-in computing system converts the voice into text. The converted keyword "Shibuya Station" is sent to a cloud server, which generates location information for "Shibuya Station" and nearby tourist information. This information is returned to the device, displayed on the screen, and navigation begins.

[1655] Example of a prompt

[1656] User input: "I would like directions to Shibuya Station and some sightseeing information."

[1657] System response:

[1658] "Calculating the route to Shibuya Station. Performing voice recognition to continue..."

[1659] "This is the route to Shibuya Station. Near Shibuya Station, you'll find tourist attractions such as the Shibuya Scramble Crossing, Center Gai, and Shibuya Hikarie."

[1660] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1661] Step 1:

[1662] The user enters a voice request into the microphone inside the autonomous vehicle, saying "I want to go to XX." The terminal acquires the voice data, and the built-in computer's voice recognition module converts it into text data. The input is voice data, and the output is text data.

[1663] Step 2:

[1664] The terminal's built-in computer analyzes the text data output from the speech recognition module and extracts keywords related to the destination. At this stage, the input is the text data obtained in step 1, and the output is the extracted keywords.

[1665] Step 3:

[1666] The terminal sends the extracted keywords to the cloud server. A communication module (e.g., the requests library) is used to send and receive data with the cloud server. The input for this step is the extracted keywords, and the output is the result of sending a request to the server.

[1667] Step 4:

[1668] The server analyzes the received keywords using generative artificial intelligence (generative AI model) and retrieves detailed information about the corresponding destination from a database or external API. After retrieving the information, the server generates optimal candidate information using the generative AI. Here, the input is the received keywords, and the output is the generated destination candidate information.

[1669] Step 5:

[1670] The server sends the generated destination candidate information to the terminal. Data is exchanged in real time using a communication module. The input is the generated candidate information, and the output is the result of the information transmission to the terminal.

[1671] Step 6:

[1672] The terminal presents the received candidate information to the user via voice or display. The user confirms the destination by checking the display or voice guidance. The input is candidate information from the server, and the output is destination candidate information presented to the user.

[1673] Step 7:

[1674] Once the user confirms the destination, the terminal sets the destination information in the autonomous vehicle's navigation system and starts navigation. The input is the destination information confirmed by the user, and the output is the autonomous vehicle's route calculation and driving instructions.

[1675] Step 8:

[1676] During navigation, if the user requests additional sightseeing information, the device sends another request to the server. The server uses a generative AI model to generate sightseeing information and sends it to the device. The input for this step is the user's sightseeing information request, and the output is the generated sightseeing information.

[1677] Step 9:

[1678] The terminal provides the user with received tourist information. It presents information to the user using a display and voice guidance, providing guidance in real time. The input is tourist information from the server, and the output is detailed tourist information presented to the user.

[1679] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1680] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[1681] System Overview

[1682] User interface (terminal)

[1683] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1684] server

[1685] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[1686] Emotional Engine

[1687] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[1688] Program Processing Overview

[1689] Request reception and analysis

[1690] When a user requests a destination via voice or text, the device receives it. The device converts the voice input into text using a speech recognition module and parses it as text data. Next, it analyzes the request content and extracts keywords related to the destination. In addition, an emotion engine analyzes the user's emotions from the voice or text.

[1691] Requests to the server and searches

[1692] The device sends request data to the server, which includes keywords extracted by the device and analyzed sentiment data. Based on the received keywords, the server retrieves destination details from its database or external APIs and generates candidate information that takes sentiment data into consideration. The generated information is then sent back to the device.

[1693] Presentation and navigation to the user

[1694] The terminal presents the user with received candidate information via voice or display. Once the user confirms the destination, the terminal sets the destination in the car navigation system and begins navigation. During navigation, the terminal guides the user while updating route information in real time. It also adjusts the tone and content of the guidance according to the user's emotional state.

[1695] Providing tourist information

[1696] During navigation, if the user requests additional information about a tourist spot, the device sends another request to the server. The server uses generative artificial intelligence to generate detailed information about the tourist spot and sends it to the device, taking sentiment data into consideration. The device then provides this information to the user and adjusts the tone and content of the guidance according to their emotions.

[1697] Specific example

[1698] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[1699] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[1700] The following describes the processing flow.

[1701] Step 1:

[1702] The user requests a destination via voice or text. For example, they might make a voice request saying, "I want to go to Tokyo Tower."

[1703] Step 2:

[1704] The device receives the user's voice input via the microphone. The speech recognition module converts the voice data into text.

[1705] Step 3:

[1706] The device analyzes the converted text, "I want to go to Tokyo Tower." A natural language processing module extracts the keyword "Tokyo Tower" from the text data. Additionally, an emotion engine analyzes the user's emotions (e.g., excitement, fatigue) from the audio and text data.

[1707] Step 4:

[1708] The terminal generates request data containing extracted keywords and analyzed sentiment data, and sends it to the server using a communication module.

[1709] Step 5:

[1710] The server analyzes the received request. Generative artificial intelligence recognizes "Tokyo Tower" and retrieves related data (location information, nearby facilities, parking information, etc.) from databases and external APIs. It also generates guidance information that takes emotional data into consideration.

[1711] Step 6:

[1712] Based on the data acquired by the server, it generates destination information and candidate information corresponding to emotions, and sends this to the terminal.

[1713] Step 7:

[1714] The terminal outputs destination information and suggested destinations received from the server to the display device. The speech synthesis module guides the user, saying, "Your destination is Tokyo Tower. Let's go." The tone and content of the guidance are adjusted according to the user's emotional data.

[1715] Step 8:

[1716] The user reviews the instructions and approves by voice input such as "yes." The device receives this voice input and converts it to text.

[1717] Step 9:

[1718] The device confirms user approval and sets "Tokyo Tower" as the destination in the car navigation system. The navigation algorithm calculates the optimal route.

[1719] Step 10:

[1720] The device starts navigation. Voice guidance gives instructions to the user, such as "Turn left at the next intersection." During navigation, the emotion engine continuously analyzes the user's emotions and adjusts the tone and content of the guidance accordingly.

[1721] Step 11:

[1722] The user requests additional information during navigation, such as "Tell me about the history of Tokyo Tower." The device receives this request.

[1723] Step 12:

[1724] The terminal converts additional requests into text using a speech recognition module and sends them to the server using a communication module.

[1725] Step 13:

[1726] The server analyzes the additional requests it receives, and a generative artificial intelligence retrieves tourist information (for example, the history of Tokyo Tower) from databases and external APIs. Based on the user's emotions analyzed by the emotion engine, the tone and content of the generated guide information are adjusted.

[1727] Step 14:

[1728] The server generates tourist information and sends corresponding guidance information to the terminal. The terminal then uses a speech synthesis module to provide this information to the user.

[1729] Step 15:

[1730] The device continues navigation, guiding the user to their destination while updating route information in real time. The tone and content of the guidance continue to be adjusted based on the user's emotions.

[1731] (Example 2)

[1732] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1733] Traditional car navigation systems simply provide route guidance to a destination and lack the ability to provide personalized guidance information tailored to the user's emotional state. This resulted in low user satisfaction and made them unsuitable for sightseeing trips intended for relaxation or excitement. Furthermore, if the requested sightseeing information did not align with the user's emotional state, the entire travel experience could become unpleasant.

[1734] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1735] In this invention, the server includes means for analyzing requested data and extracting destination information, means for transmitting the extracted destination information and analyzed user emotion data to the server, and means for searching for details of the relevant destination based on the received destination information and emotion data, and generating candidate information corresponding to the emotion using generative artificial intelligence. This enables personalized sightseeing guidance and navigation that is tailored to the user's emotional state.

[1736] A "user" refers to a person who uses a system to request destinations or tourist information using voice or text.

[1737] A "terminal" refers to a device installed as part of a vehicle system that receives voice and text input from the user and processes the requested data.

[1738] A "speech recognition module" refers to software or hardware that converts a user's voice input into text data.

[1739] An "emotion engine" refers to a technology that analyzes a user's emotional state from their voice or text input and outputs the results of that analysis.

[1740] A "server" refers to a device located in a central computing environment that processes requested data and generates detailed destination information and tourist guides.

[1741] A "generative AI model" refers to artificial intelligence technology that generates personalized tourist information and navigation information based on the user's emotional state.

[1742] A "database" refers to a data management system that stores detailed information and tourist information about a destination, and allows users to search for and retrieve that information as needed.

[1743] An "external API" refers to an application program interface used to retrieve information by interacting with other services or databases.

[1744] A "navigation system" refers to a device and software that provides route guidance to a destination and updates route information in real time.

[1745] "Personalized guidance" refers to guidance information whose content and tone are adjusted based on the user's current emotional state.

[1746] This invention relates to a system that allows users to request destinations via voice or text, provides car navigation and tourist information based on those requests, analyzes the user's emotions, and provides personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[1747] User interface (terminal)

[1748] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and has built-in speech recognition and text conversion capabilities. It can also exchange data with a server in real time via a communication module. Specifically, the terminal uses the Google Speech-to-Text API to convert speech to text and further uses an emotion engine to analyze the user's emotions from the speech and text.

[1749] server

[1750] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases. Specifically, destination data is obtained using the Google Maps API and MongoDB, and tourist information is generated using generative AI models such as OpenAI's GPT-4.

[1751] Emotional Engine

[1752] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[1753] Specific example

[1754] For example, if a user requests by voice, "I want to go to Tokyo Tower," and the emotion engine detects "excitement" from that voice, the server will generate information about "Tokyo Tower" along with energetic guidance suitable for the excited user. The terminal will then present the information to the user and provide further exciting guidance, such as, "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!"

[1755] Furthermore, if the emotion engine detects "fatigue" when a user requests "Tell me about the history of Tokyo Tower," the server will generate a relaxed and gentle guided message. The device will then guide the user with a message such as, "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure." In this way, by responding to the user's emotions, it is possible to provide a more personalized experience.

[1756] Prompt example

[1757] The prompt message when a user requests to go to Tokyo Tower.

[1758] "The user requested to visit Tokyo Tower, and the emotion engine detected excitement. Please generate an energetic sightseeing guide suitable for the excited user."

[1759] The prompt text when a user requests, "Tell me about the history of Tokyo Tower."

[1760] "The user requested information about the history of Tokyo Tower, and the emotion engine detected fatigue. Please generate a gentle, user-friendly tourist guide."

[1761] In this way, the present invention allows users to receive more satisfying, personalized navigation and tourist information.

[1762] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1763] Step 1:

[1764] The user requests the destination via voice or text.

[1765] Input: User voice or text request (e.g., "I want to go to Tokyo Tower").

[1766] Specific action: The user speaks to the in-car system, saying, "I want to go to Tokyo Tower."

[1767] Output: The device receives the audio data.

[1768] Step 2:

[1769] The device converts voice input into text using a speech recognition module.

[1770] Input: User voice data received by the terminal.

[1771] Specific operation: The speech recognition module installed in the device (e.g., Google Speech-to-Text API) converts speech to text.

[1772] Output: Text data (e.g., "I want to go to Tokyo Tower").

[1773] Step 3:

[1774] The terminal analyzes the request content and extracts destination information.

[1775] Input: Converted text data.

[1776] Specific operation: The terminal uses a text analysis module to extract keywords related to the destination (e.g., "Tokyo Tower").

[1777] Output: Extracted keywords (e.g., "Tokyo Tower").

[1778] Step 4:

[1779] The device sends the converted text and audio data to the emotion engine for sentiment analysis.

[1780] Input: Converted text data and audio data.

[1781] Specific operation: The emotion engine analyzes the voice tone and text content to determine the user's emotional state (e.g., "excited").

[1782] Output: Sentiment data (e.g., "excited").

[1783] Step 5:

[1784] The device sends the extracted keywords and analyzed sentiment data to the server.

[1785] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[1786] Specific operation: The terminal structures keyword and sentiment data, generates data packets, and sends them to the server.

[1787] Output: The server receives the request data.

[1788] Step 6:

[1789] The server retrieves destination information from external APIs and databases based on the keywords it receives, and uses a generative AI model to generate candidate information that reflects emotions.

[1790] Input: Keywords (e.g., "Tokyo Tower") and sentiment data (e.g., "excited").

[1791] Specific operation: The server uses the Google Maps API or similar to retrieve detailed information about "Tokyo Tower," and then uses a generative AI model like OpenAI's GPT-4 to generate a guide text appropriate for the emotion "excitement."

[1792] Output: Generated candidate information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[1793] Step 7:

[1794] The server sends the candidate information it generates to the terminal.

[1795] Input: Generated candidate information (Example: Guide text "Tokyo Tower is a very popular tourist spot, and the night view is especially wonderful!").

[1796] Specific operation: The server structures the candidate information as a data packet and sends it to the terminal.

[1797] Output: The terminal receives candidate information.

[1798] Step 8:

[1799] The terminal presents the received candidate information to the user via voice or display.

[1800] Input: Received suggested information (e.g., "Tokyo Tower is a very popular tourist spot, and the night view is especially amazing!").

[1801] Specific operation: The device uses a speech synthesis module and display function to present instructions to the user.

[1802] Output: The user reviews the suggested information.

[1803] Step 9:

[1804] The user confirms the destination, the device sets the destination in the car's navigation system, and starts navigation.

[1805] Input: User confirmation instructions (e.g., "Set Tokyo Tower as destination").

[1806] Specific operation: The device receives user instructions, sets the destination "Tokyo Tower" in the car navigation system, and starts navigation.

[1807] Output: The car navigation system begins route guidance.

[1808] Step 10:

[1809] During navigation, the device guides the user while updating route information in real time.

[1810] Input: Real-time data from the car navigation system.

[1811] Specific operation: The terminal receives real-time data from the navigation system and guides the user while updating route information.

[1812] Output: Guidance based on the latest route information.

[1813] Step 11:

[1814] During navigation, the user requests additional information about tourist attractions.

[1815] Input: User voice or text request (e.g., "Tell me about the history of Tokyo Tower").

[1816] Specific operation: The user requests additional information from the in-vehicle system via voice or text.

[1817] Output: The device receives voice or text data.

[1818] Step 12:

[1819] The device receives the request, converts it to text using a speech recognition module, and performs sentiment analysis using an emotion engine.

[1820] Input: User voice data (e.g., "Tell me about the history of Tokyo Tower").

[1821] Specific operation: The device converts the voice data into text using a speech recognition module, and then performs sentiment analysis using an emotion engine (e.g., "fatigue").

[1822] Output: Text data and sentiment data (e.g., "fatigue").

[1823] Step 13:

[1824] The terminal sends the request data to the server again along with the analysis results.

[1825] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[1826] Specific operation: The device structures text data and sentiment data and sends them to the server as data packets.

[1827] Output: The server receives the request data.

[1828] Step 14:

[1829] The server uses an AI model to generate tourist information that responds to emotions.

[1830] Input: Text data (e.g., "Tell me about the history of Tokyo Tower") and sentiment data (e.g., "Fatigue").

[1831] Specific operation: The server retrieves information from an internal database or external API, and a generative AI model generates a gentle, user-friendly guidance message suitable for "fatigue."

[1832] Output: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[1833] Step 15:

[1834] The server generates a message and sends it to the terminal.

[1835] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[1836] Specific operation: The server structures the generated guidance message into a data packet and sends it to the terminal.

[1837] Output: The terminal receives the generated guidance message.

[1838] Step 16:

[1839] The terminal provides the user with the received notification message, and adjusts the tone and content of the message according to the user's emotions.

[1840] Input: Generated tourist information text (Example: "Tokyo Tower was completed in 1958 and has a long history. Please enjoy it at your leisure.").

[1841] Specific operation: The device uses a speech synthesis module and display function to present guidance text to the user in an appropriate tone.

[1842] Output: Users understand and become interested in tourist information.

[1843] (Application Example 2)

[1844] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1845] Conventional car navigation and tourist information systems only provide destinations and tourist information based on user requests, and are unable to provide personalized guidance that takes into account the user's emotional state. This results in a lack of improved user experience. Furthermore, methods for integrating and providing multiple pieces of information are insufficient, making it difficult to provide appropriate information based on the user's emotions.

[1846] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing emotions from the user's voice or text input, means for adjusting the tone and content of the guidance based on the emotion analysis results, and means for generating tourist information using generative artificial intelligence on the server and providing the user with personalized guidance based on emotion data. This makes it possible to provide personalized guidance according to the user's emotional state.

[1847] A "user" refers to a person who uses the system to request destinations or receive tourist information.

[1848] "Voice or text" refers to a means by which a user inputs instructions to the system, and includes both voice input and text input.

[1849] A "request" refers to an instruction or request that a user makes to a system.

[1850] "Analyzing data" refers to the process of understanding user requests and extracting their content.

[1851] "Destination information" refers to detailed data related to the destination specified by the user through the system.

[1852] A "server" is a central computer system that uses data processing and generation artificial intelligence to retrieve and generate information based on user requests.

[1853] "Extraction" refers to the act of selecting destination information from request data.

[1854] "Searching" refers to the process by which a server retrieves relevant information from external databases or APIs based on destination information it has received.

[1855] "Candidate information" refers to multiple suggestions or options generated by the server to respond to a user's request.

[1856] A "terminal" refers to a device that provides a user interface and allows users to access a system.

[1857] "Navigation" refers to a guide function that directs the user to a specified destination.

[1858] "Tourism information" refers to additional information about tourist attractions and facilities related to the destination.

[1859] "Emotion" refers to the emotional state analyzed from the user's voice and text.

[1860] "Emotion analysis" refers to the process of identifying a user's emotional state from their input data.

[1861] "Generative artificial intelligence" refers to artificial intelligence technology that generates tourist information based on user requests and sentiment data.

[1862] "Personalization" refers to adjusting the content of information and guidance provided according to each user's emotional state and preferences.

[1863] "The tone and content of the guidance" refers to the tone and specific content of the information provided to the user, and is adjusted based on their emotional state.

[1864] This invention is a system in which a user requests a destination by voice or text, and based on that request, provides car navigation and tourist information, while also analyzing the user's emotions and providing personalized guidance based on those emotions. The system consists of a terminal as a user interface, a server including data processing and generative artificial intelligence, and an emotion engine.

[1865] User interface (terminal)

[1866] The terminal is installed as part of the in-vehicle system and is responsible for receiving voice and text input from the user. The terminal includes a microphone and display, and incorporates voice recognition and text conversion functions. Furthermore, a communication module allows for real-time data exchange with a server.

[1867] server

[1868] The servers are located in a central cloud computing environment and process data based on user requests. Generative artificial intelligence is incorporated to perform tasks such as searching for destination information, generating candidate information, and providing tourist information. Detailed destination information and tourist information are obtained using external APIs and internal databases.

[1869] Emotional Engine

[1870] The emotion engine has the ability to analyze emotions from the user's voice or text input. Based on factors such as voice tone, word choice, and sentence structure, this engine identifies the user's emotional state and reflects the results in the operation of the entire system.

[1871] Specific examples

[1872] Processing voice requests

[1873] If a user requests to go to Senso-ji Temple by voice while inside the vehicle, the voice is input through the device's microphone and converted into text by a voice recognition module. Next, this text data is sent to an emotion engine for emotion analysis. Let's say "joy" is detected at this point.

[1874] Data processing on the server

[1875] The analyzed sentiment data and text data are integrated and sent to the server. The server identifies the destination "Senso-ji Temple" from the received text data and generates personalized tourist information based on the sentiment data. For example, it might generate information such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple."

[1876] Provide feedback

[1877] The generated personalized guidance information is sent back to the device, which then presents the guidance to the user via voice or display. At this time, the tone and content of the guidance are adjusted based on emotional data.

[1878] As described above, this system allows users to receive not only navigation to their destination, but also optimal sightseeing information tailored to their mood at the time.

[1879] Example of a prompt

[1880] "Please analyze the following voice request and generate appropriate, emotion-sensitive tourist information:

[1881] Voice request: "I want to go to Senso-ji Temple."

[1882] Emotion analysis result: "Joy"

[1883] In this way, by providing guidance tailored to the user's emotions, it becomes possible to offer a more personalized navigation experience.

[1884] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1885] Step 1:

[1886] The user requests a destination by voice or text. When the user speaks into the in-car microphone, saying "I want to go to Senso-ji Temple," the voice is picked up by the device's microphone.

[1887] Input: Voice Request

[1888] Output: Audio data

[1889] Step 2:

[1890] The device's speech recognition module converts the audio data into text data. The speech recognition module (for example, the Vosk library) analyzes the audio data and generates the text "I want to go to Senso-ji Temple."

[1891] Input: Audio data

[1892] Output: Text data

[1893] Step 3:

[1894] The emotion engine analyzes the user's emotions based on text data. The emotion engine (for example, EmotionRecognizer) analyzes text data and voice tone information to identify the emotion of "joy" in this case.

[1895] Input: Text data, voice tone

[1896] Output: Sentiment data

[1897] Step 4:

[1898] The terminal sends generated text data and sentiment data to the server. A communication module is used to send request data to the server.

[1899] Input: Text data, sentiment data

[1900] Output: Request data

[1901] Step 5:

[1902] The server processes the request data and generates tourist information based on destination and sentiment data. The server's data processing module analyzes the request data, retrieves tourist information for "Senso-ji Temple" from an external API, and then a generative artificial intelligence adjusts the personalized guidance.

[1903] Input: Request data

[1904] Output: Personalized tourist information

[1905] Step 6:

[1906] The server sends the generated personalized tourist information to the terminal. The server sends the tourist information to the terminal via a communication module.

[1907] Input: Personalized tourist information

[1908] Output: Transmitted data

[1909] Step 7:

[1910] The terminal presents tourist information received by the user via voice or display. A speech synthesis engine (e.g., pyttsx3) generates a guidance message, such as, "Senso-ji Temple is a tourist spot in Taito Ward, Tokyo, and there are many beautiful sights around the temple," which is presented aloud.

[1911] Input: Data to send

[1912] Output: Audio or display

[1913] Step 8:

[1914] The system checks the information received by the user and inputs instructions to start navigation. When the user gives a voice command such as "Start navigation," that command is entered into the terminal.

[1915] Input: User instructions

[1916] Output: Navigation start instruction

[1917] Step 9:

[1918] The terminal activates the navigation system and provides directions to the destination. The navigation module calculates the optimal route based on the current location and destination, and provides real-time guidance.

[1919] Input: Navigation start command

[1920] Output: Navigation Guide

[1921] In this way, the system analyzes the user's emotions based on their request, provides optimal sightseeing information, and initiates navigation.

[1922] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1923] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1924] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1925] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1926] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1927] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1928] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1929] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1930] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1931] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1932] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1933] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1934] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1935] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1936] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1937] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1938] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1939] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1940] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1941] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1942] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1943] The following is further disclosed regarding the embodiments described above.

[1944] (Claim 1)

[1945] A means for users to request a destination by voice or text,

[1946] A means of analyzing the requested data and extracting destination information,

[1947] A means for sending the extracted destination information to the server,

[1948] A means for searching for details of the relevant destination and generating candidate information based on destination information received by the server,

[1949] A means for sending the generated candidate information to the terminal,

[1950] A means for the terminal to present the received candidate information to the user via voice or display and initiate navigation,

[1951] During navigation, a means of receiving tourist information requests from users and providing further information,

[1952] A system that includes this.

[1953] (Claim 2)

[1954] The system according to claim 1, further comprising means for providing route guidance to a destination in cooperation with a car navigation system.

[1955] (Claim 3)

[1956] The system according to claim 1, further comprising means for generating tourist information using generative artificial intelligence on a server and providing it to a user.

[1957] "Example 1"

[1958] (Claim 1)

[1959] A means for users to request a destination by voice or text,

[1960] A means of analyzing the requested data and extracting destination information,

[1961] A means for sending the extracted destination information to the server,

[1962] A means for searching for details of the relevant destination and generating candidate information based on destination information received by the server,

[1963] A means for sending the generated candidate information to the terminal,

[1964] A means for the terminal to present the received candidate information to the user via voice or display device and to initiate navigation,

[1965] A means of converting speech to text using a speech recognition module,

[1966] A method for extracting keywords using a natural language processing algorithm,

[1967] A means for searching and generating detailed information using servers located in a cloud computing environment,

[1968] Means of obtaining detailed information via an external API,

[1969] A means by which the navigation system updates and provides routes in real time,

[1970] During navigation, the system receives requests for tourist information from users and uses generative artificial intelligence to provide further information.

[1971] A system that includes this.

[1972] (Claim 2)

[1973] The system according to claim 1, further comprising means for providing route guidance to a destination in cooperation with a car navigation system.

[1974] (Claim 3)

[1975] The system according to claim 1, further comprising means for generating tourist information using generative artificial intelligence on a server and providing it to a user.

[1976] "Application Example 1"

[1977] (Claim 1)

[1978] A means for users to request a destination by voice or text,

[1979] A means of analyzing the requested data and extracting destination information,

[1980] A means for sending the extracted destination information to the server,

[1981] A means for searching for details of the relevant destination and generating candidate information based on destination information received by the server,

[1982] A means for sending the generated candidate information to the terminal,

[1983] A means for the terminal to present the received candidate information to the user via voice or display and initiate navigation,

[1984] During navigation, a means of receiving tourist information requests from users and providing further information,

[1985] The system is installed in an autonomous vehicle, and the vehicle processes voice input using a terminal as a user interface,

[1986] A means of communicating with a cloud server to generate destination information and tourist guides and providing them to users,

[1987] A system that includes this.

[1988] (Claim 2)

[1989] The system according to claim 1, further comprising means for providing route guidance to a destination in cooperation with a car navigation system.

[1990] (Claim 3)

[1991] The system according to claim 1, further comprising means for generating tourist information using generative artificial intelligence on a server and providing it to a user.

[1992] "Example 2 of combining an emotion engine"

[1993] (Claim 1)

[1994] A means for users to request a destination by voice or text,

[1995] A means of analyzing the requested data and extracting destination information,

[1996] A means for transmitting extracted destination information and analyzed user sentiment data to a server,

[1997] A means for searching for details of a relevant destination based on destination information and sentiment data received by a server, and generating candidate information corresponding to the sentiment using generative artificial intelligence,

[1998] A means for sending the generated candidate information to the terminal,

[1999] A means for the terminal to present the received candidate information to the user via voice or display and initiate navigation,

[2000] A means of providing tourist information while adjusting the tone and content of the guidance according to the user's emotional state during navigation,

[2001] A system that includes this.

[2002] (Claim 2)

[2003] The system according to claim 1, further comprising means for providing route guidance to a destination in cooperation with a car navigation system and updating route information in real time.

[2004] (Claim 3)

[2005] The system according to claim 1, further comprising means for generating personalized tourist information based on the user's emotions using artificial intelligence on a server and providing it to the user.

[2006] "Application example 2 when combining with an emotional engine"

[2007] (Claim 1)

[2008] A means for users to request a destination by voice or text,

[2009] A means of analyzing the requested data and extracting destination information,

[2010] A means for sending the extracted destination information to the server,

[2011] A means for searching for details of the relevant destination and generating candidate information based on destination information received by the server,

[2012] A means for sending the generated candidate information to the terminal,

[2013] A means for the terminal to present the received candidate information to the user via voice or display and initiate navigation,

[2014] During navigation, a means of receiving tourist information requests from users and providing further information,

[2015] A means of analyzing emotions from a user's voice or text input,

[2016] Based on the results of the emotion analysis, a means to adjust the tone and content of the guidance,

[2017] A system that includes this.

[2018] (Claim 2)

[2019] The system according to claim 1, further comprising means for providing route guidance to a destination in cooperation with a car navigation system.

[2020] (Claim 3)

[2021] The system according to claim 1, further comprising means for generating tourist information using generative artificial intelligence on a server and providing personalized guidance to the user based on emotional data. [Explanation of Symbols]

[2022] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for users to request a destination by voice or text, A means of analyzing the requested data and extracting destination information, A means for sending the extracted destination information to the server, A means for searching for details of the relevant destination and generating candidate information based on destination information received by the server, A means for sending the generated candidate information to the terminal, A means for the terminal to present the received candidate information to the user via voice or display and initiate navigation, During navigation, a means of receiving tourist information requests from users and providing further information, A system that includes this.

2. The system according to claim 1, further comprising means for providing route guidance to a destination in cooperation with a car navigation system.

3. The system according to claim 1, further comprising means for generating tourist information using generative artificial intelligence on a server and providing it to a user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A