system
A system using a generative model and emotion recognition to create personalized travel plans and provide real-time, emotionally responsive guidance addresses the challenges of insufficient information and language barriers, enhancing the travel experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Modern travelers, especially tourists, face challenges with insufficient information, language barriers, and the inability to provide personalized travel experiences tailored to their interests and emotional states, leading to suboptimal travel quality.
A system utilizing a generative model to create personalized travel plans based on destination and user interests, combined with real-time speech synthesis and emotion recognition to provide dynamic guidance and interaction, enhancing the travel experience.
Enriches the travel experience by providing personalized, real-time, and emotionally responsive guidance, improving the quality of tourism by addressing individual preferences and emotional states.
Smart Images

Figure 2026069178000001_ABST
Abstract
Description
Technical Field
[0005] ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Modern travelers, especially tourists visiting tourist destinations, want to obtain sufficient information about the place and enjoy a satisfactory tourism experience. However, in reality, there are problems that the quality of travel cannot be improved due to lack of information and language barriers. In addition, it is difficult to provide information according to the individual interests and preferences of travelers, and it is difficult to provide assistance suitable for all travelers. In particular, the lack of information for tourists visiting Japan and the shortage of guides for individual travelers are regarded as problems.
Means for Solving the Problems
[0006] A "generative model" is a set of algorithms used to automatically generate optimal travel plans based on a traveler's destination information and interests.
[0007] "Speech synthesis means" refers to a technology that converts text information into speech using a character's voice and provides guidance to the user in real time.
[0008] A "user's device" is a device carried and used by travelers, and is used to provide travel information, allow interaction with characters, and obtain location information.
[0009] A "travel plan" refers to an itinerary that includes the places a traveler will visit, their mode of transportation, and the length of stay, and is a plan that is individually optimized according to the traveler's interests and requests.
[0010] "Real-time" refers to the immediate processing and provision of information within the current timeframe.
[0011] "Feedback" refers to information entered by travelers as questions or reactions to characters, and it serves as fundamental data for the system to adjust its responses and guidance content.
[0012] "Dynamic adjustment" is a process that appropriately modifies the information provided and the travel plan in response to changes in the traveler's current location and environment, in order to provide the most suitable information. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0014] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the language used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system that provides professional travel guidance using an AI character, generating personalized travel plans based on the traveler's destination and interests, and providing real-time voice guidance. The system mainly consists of a server, a user terminal, and an AI character, and functions as follows:
[0035] 1. Server-driven plan generation
[0036] The server uses a generative model to create customized travel plans based on destination and interest information received from the user. This generative model has learned from a large amount of tourist destination data and guide expertise, and makes suggestions that take into account the optimal order of visits and modes of transportation. The server also reflects the user's past travel history and preferences to personalize the plan.
[0037] 2. Speech synthesis and real-time guidance
[0038] Based on a travel plan generated on the server, talk segments about each tourist destination in the plan are created and converted into a character's voice using speech synthesis technology. The user terminal receives this voice data and provides real-time voice guidance during the trip. Speech synthesis technology enables natural-sounding guidance at visited locations, supporting a smooth traveler experience.
[0039] 3. User Interaction
[0040] Users can interact with an AI character through their device to get answers to questions that arise during their sightseeing and to request additional information. The server analyzes the user's questions and requests and returns appropriate responses in real time using a generative model. This allows the character to support travelers as if they were a real-life guide.
[0041] Specific example
[0042] For example, consider a user who is a traveler interested in Japanese culture and is visiting Kyoto. The user sends a request from their device saying, "I want to visit traditional temples in Kyoto." The server uses a generative model to generate a travel plan that includes representative locations such as Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. During the trip, the user's device provides voice guidance in a character's voice, such as, "Next is Kinkaku-ji. Enjoy the iconic architectural style of the Muromachi period." If the user then asks, "I want to know more about the history," the server provides detailed historical background information based on that request.
[0043] In this way, this system utilizes AI technology to provide travelers with high-value tourism experiences and aims to improve the quality of their travels.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user enters their travel destination and areas of interest into the device. The device then sends the entered information to the server.
[0047] Step 2:
[0048] The server receives information from the user and starts a generative model to generate a travel plan. The server consults a database to collect information about relevant tourist destinations.
[0049] Step 3:
[0050] The server uses collected tourist destination information and a generative model to create a travel plan optimized for the user. The plan includes the order of visits, modes of transportation, and time allocation.
[0051] Step 4:
[0052] Based on the travel plan created by the server, the content of the guided conversation is constructed. The server then converts the conversation content into a character's voice using speech synthesis technology.
[0053] Step 5:
[0054] The device receives audio data sent from the server. The device then uses this data to begin providing real-time audio guidance to the user.
[0055] Step 6:
[0056] The user interacts with the character while traveling. The device forwards user questions and requests to the server via voice or text.
[0057] Step 7:
[0058] The server receives interaction from the user and generates an appropriate response using a generative model. The server converts the response into audio data and sends it to the terminal.
[0059] Step 8:
[0060] The device relays responses sent from the server to the user. The device continuously provides updated plans and information in real time as needed.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] Modern travelers tend to demand personalized information and real-time guidance, but traditional guide systems are unable to adequately meet these demands. Specifically, it has been difficult to provide personalized itineraries based on travelers' preferences and past travel history, as well as effective support including immediate responses to questions and requests.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for using a speech synthesis device for a character to provide voice guidance based on the generated travel plan, and means for personalizing the travel plan based on the traveler's past visit history and preferences. This enables more personalized and real-time information and guidance for the traveler.
[0066] A "traveler" refers to an individual who visits a tourist destination or place of interest and seeks information and guidance regarding their itinerary.
[0067] "Destination information" refers to information about specific tourist destinations or cities that travelers wish to visit.
[0068] A "generative model" refers to methods and algorithms, including artificial intelligence technology, used to create travel plans, which make suggestions based on tourist destination data and guide expertise.
[0069] A "speech synthesis device" refers to a technology that converts text data into speech data, enabling information to be transmitted using a human-like voice.
[0070] "Personalization" refers to the process of adjusting services and information based on the individual traveler's preferences and past visit history.
[0071] "Real-time" refers to a state where information and services can be provided immediately, meaning that travelers can resolve their questions and requests on the spot.
[0072] This invention is a system for providing travelers with personalized travel plans and real-time voice guidance. The main components of the invention are a server, a user terminal, and a generative AI model.
[0073] The server analyzes destination information and interest prompts provided by travelers and generates personalized travel plans based on this analysis. Natural language processing technology is used for this analysis, and a generating AI model utilizes tourist data and guide expertise to create the optimal plan. For example, if a user sends the prompt "I want to visit traditional temples in Kyoto," the server will generate a plan including Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. In doing so, it also considers past travel history and preferences to provide the most fulfilling travel experience for the user.
[0074] The device provides travelers with audio data based on their travel plan, received from the server. Utilizing speech synthesis technology, an AI character delivers information about each tourist spot as if it were a real-life guide. For example, as a user is on their way to Kinkaku-ji Temple, the device will play an audio announcement saying, "Next stop is Kinkaku-ji Temple. Please enjoy this iconic architectural style from the Muromachi period."
[0075] Users can interact with an AI character via their device during their trip, getting their questions answered and seeking more in-depth information. The server analyzes the received requests in real time, uses a generative AI model to quickly generate appropriate responses, and provides them to the user. This process allows travelers to gain a more valuable experience than they could have planned on their own.
[0076] Thus, this invention uses AI technology to improve the quality of travel and ensure consistency and convenience in providing information to travelers.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The user uses the terminal to enter prompts about their desired destination or interests. Specifically, if the user enters "I want to visit traditional temples in Kyoto," a prompt will be generated. This is sent to the server and becomes input data for the next processing step.
[0080] Step 2:
[0081] The server parses the received prompt message. This parsing uses natural language processing to understand the user's intent and desired destinations. Based on the data analysis, the server utilizes a generative AI model to generate a personalized travel plan from tourist destination data and guide expertise. The generated travel plan becomes the output data used in the next step.
[0082] Step 3:
[0083] Based on the travel plan generated on the server, audio guides for each tourist destination are created. The server uses speech synthesis technology to convert explanations about the places visited in the plan into natural-sounding speech. This audio data is generated and becomes the output that is sent to the user's terminal.
[0084] Step 4:
[0085] The device plays audio data received from the server, providing real-time guidance to travelers. When a user visits Kinkaku-ji Temple, the device plays an audio announcement such as, "Next stop is Kinkaku-ji Temple. Please enjoy the iconic architectural style of the Muromachi period." This allows users to experience a realistic guided tour.
[0086] Step 5:
[0087] Users can request additional information or ask questions during their trip. These requests are sent to the server via their device. Specifically, the user might say something like, "I want to know more about the history of this temple," and that request becomes new input data.
[0088] Step 6:
[0089] The server analyzes user requests in real time and uses a generative AI model to create appropriate responses. Based on the analysis results, the newly created response is generated as audio data and sent to the terminal. This allows the user to instantly obtain the information they need.
[0090] (Application Example 1)
[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0092] Traditional travel information systems have struggled to provide personalized travel plans for individual travelers, and the quality of real-time guidance and interaction has been limited. Furthermore, the lack of means to offer virtual sightseeing experiences before actual travel meant that users were not adequately provided with engaging experiences during the travel planning stage. This leaves challenges remaining, such as providing flexible guidance tailored to travelers' needs and increasing travel motivation through simulated experiences before the trip.
[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0094] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for voice synthesis to have a character provide voice guidance based on the generated travel plan, and means for providing guidance in a virtual environment to the user via visual augmentation technology, and for presenting guidance related to each scene in audio and visual form. This enables the provision of personalized travel plans that meet the individual needs of travelers, as well as real-time and interactive sightseeing guidance in a virtual environment.
[0095] The term "traveler" refers to individuals or groups who visit a specific destination and travel for various purposes, such as tourism or business.
[0096] "Destination information" refers to information about the geographical location and related tourist resources that travelers intend to visit, and includes specific place names and tourist attractions.
[0097] A "generative model" refers to an algorithm or technology used to automatically create travel plans based on input information. It is a system that learns from large amounts of data and patterns to generate responses.
[0098] "Speech synthesis means" refers to technology that converts text data into speech signals, and includes devices and software for artificially generating and providing speech to users.
[0099] "Visual terminals" refer to devices that allow users to receive visual information, and include smart glasses and head-mounted displays.
[0100] "Feedback" refers to the opinions, evaluations, and questions that users provide to a system, and is used to improve the system and generate responses.
[0101] "Visual augmentation technology" refers to the technology of overlaying virtual information onto the real world, and is a means of improving the user's experience of the real world.
[0102] A "virtual environment" refers to a virtual space or situation created using computer technology, allowing users to experience a space that does not exist in reality.
[0103] "Interactive tourist information" refers to a form of information gathering where travelers obtain information through two-way communication with the system, and it is a dynamic guidance method that changes its response according to the user's input.
[0104] This invention is a system that provides personalized travel guidance for travelers, and mainly consists of a server, a terminal, and user interaction.
[0105] The server uses a generative model to generate travel plans based on the traveler's destination information and interests. The generative model learns from a large amount of tourist destination data and guide expertise, and provides plans that consider the optimal order of visits and means of transportation. Based on the generated travel plan, the server uses speech synthesis to create talk about each tourist destination in the plan, converts it into the voice of an AI character, and sends it to the terminal.
[0106] The device plays back received audio data and provides travelers with real-time tourist information. It also enables travelers to experience visual information within the virtual environment through visual augmentation technology. Smart glasses and head-mounted displays are used in this process.
[0107] Users can interact with AI characters through these devices. Through this interaction, users can request feedback or additional information, which is then sent from the device to the server. The server analyzes the received feedback and reuses the generative model to generate an appropriate response.
[0108] For example, if a user requests to "look for summer clothes," the server creates a virtual tour plan of stores that carry the latest summer clothing. Based on this, the device will guide the user by saying, "This is a store showcasing the latest summer clothing collection."
[0109] An example of a prompt in a generative AI model is: "The user wants to buy summer clothes. Introduce stores that feature the latest trends and generate an audio guide."
[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0111] Step 1:
[0112] The user uses a terminal to input travel destination information and interests. This information is sent to a server. The server analyzes the data and uses a generative AI model to create a customized travel plan. The user's past travel history and preferences are also considered to determine the optimal order of visits and modes of transportation. The output is an optimized travel plan.
[0113] Step 2:
[0114] The server creates guided tours for each tourist spot based on the generated travel plan. Using speech synthesis technology, this talk is converted into the voice of an AI character. The input is the text data of the generated travel plan, and the output is an audio file. This audio data is then sent to the terminal.
[0115] Step 3:
[0116] The terminal plays audio data received from the server, providing real-time guidance to travelers. Simultaneously, users experience information within the virtual environment visually using visual augmentation technology. Specifically, this involves displaying information using smart glasses or head-mounted displays.
[0117] Step 4:
[0118] The user inputs questions and requests for additional information during the tour via voice. The input is sent from the terminal to the server. The server analyzes this information and uses a generated AI model to produce appropriate responses to the user's questions and requests. The output is voice data for the AI character to provide answers.
[0119] Step 5:
[0120] The terminal plays audio data generated by the server and provides responses to the user. Travelers can enjoy an interactive sightseeing experience through this process. The use of prompts in this sequence improves the accuracy of the generated AI model and user satisfaction.
[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0122] This invention is a travel guidance system using an AI character that incorporates an emotion engine to recognize the traveler's emotions and improve the travel experience. The system mainly consists of a server, a user terminal, an AI character, and an emotion engine, and functions as follows:
[0123] 1. Server-based plan generation and sentiment analysis
[0124] Users input their travel destinations and interests into a device, and based on this, the server uses a generative model to create a customized travel plan. Furthermore, an emotion engine is used to analyze data obtained from the user's voice and facial expressions to determine the user's emotional state. The server uses this information to dynamically adjust the travel plan and guidance content, personalizing the user's travel experience.
[0125] 2. Speech synthesis and emotion-based real-time guidance
[0126] Based on the generated plan, the server creates guidance dialogue that reflects the output of the emotion engine and converts it into a character's voice using speech synthesis. The terminal receives this audio data and provides the user with real-time, emotion-responsive voice guidance. The voice guidance is delivered in a tone and speed that matches the user's emotions, resulting in a more natural and engaging guided experience.
[0127] 3. Emotional interaction with the user
[0128] Users interact with an AI character via their device while traveling. The server utilizes an emotion engine to analyze the user's voice and facial expressions, generating real-time responses based on their current emotional state. For example, if the user appears tired, it suggests rest stops; if they seem to be enjoying themselves, it recommends further activities.
[0129] Specific example
[0130] As a concrete example, imagine a user is a traveler interested in history who wants to tour historical landmarks in New York City. The user enters "I want to visit historic buildings in New York" into their device. The server uses a generative model and an emotion engine to create a plan that includes landmarks such as Central Park and the Flatiron Building. During the tour, the character will announce in real time, "Next is the Flatiron Building," and if the emotion engine recognizes that the user's expression shows surprise, it will add a more detailed explanation, such as, "This building was constructed in 1902 and is characterized by its triangular exterior." In this way, the guidance content is flexibly adjusted according to the user's emotional state, enhancing the quality of the trip.
[0131] This system aims to provide travelers with innovative and engaging guidance services by deeply understanding their emotions and individually optimizing their travel experience.
[0132] The following describes the processing flow.
[0133] Step 1:
[0134] The user uses their device to input their travel destination and topics of interest. The device then sends the entered information to the server.
[0135] Step 2:
[0136] The server receives user input and uses a generative model to create a travel plan tailored to the user's preferences. The server also references a database to collect tourist destination information.
[0137] Step 3:
[0138] The server starts the emotion engine and prepares to analyze the user's voice and facial expression data. The user grants permission for the device to access the camera and microphone.
[0139] Step 4:
[0140] The device captures the user's voice and facial expressions in real time and sends them to a server as emotion data. The data is processed while protecting the user's privacy.
[0141] Step 5:
[0142] The server's emotion engine analyzes the user's emotional state and dynamically adjusts the generated travel plan and information based on the results.
[0143] Step 6:
[0144] The server prepares a pre-arranged guided speech and converts it into a character's voice using speech synthesis technology. The generated audio data is then sent to the terminal.
[0145] Step 7:
[0146] The device receives audio data from the server and provides real-time, emotion-responsive voice guidance to the user. The guidance content is tailored to the user's emotions.
[0147] Step 8:
[0148] The user enters additional questions or requests into the device during the guidance process. The device then sends the interaction data to the server.
[0149] Step 9:
[0150] The server processes additional user interactions and generates an appropriate response, taking into account the results of the emotion engine. The response is then converted into speech and sent to the terminal.
[0151] Step 10:
[0152] The device provides the user with generated responses and offers continuous guidance to enrich the travel experience.
[0153] (Example 2)
[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0155] Traditional travel guidance systems have struggled to individually optimize the travel experience based on travelers' emotions and real-time interests. In particular, the one-way nature of voice guidance is problematic, as it lacks the flexibility to respond to users' immediate emotional changes. Furthermore, personalization using information specific to individual travelers, such as past travel history, is insufficient.
[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0157] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for providing a speech synthesis device for a virtual character to provide voice guidance, and means for using an emotion analysis device to analyze the user's emotional state. This makes it possible to dynamically adjust guidance according to the traveler's emotions and provide a personalized travel experience.
[0158] A "travel plan" is a series of guided information that suggests tourist attractions and activities based on the traveler's destination and interests.
[0159] A "generative model" is an artificial intelligence technique used to generate new information from given inputs, based on a large dataset.
[0160] A "virtual character" is a computer-generated character, whether human or animal, that interacts with or guides the user.
[0161] A "speech synthesis device" is a device that has the function of converting text data into speech data and providing information to the user as speech.
[0162] "User's emotional state" refers to the user's psychological or emotional condition at a given time, analyzed based on data such as the user's voice and facial expressions.
[0163] An "emotion analysis device" is a device used to analyze a user's facial expressions and voice, and to estimate their emotions from that information.
[0164] A "dynamic adjustment mechanism" refers to a system that can change the content of guidance and the services provided in response to real-time changes in the situation.
[0165] A "personalized travel experience" refers to a personalized travel experience that is optimized for the individual interests and emotional state of the traveler.
[0166] This invention is an interactive travel guidance system that analyzes travelers' emotional states in real time and personalizes their travel experience. This system primarily consists of three components: a server, a terminal, and a user. The following describes how each component functions.
[0167] The server uses a generative AI model to generate a customized travel plan based on destination information and interests sent from the user via their device. The generative AI model used is based on a large dataset, analyzes text-based input, and provides a new travel plan in response to the prompt "Provide a prompt for the user to enter places they want to visit and generate a travel plan."
[0168] The device is equipped with a camera and microphone to record the user's voice and facial expressions. This data is transmitted to a server in real time, and the server estimates the user's emotional state through an emotion analysis device. This analysis uses software incorporating existing emotion analysis algorithms to evaluate the user's psychological state, such as whether they are enjoying themselves or feeling tired.
[0169] Users interact with an AI character using a smartphone or other mobile device. This character plays back guidance speech generated on a server using a speech synthesis device. The speech synthesis uses common software for converting text data into speech, and the tone and speed are adjusted according to the user's emotional state.
[0170] For example, if a user enters "I want to visit a quiet tourist spot" into their device, the server uses a generative AI model to generate a plan suggesting quiet tourist spots. If the user appears tired during their visit, the system will provide voice guidance such as, "There's a cafe nearby. Why don't you take a break?" In this way, by providing personalized travel guidance based on the user's emotional state in real time, the quality of the travel experience can be improved.
[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0172] Step 1:
[0173] The user enters their travel destination and interests into the device. The device formats this data to send it to the server. This input includes specific keywords related to the places they want to visit and their interests. The device formats this text data and generates text packets to send to the server.
[0174] Step 2:
[0175] The server generates a travel plan using a generative AI model based on destination information and interests received from the terminal. The server inputs a prompt message into the AI model: "Provide a prompt that allows the user to input places they want to visit and generate a travel plan." The AI model analyzes this input and outputs a travel plan that meets the user's needs. This output includes information on tourist spots to visit and recommended routes.
[0176] Step 3:
[0177] When a user begins interacting with the device, the device uses its camera and microphone to collect facial and audio data from the user. This data is used to analyze the user's emotional state. The device collects this raw data, converts it into a format suitable for emotion analysis, and sends it to the server.
[0178] Step 4:
[0179] The server passes the received facial and audio data to an emotion analysis device to analyze the user's emotional state. The analysis engine evaluates whether the user is surprised, amused, or tired based on factors such as voice tone, volume, and changes in facial expression. Based on this, it outputs parameters indicating the user's psychological state.
[0180] Step 5:
[0181] The server dynamically adjusts the guidance content using the generated travel plan and the results of the emotion analysis. The server generates guidance talk tailored to the user's emotional state and converts it into audio data using a speech synthesizer. For example, if the analysis indicates that the user is surprised, it incorporates explanations with more detailed background information. Audio data based on this dynamic adjustment is output.
[0182] Step 6:
[0183] The device receives audio data from the server and plays audio guidance to the user in real time. The device applies settings to play the audio at a tone and speed that matches the user's emotional state, providing guidance that naturally appeals to both sight and hearing. The played audio data is delivered to the user, complementing their travel experience.
[0184] (Application Example 2)
[0185] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0186] Conventional travel guidance systems often provide uniform plans and information without considering the emotional state of travelers, making it difficult to enhance the experience according to the individual emotions and preferences of travelers. Furthermore, when using autonomous vehicles as a means of transportation, dynamic adjustment of guidance content in response to the emotions of travelers during operation is required, but there is a lack of technology to achieve this.
[0187] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0188] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests; means for a proxy to provide voice guidance based on the generated travel plan; means for playing the voice guidance on the user's device and providing the traveler with real-time tourist information; means for analyzing the passenger's emotional state and dynamically adjusting the travel plan and guidance content using that information; and means for receiving feedback and questions from the user and generating appropriate responses using the generative model. This makes it possible to provide a personalized travel experience that responds to the traveler's emotions.
[0189] A "travel plan" is a detailed itinerary that outlines the places to visit and activities to do, based on the traveler's destination information and interests.
[0190] A "generative model" refers to an algorithm or data processing system that calculates, selects, and proposes the optimal travel plan based on the traveler's input information.
[0191] "Speech synthesis means" refers to technologies and devices for converting text information into speech and outputting it.
[0192] "Emotional analysis methods" are technologies that use data such as a traveler's voice and video to determine their emotional state at a given time.
[0193] "Dynamic adjustment" refers to the process of changing the itinerary and guidance content in real time according to the situation.
[0194] "Personalization" refers to optimizing the experience according to each traveler's individual characteristics, preferences, and emotions.
[0195] To implement this invention, a system is needed that generates travel plans based on travelers' destination information and interests. The server employs a generation AI model to create a travel plan based on this information and automatically suggests optimal destinations and activities. This allows for the provision of personalized travel plans for each traveler.
[0196] The server generates voice guidance based on a plan created using speech synthesis technology and transmits it to the user's mobile device or the navigation system of the autonomous vehicle. The device plays the voice guidance received from the server as is, providing the traveler with real-time tourist information. This voice guidance is adjusted in real time, taking into account the traveler's current emotional state.
[0197] Furthermore, cameras and microphones installed inside the train cars are used to analyze passengers' emotions. Through these input devices, the emotion analysis system analyzes the passengers' emotional state and sends the data to a server. The server uses this information to dynamically adjust the travel plan and information provided.
[0198] For example, if a traveler's emotions are analyzed while visiting a tourist attraction, and they appear tired, the server might suggest a place to rest as their next destination. Conversely, if they appear excited, it could suggest more active activities.
[0199] Example of a prompt:
[0200] "Are the passengers excited? Please generate information about the next tourist attraction accordingly."
[0201] "If the passengers are tired, please suggest a route that allows them to take a break."
[0202] These features allow the system to provide an innovative travel experience that incorporates the traveler's emotions.
[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0204] Step 1:
[0205] The server receives the traveler's destination information and interests. User information is sent via the terminal as input. The server uses this input data to generate prompts for the AI model and constructs a travel plan based on that information. Through this process, suggestions tailored to the traveler's interests are formed.
[0206] Step 2:
[0207] The terminal receives travel plan information sent from the server. The terminal passes this information to a speech synthesis system, which converts the text information into audio data to provide guidance to the traveler. The input is text-based guidance information from the server, and the output is generated audio guidance. In this step, acoustic data processing is performed with the aim of conveying the guide content in a natural tone and speaking speed.
[0208] Step 3:
[0209] The server analyzes the user's emotions in real time via cameras and microphones inside the vehicle. Input includes audio and video data from the emotion analysis system, and the server processes this data to infer the emotional state. The output is the inferred emotional state. This information is used in the next step to adjust the travel plan.
[0210] Step 4:
[0211] The server dynamically adjusts travel plans and guidance based on the emotion analysis results. The input is generated emotion state data, which is used to create new guidance and routes optimized for the traveler's emotions. The output is an updated travel plan, which is then sent back to the terminal for speech synthesis. This process personalizes the traveler's experience.
[0212] Step 5:
[0213] The user sends feedback and additional questions to the server via their device. Based on this input, the server uses a generative AI model to generate an appropriate response and sends it back to the device. The output here is a customized answer or suggestion to the user's inquiry, which is then provided to the traveler again as voice guidance.
[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0217] [Second Embodiment]
[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0230] This invention is a system that provides professional travel guidance using an AI character, generating personalized travel plans based on the traveler's destination and interests, and providing real-time voice guidance. The system mainly consists of a server, a user terminal, and an AI character, and functions as follows:
[0231] 1. Server-driven plan generation
[0232] The server uses a generative model to create customized travel plans based on destination and interest information received from the user. This generative model has learned from a large amount of tourist destination data and guide expertise, and makes suggestions that take into account the optimal order of visits and modes of transportation. The server also reflects the user's past travel history and preferences to personalize the plan.
[0233] 2. Speech synthesis and real-time guidance
[0234] Based on a travel plan generated on the server, talk segments about each tourist destination in the plan are created and converted into a character's voice using speech synthesis technology. The user terminal receives this voice data and provides real-time voice guidance during the trip. Speech synthesis technology enables natural-sounding guidance at visited locations, supporting a smooth traveler experience.
[0235] 3. User Interaction
[0236] Users can interact with an AI character through their device to get answers to questions that arise during their sightseeing and to request additional information. The server analyzes the user's questions and requests and returns appropriate responses in real time using a generative model. This allows the character to support travelers as if they were a real-life guide.
[0237] Specific example
[0238] For example, consider a user who is a traveler interested in Japanese culture and is visiting Kyoto. The user sends a request from their device saying, "I want to visit traditional temples in Kyoto." The server uses a generative model to generate a travel plan that includes representative locations such as Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. During the trip, the user's device provides voice guidance in a character's voice, such as, "Next is Kinkaku-ji. Enjoy the iconic architectural style of the Muromachi period." If the user then asks, "I want to know more about the history," the server provides detailed historical background information based on that request.
[0239] In this way, this system utilizes AI technology to provide travelers with high-value tourism experiences and aims to improve the quality of their travels.
[0240] The following describes the processing flow.
[0241] Step 1:
[0242] The user enters their travel destination and areas of interest into the device. The device then sends the entered information to the server.
[0243] Step 2:
[0244] The server receives information from the user and starts a generative model to generate a travel plan. The server consults a database to collect information about relevant tourist destinations.
[0245] Step 3:
[0246] The server uses collected tourist destination information and a generative model to create a travel plan optimized for the user. The plan includes the order of visits, modes of transportation, and time allocation.
[0247] Step 4:
[0248] Based on the travel plan created by the server, the content of the guided conversation is constructed. The server then converts the conversation content into a character's voice using speech synthesis technology.
[0249] Step 5:
[0250] The device receives audio data sent from the server. The device then uses this data to begin providing real-time audio guidance to the user.
[0251] Step 6:
[0252] The user interacts with the character while traveling. The device forwards user questions and requests to the server via voice or text.
[0253] Step 7:
[0254] The server receives interaction from the user and generates an appropriate response using a generative model. The server converts the response into audio data and sends it to the terminal.
[0255] Step 8:
[0256] The device relays responses sent from the server to the user. The device continuously provides updated plans and information in real time as needed.
[0257] (Example 1)
[0258] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0259] Modern travelers tend to demand personalized information and real-time guidance, but traditional guide systems are unable to adequately meet these demands. Specifically, it has been difficult to provide personalized itineraries based on travelers' preferences and past travel history, as well as effective support including immediate responses to questions and requests.
[0260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0261] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for using a speech synthesis device for a character to provide voice guidance based on the generated travel plan, and means for personalizing the travel plan based on the traveler's past visit history and preferences. This enables more personalized and real-time information and guidance for the traveler.
[0262] A "traveler" refers to an individual who visits a tourist destination or place of interest and seeks information and guidance regarding their itinerary.
[0263] "Destination information" refers to information about specific tourist destinations or cities that travelers wish to visit.
[0264] A "generative model" refers to methods and algorithms, including artificial intelligence technology, used to create travel plans, which make suggestions based on tourist destination data and guide expertise.
[0265] A "speech synthesis device" refers to a technology that converts text data into speech data, enabling information to be transmitted using a human-like voice.
[0266] "Personalization" refers to the process of adjusting services and information based on the individual traveler's preferences and past visit history.
[0267] "Real-time" refers to a state where information and services can be provided immediately, meaning that travelers can resolve their questions and requests on the spot.
[0268] This invention is a system for providing travelers with personalized travel plans and real-time voice guidance. The main components of the invention are a server, a user terminal, and a generative AI model.
[0269] The server analyzes destination information and interest prompts provided by travelers and generates personalized travel plans based on this analysis. Natural language processing technology is used for this analysis, and a generating AI model utilizes tourist data and guide expertise to create the optimal plan. For example, if a user sends the prompt "I want to visit traditional temples in Kyoto," the server will generate a plan including Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. In doing so, it also considers past travel history and preferences to provide the most fulfilling travel experience for the user.
[0270] The device provides travelers with audio data based on their travel plan, received from the server. Utilizing speech synthesis technology, an AI character delivers information about each tourist spot as if it were a real-life guide. For example, as a user is on their way to Kinkaku-ji Temple, the device will play an audio announcement saying, "Next stop is Kinkaku-ji Temple. Please enjoy this iconic architectural style from the Muromachi period."
[0271] Users can interact with an AI character via their device during their trip, getting their questions answered and seeking more in-depth information. The server analyzes the received requests in real time, uses a generative AI model to quickly generate appropriate responses, and provides them to the user. This process allows travelers to gain a more valuable experience than they could have planned on their own.
[0272] Thus, this invention uses AI technology to improve the quality of travel and ensure consistency and convenience in providing information to travelers.
[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0274] Step 1:
[0275] The user uses the terminal to enter prompts about their desired destination or interests. Specifically, if the user enters "I want to visit traditional temples in Kyoto," a prompt will be generated. This is sent to the server and becomes input data for the next processing step.
[0276] Step 2:
[0277] The server analyzes the received prompt text. Natural language processing is used for this analysis to understand the user's intention and desired destination. Based on the data analysis, the server utilizes a generative AI model to generate a personalized travel plan from tourist destination data and guide know-how. The generated travel plan becomes the output data to be used in the next step.
[0278] Step 3:
[0279] Based on the travel plan generated by the server, the content of the voice guidance for each tourist destination is created. The server uses voice synthesis technology to convert the explanations about the destinations in the plan into natural voices. This voice data is generated and becomes the output to be sent to the user terminal.
[0280] Step 4:
[0281] The terminal plays the voice data received from the server and provides real-time guidance to the traveler. When the user visits the Golden Pavilion, a voice guidance such as "Next is the Golden Pavilion. Please enjoy the symbolic architectural style of the Muromachi period" flows from the terminal. As a result, the user can receive a real guide experience.
[0282] Step 5:
[0283] The user can request additional information or ask questions during the trip. These requests are sent to the server through the terminal. The user specifically says something like "I want to know more about the history of this temple", and that request becomes new input data.
[0284] Step 6:
[0285] The server analyzes requests from users in real time and creates appropriate responses using a generative AI model. Based on the analysis results, newly created responses are generated as voice data and sent to the terminal. This enables users to obtain the necessary information immediately.
[0286] (Application Example 1)
[0287] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0288] In conventional travel guidance systems, it has been difficult to provide travel plans tailored to individual travelers, and the quality of real-time guidance and interaction has also been limited. In addition, since there has been a lack of means to provide virtual tourism experiences before an actual trip, it has not been sufficient to offer attractive experiences to users during the trip planning stage. As a result, there remain issues such as flexible guidance according to travelers' needs and improvement of travel motivation through pre-trip virtual experiences.
[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0290] In this invention, the server includes means for using a generative model to generate a travel plan based on travelers' destination information and interests, voice synthesis means for a character to provide voice guidance based on the generated travel plan, and means for providing guidance to the user in a virtual environment via visual augmentation technology and presenting guidance related to each scene both audibly and visually. This enables the provision of personalized travel plans according to individual needs of travelers and real-time and interactive tourism guidance in a virtual environment.
[0291] "Travelers" refers to people who visit a specific destination and are individuals or groups moving for various purposes such as tourism or business.
[0292] "Destination information" refers to information about the geographical location and related tourist resources that travelers intend to visit, and includes specific place names and tourist attractions.
[0293] A "generative model" refers to an algorithm or technology used to automatically create travel plans based on input information. It is a system that learns from large amounts of data and patterns to generate responses.
[0294] "Speech synthesis means" refers to technology that converts text data into speech signals, and includes devices and software for artificially generating and providing speech to users.
[0295] "Visual terminals" refer to devices that allow users to receive visual information, and include smart glasses and head-mounted displays.
[0296] "Feedback" refers to the opinions, evaluations, and questions that users provide to a system, and is used to improve the system and generate responses.
[0297] "Visual augmentation technology" refers to the technology of overlaying virtual information onto the real world, and is a means of improving the user's experience of the real world.
[0298] A "virtual environment" refers to a virtual space or situation created using computer technology, allowing users to experience a space that does not exist in reality.
[0299] "Interactive tourist information" refers to a form of information gathering where travelers obtain information through two-way communication with the system, and it is a dynamic guidance method that changes its response according to the user's input.
[0300] This invention is a system that provides personalized travel guidance for travelers, and mainly consists of a server, a terminal, and user interaction.
[0301] The server uses a generation model to generate travel plans based on travelers' destination information and interests. The generation model has learned a large amount of tourist destination data and guides' know-how, and provides a plan considering the optimal visiting order and means of transportation. Then, based on the generated travel plan, the server uses voice synthesis means to create talks about each tourist destination in the plan, converts them into the voices of AI characters, and sends them to the terminal.
[0302] The terminal plays the received voice data and provides tourist destination information to travelers in real time. Also, through visual augmentation technology, it enables travelers to experience visual information in a virtual environment. In this process, smart glasses and head-mounted displays are used.
[0303] The user can interact with the AI character through these devices. Through the interaction, when the user makes a request for feedback or additional information, that information is sent from the terminal to the server. The server analyzes the received feedback and reuses the generation model to generate an appropriate response.
[0304] As a specific example, when the user requests "looking for summer clothes", the server creates a virtual tour plan for stores handling the latest summer clothes. Based on this, the terminal guides, "This is a store that showcases the latest summer clothing collection."
[0305] An example of the prompt sentence in the generation AI model is "The user hopes to purchase summer clothes. Introduce stores incorporating the latest fashion trends and generate an audio guide."
[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0307] Step 1:
[0308] The user uses a terminal to input travel destination information and interests. This information is sent to a server. The server analyzes the data and uses a generative AI model to create a customized travel plan. The user's past travel history and preferences are also considered to determine the optimal order of visits and modes of transportation. The output is an optimized travel plan.
[0309] Step 2:
[0310] The server creates guided tours for each tourist spot based on the generated travel plan. Using speech synthesis technology, this talk is converted into the voice of an AI character. The input is the text data of the generated travel plan, and the output is an audio file. This audio data is then sent to the terminal.
[0311] Step 3:
[0312] The terminal plays audio data received from the server, providing real-time guidance to travelers. Simultaneously, users experience information within the virtual environment visually using visual augmentation technology. Specifically, this involves displaying information using smart glasses or head-mounted displays.
[0313] Step 4:
[0314] The user inputs questions and requests for additional information during the tour via voice. The input is sent from the terminal to the server. The server analyzes this information and uses a generated AI model to produce appropriate responses to the user's questions and requests. The output is voice data for the AI character to provide answers.
[0315] Step 5:
[0316] The terminal plays audio data generated by the server and provides responses to the user. Travelers can enjoy an interactive sightseeing experience through this process. The use of prompts in this sequence improves the accuracy of the generated AI model and user satisfaction.
[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0318] This invention is a travel guidance system using an AI character that incorporates an emotion engine to recognize the traveler's emotions and improve the travel experience. The system mainly consists of a server, a user terminal, an AI character, and an emotion engine, and functions as follows:
[0319] 1. Server-based plan generation and sentiment analysis
[0320] Users input their travel destinations and interests into a device, and based on this, the server uses a generative model to create a customized travel plan. Furthermore, an emotion engine is used to analyze data obtained from the user's voice and facial expressions to determine the user's emotional state. The server uses this information to dynamically adjust the travel plan and guidance content, personalizing the user's travel experience.
[0321] 2. Speech synthesis and emotion-based real-time guidance
[0322] Based on the generated plan, the server creates guidance dialogue that reflects the output of the emotion engine and converts it into a character's voice using speech synthesis. The terminal receives this audio data and provides the user with real-time, emotion-responsive voice guidance. The voice guidance is delivered in a tone and speed that matches the user's emotions, resulting in a more natural and engaging guided experience.
[0323] 3. Emotional interaction with the user
[0324] Users interact with an AI character via their device while traveling. The server utilizes an emotion engine to analyze the user's voice and facial expressions, generating real-time responses based on their current emotional state. For example, if the user appears tired, it suggests rest stops; if they seem to be enjoying themselves, it recommends further activities.
[0325] Specific example
[0326] As a concrete example, imagine a user is a traveler interested in history who wants to tour historical landmarks in New York City. The user enters "I want to visit historic buildings in New York" into their device. The server uses a generative model and an emotion engine to create a plan that includes landmarks such as Central Park and the Flatiron Building. During the tour, the character will announce in real time, "Next is the Flatiron Building," and if the emotion engine recognizes that the user's expression shows surprise, it will add a more detailed explanation, such as, "This building was constructed in 1902 and is characterized by its triangular exterior." In this way, the guidance content is flexibly adjusted according to the user's emotional state, enhancing the quality of the trip.
[0327] This system aims to provide travelers with innovative and engaging guidance services by deeply understanding their emotions and individually optimizing their travel experience.
[0328] The following describes the processing flow.
[0329] Step 1:
[0330] The user uses their device to input their travel destination and topics of interest. The device then sends the entered information to the server.
[0331] Step 2:
[0332] The server receives user input and uses a generative model to create a travel plan tailored to the user's preferences. The server also references a database to collect tourist destination information.
[0333] Step 3:
[0334] The server starts the emotion engine and prepares to analyze the user's voice and facial expression data. The user grants permission for the device to access the camera and microphone.
[0335] Step 4:
[0336] The device captures the user's voice and facial expressions in real time and sends them to a server as emotion data. The data is processed while protecting the user's privacy.
[0337] Step 5:
[0338] The server's emotion engine analyzes the user's emotional state and dynamically adjusts the generated travel plan and information based on the results.
[0339] Step 6:
[0340] The server prepares a pre-arranged guided speech and converts it into a character's voice using speech synthesis technology. The generated audio data is then sent to the terminal.
[0341] Step 7:
[0342] The device receives audio data from the server and provides real-time, emotion-responsive voice guidance to the user. The guidance content is tailored to the user's emotions.
[0343] Step 8:
[0344] The user enters additional questions or requests into the device during the guidance process. The device then sends the interaction data to the server.
[0345] Step 9:
[0346] The server processes additional user interactions and generates an appropriate response, taking into account the results of the emotion engine. The response is then converted into speech and sent to the terminal.
[0347] Step 10:
[0348] The device provides the user with generated responses and offers continuous guidance to enrich the travel experience.
[0349] (Example 2)
[0350] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0351] Traditional travel guidance systems have struggled to individually optimize the travel experience based on travelers' emotions and real-time interests. In particular, the one-way nature of voice guidance is problematic, as it lacks the flexibility to respond to users' immediate emotional changes. Furthermore, personalization using information specific to individual travelers, such as past travel history, is insufficient.
[0352] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0353] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for providing a speech synthesis device for a virtual character to provide voice guidance, and means for using an emotion analysis device to analyze the user's emotional state. This makes it possible to dynamically adjust guidance according to the traveler's emotions and provide a personalized travel experience.
[0354] A "travel plan" is a series of guided information that suggests tourist attractions and activities based on the traveler's destination and interests.
[0355] A "generative model" is an artificial intelligence technique used to generate new information from given inputs, based on a large dataset.
[0356] A "virtual character" is a computer-generated character, whether human or animal, that interacts with or guides the user.
[0357] A "speech synthesis device" is a device that has the function of converting text data into speech data and providing information to the user as speech.
[0358] "User's emotional state" refers to the user's psychological or emotional condition at a given time, analyzed based on data such as the user's voice and facial expressions.
[0359] An "emotion analysis device" is a device used to analyze a user's facial expressions and voice, and to estimate their emotions from that information.
[0360] A "dynamic adjustment mechanism" refers to a system that can change the content of guidance and the services provided in response to real-time changes in the situation.
[0361] A "personalized travel experience" refers to a personalized travel experience that is optimized for the individual interests and emotional state of the traveler.
[0362] This invention is an interactive travel guidance system that analyzes travelers' emotional states in real time and personalizes their travel experience. This system primarily consists of three components: a server, a terminal, and a user. The following describes how each component functions.
[0363] The server uses a generative AI model to generate a customized travel plan based on destination information and interests sent from the user via their device. The generative AI model used is based on a large dataset, analyzes text-based input, and provides a new travel plan in response to the prompt "Provide a prompt for the user to enter places they want to visit and generate a travel plan."
[0364] The device is equipped with a camera and microphone to record the user's voice and facial expressions. This data is transmitted to a server in real time, and the server estimates the user's emotional state through an emotion analysis device. This analysis uses software incorporating existing emotion analysis algorithms to evaluate the user's psychological state, such as whether they are enjoying themselves or feeling tired.
[0365] Users interact with an AI character using a smartphone or other mobile device. This character plays back guidance speech generated on a server using a speech synthesis device. The speech synthesis uses common software for converting text data into speech, and the tone and speed are adjusted according to the user's emotional state.
[0366] For example, if a user enters "I want to visit a quiet tourist spot" into their device, the server uses a generative AI model to generate a plan suggesting quiet tourist spots. If the user appears tired during their visit, the system will provide voice guidance such as, "There's a cafe nearby. Why don't you take a break?" In this way, by providing personalized travel guidance based on the user's emotional state in real time, the quality of the travel experience can be improved.
[0367] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0368] Step 1:
[0369] The user enters their travel destination and interests into the device. The device formats this data to send it to the server. This input includes specific keywords related to the places they want to visit and their interests. The device formats this text data and generates text packets to send to the server.
[0370] Step 2:
[0371] The server generates a travel plan using a generative AI model based on destination information and interests received from the terminal. The server inputs a prompt message into the AI model: "Provide a prompt that allows the user to input places they want to visit and generate a travel plan." The AI model analyzes this input and outputs a travel plan that meets the user's needs. This output includes information on tourist spots to visit and recommended routes.
[0372] Step 3:
[0373] When a user begins interacting with the device, the device uses its camera and microphone to collect facial and audio data from the user. This data is used to analyze the user's emotional state. The device collects this raw data, converts it into a format suitable for emotion analysis, and sends it to the server.
[0374] Step 4:
[0375] The server passes the received facial and audio data to an emotion analysis device to analyze the user's emotional state. The analysis engine evaluates whether the user is surprised, amused, or tired based on factors such as voice tone, volume, and changes in facial expression. Based on this, it outputs parameters indicating the user's psychological state.
[0376] Step 5:
[0377] The server dynamically adjusts the guidance content using the generated travel plan and the results of the emotion analysis. The server generates guidance talk tailored to the user's emotional state and converts it into audio data using a speech synthesizer. For example, if the analysis indicates that the user is surprised, it incorporates explanations with more detailed background information. Audio data based on this dynamic adjustment is output.
[0378] Step 6:
[0379] The device receives audio data from the server and plays audio guidance to the user in real time. The device applies settings to play the audio at a tone and speed that matches the user's emotional state, providing guidance that naturally appeals to both sight and hearing. The played audio data is delivered to the user, complementing their travel experience.
[0380] (Application Example 2)
[0381] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0382] Conventional travel guidance systems often provide uniform plans and information without considering the emotional state of travelers, making it difficult to enhance the experience according to the individual emotions and preferences of travelers. Furthermore, when using autonomous vehicles as a means of transportation, dynamic adjustment of guidance content in response to the emotions of travelers during operation is required, but there is a lack of technology to achieve this.
[0383] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0384] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests; means for a proxy to provide voice guidance based on the generated travel plan; means for playing the voice guidance on the user's device and providing the traveler with real-time tourist information; means for analyzing the passenger's emotional state and dynamically adjusting the travel plan and guidance content using that information; and means for receiving feedback and questions from the user and generating appropriate responses using the generative model. This makes it possible to provide a personalized travel experience that responds to the traveler's emotions.
[0385] A "travel plan" is a detailed itinerary that outlines the places to visit and activities to do, based on the traveler's destination information and interests.
[0386] A "generative model" refers to an algorithm or data processing system that calculates, selects, and proposes the optimal travel plan based on the traveler's input information.
[0387] "Speech synthesis means" refers to technologies and devices for converting text information into speech and outputting it.
[0388] "Emotional analysis methods" are technologies that use data such as a traveler's voice and video to determine their emotional state at a given time.
[0389] "Dynamic adjustment" refers to the process of changing the itinerary and guidance content in real time according to the situation.
[0390] "Personalization" refers to optimizing the experience according to each traveler's individual characteristics, preferences, and emotions.
[0391] To implement this invention, a system is needed that generates travel plans based on travelers' destination information and interests. The server employs a generation AI model to create a travel plan based on this information and automatically suggests optimal destinations and activities. This allows for the provision of personalized travel plans for each traveler.
[0392] The server generates voice guidance based on a plan created using speech synthesis technology and transmits it to the user's mobile device or the navigation system of the autonomous vehicle. The device plays the voice guidance received from the server as is, providing the traveler with real-time tourist information. This voice guidance is adjusted in real time, taking into account the traveler's current emotional state.
[0393] Furthermore, cameras and microphones installed inside the train cars are used to analyze passengers' emotions. Through these input devices, the emotion analysis system analyzes the passengers' emotional state and sends the data to a server. The server uses this information to dynamically adjust the travel plan and information provided.
[0394] For example, if a traveler's emotions are analyzed while visiting a tourist attraction, and they appear tired, the server might suggest a place to rest as their next destination. Conversely, if they appear excited, it could suggest more active activities.
[0395] Example of a prompt:
[0396] "Are the passengers excited? Please generate information about the next tourist attraction accordingly."
[0397] "If the passengers are tired, please suggest a route that allows them to take a break."
[0398] These features allow the system to provide an innovative travel experience that incorporates the traveler's emotions.
[0399] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0400] Step 1:
[0401] The server receives the traveler's destination information and interests. User information is sent via the terminal as input. The server uses this input data to generate prompts for the AI model and constructs a travel plan based on that information. Through this process, suggestions tailored to the traveler's interests are formed.
[0402] Step 2:
[0403] The terminal receives travel plan information sent from the server. The terminal passes this information to a speech synthesis system, which converts the text information into audio data to provide guidance to the traveler. The input is text-based guidance information from the server, and the output is generated audio guidance. In this step, acoustic data processing is performed with the aim of conveying the guide content in a natural tone and speaking speed.
[0404] Step 3:
[0405] The server analyzes the user's emotions in real time via cameras and microphones inside the vehicle. Input includes audio and video data from the emotion analysis system, and the server processes this data to infer the emotional state. The output is the inferred emotional state. This information is used in the next step to adjust the travel plan.
[0406] Step 4:
[0407] The server dynamically adjusts travel plans and guidance based on the emotion analysis results. The input is generated emotion state data, which is used to create new guidance and routes optimized for the traveler's emotions. The output is an updated travel plan, which is then sent back to the terminal for speech synthesis. This process personalizes the traveler's experience.
[0408] Step 5:
[0409] The user sends feedback and additional questions to the server via their device. Based on this input, the server uses a generative AI model to generate an appropriate response and sends it back to the device. The output here is a customized answer or suggestion to the user's inquiry, which is then provided to the traveler again as voice guidance.
[0410] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0411] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0412] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0413] [Third Embodiment]
[0414] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0415] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0416] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0417] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0418] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0420] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0421] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0422] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0423] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0424] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0425] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0426] This invention is a system that provides professional travel guidance using an AI character, generating personalized travel plans based on the traveler's destination and interests, and providing real-time voice guidance. The system mainly consists of a server, a user terminal, and an AI character, and functions as follows:
[0427] 1. Server-driven plan generation
[0428] The server uses a generative model to create customized travel plans based on destination and interest information received from the user. This generative model has learned from a large amount of tourist destination data and guide expertise, and makes suggestions that take into account the optimal order of visits and modes of transportation. The server also reflects the user's past travel history and preferences to personalize the plan.
[0429] 2. Speech synthesis and real-time guidance
[0430] Based on a travel plan generated on the server, talk segments about each tourist destination in the plan are created and converted into a character's voice using speech synthesis technology. The user terminal receives this voice data and provides real-time voice guidance during the trip. Speech synthesis technology enables natural-sounding guidance at visited locations, supporting a smooth traveler experience.
[0431] 3. User Interaction
[0432] Users can interact with an AI character through their device to get answers to questions that arise during their sightseeing and to request additional information. The server analyzes the user's questions and requests and returns appropriate responses in real time using a generative model. This allows the character to support travelers as if they were a real-life guide.
[0433] Specific example
[0434] For example, consider a user who is a traveler interested in Japanese culture and is visiting Kyoto. The user sends a request from their device saying, "I want to visit traditional temples in Kyoto." The server uses a generative model to generate a travel plan that includes representative locations such as Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. During the trip, the user's device provides voice guidance in a character's voice, such as, "Next is Kinkaku-ji. Enjoy the iconic architectural style of the Muromachi period." If the user then asks, "I want to know more about the history," the server provides detailed historical background information based on that request.
[0435] In this way, this system utilizes AI technology to provide travelers with high-value tourism experiences and aims to improve the quality of their travels.
[0436] The following describes the processing flow.
[0437] Step 1:
[0438] The user enters their travel destination and areas of interest into the device. The device then sends the entered information to the server.
[0439] Step 2:
[0440] The server receives information from the user and starts a generative model to generate a travel plan. The server consults a database to collect information about relevant tourist destinations.
[0441] Step 3:
[0442] The server uses collected tourist destination information and a generative model to create a travel plan optimized for the user. The plan includes the order of visits, modes of transportation, and time allocation.
[0443] Step 4:
[0444] Based on the travel plan created by the server, the content of the guided conversation is constructed. The server then converts the conversation content into a character's voice using speech synthesis technology.
[0445] Step 5:
[0446] The device receives audio data sent from the server. The device then uses this data to begin providing real-time audio guidance to the user.
[0447] Step 6:
[0448] The user interacts with the character while traveling. The device forwards user questions and requests to the server via voice or text.
[0449] Step 7:
[0450] The server receives interaction from the user and generates an appropriate response using a generative model. The server converts the response into audio data and sends it to the terminal.
[0451] Step 8:
[0452] The device relays responses sent from the server to the user. The device continuously provides updated plans and information in real time as needed.
[0453] (Example 1)
[0454] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0455] Modern travelers tend to demand personalized information and real-time guidance, but traditional guide systems are unable to adequately meet these demands. Specifically, it has been difficult to provide personalized itineraries based on travelers' preferences and past travel history, as well as effective support including immediate responses to questions and requests.
[0456] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0457] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for using a speech synthesis device for a character to provide voice guidance based on the generated travel plan, and means for personalizing the travel plan based on the traveler's past visit history and preferences. This enables more personalized and real-time information and guidance for the traveler.
[0458] A "traveler" refers to an individual who visits a tourist destination or place of interest and seeks information and guidance regarding their itinerary.
[0459] "Destination information" refers to information about specific tourist destinations or cities that travelers wish to visit.
[0460] A "generative model" refers to methods and algorithms, including artificial intelligence technology, used to create travel plans, which make suggestions based on tourist destination data and guide expertise.
[0461] A "speech synthesis device" refers to a technology that converts text data into speech data, enabling information to be transmitted using a human-like voice.
[0462] "Personalization" refers to the process of adjusting services and information based on the individual traveler's preferences and past visit history.
[0463] "Real-time" refers to a state where information and services can be provided immediately, meaning that travelers can resolve their questions and requests on the spot.
[0464] This invention is a system for providing travelers with personalized travel plans and real-time voice guidance. The main components of the invention are a server, a user terminal, and a generative AI model.
[0465] The server analyzes destination information and interest prompts provided by travelers and generates personalized travel plans based on this analysis. Natural language processing technology is used for this analysis, and a generating AI model utilizes tourist data and guide expertise to create the optimal plan. For example, if a user sends the prompt "I want to visit traditional temples in Kyoto," the server will generate a plan including Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. In doing so, it also considers past travel history and preferences to provide the most fulfilling travel experience for the user.
[0466] The device provides travelers with audio data based on their travel plan, received from the server. Utilizing speech synthesis technology, an AI character delivers information about each tourist spot as if it were a real-life guide. For example, as a user is on their way to Kinkaku-ji Temple, the device will play an audio announcement saying, "Next stop is Kinkaku-ji Temple. Please enjoy this iconic architectural style from the Muromachi period."
[0467] Users can interact with an AI character via their device during their trip, getting their questions answered and seeking more in-depth information. The server analyzes the received requests in real time, uses a generative AI model to quickly generate appropriate responses, and provides them to the user. This process allows travelers to gain a more valuable experience than they could have planned on their own.
[0468] Thus, this invention uses AI technology to improve the quality of travel and ensure consistency and convenience in providing information to travelers.
[0469] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0470] Step 1:
[0471] The user uses the terminal to enter prompts about their desired destination or interests. Specifically, if the user enters "I want to visit traditional temples in Kyoto," a prompt will be generated. This is sent to the server and becomes input data for the next processing step.
[0472] Step 2:
[0473] The server parses the received prompt message. This parsing uses natural language processing to understand the user's intent and desired destinations. Based on the data analysis, the server utilizes a generative AI model to generate a personalized travel plan from tourist destination data and guide expertise. The generated travel plan becomes the output data used in the next step.
[0474] Step 3:
[0475] Based on the travel plan generated on the server, audio guides for each tourist destination are created. The server uses speech synthesis technology to convert explanations about the places visited in the plan into natural-sounding speech. This audio data is generated and becomes the output that is sent to the user's terminal.
[0476] Step 4:
[0477] The device plays audio data received from the server, providing real-time guidance to travelers. When a user visits Kinkaku-ji Temple, the device plays an audio announcement such as, "Next stop is Kinkaku-ji Temple. Please enjoy the iconic architectural style of the Muromachi period." This allows users to experience a realistic guided tour.
[0478] Step 5:
[0479] Users can request additional information or ask questions during their trip. These requests are sent to the server via their device. Specifically, the user might say something like, "I want to know more about the history of this temple," and that request becomes new input data.
[0480] Step 6:
[0481] The server analyzes user requests in real time and uses a generative AI model to create appropriate responses. Based on the analysis results, the newly created response is generated as audio data and sent to the terminal. This allows the user to instantly obtain the information they need.
[0482] (Application Example 1)
[0483] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0484] Traditional travel information systems have struggled to provide personalized travel plans for individual travelers, and the quality of real-time guidance and interaction has been limited. Furthermore, the lack of means to offer virtual sightseeing experiences before actual travel meant that users were not adequately provided with engaging experiences during the travel planning stage. This leaves challenges remaining, such as providing flexible guidance tailored to travelers' needs and increasing travel motivation through simulated experiences before the trip.
[0485] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0486] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for voice synthesis to have a character provide voice guidance based on the generated travel plan, and means for providing guidance in a virtual environment to the user via visual augmentation technology, and for presenting guidance related to each scene in audio and visual form. This enables the provision of personalized travel plans that meet the individual needs of travelers, as well as real-time and interactive sightseeing guidance in a virtual environment.
[0487] The term "traveler" refers to individuals or groups who visit a specific destination and travel for various purposes, such as tourism or business.
[0488] "Destination information" refers to information about the geographical location and related tourist resources that travelers intend to visit, and includes specific place names and tourist attractions.
[0489] A "generative model" refers to an algorithm or technology used to automatically create travel plans based on input information. It is a system that learns from large amounts of data and patterns to generate responses.
[0490] "Speech synthesis means" refers to technology that converts text data into speech signals, and includes devices and software for artificially generating and providing speech to users.
[0491] "Visual terminals" refer to devices that allow users to receive visual information, and include smart glasses and head-mounted displays.
[0492] "Feedback" refers to the opinions, evaluations, and questions that users provide to a system, and is used to improve the system and generate responses.
[0493] "Visual augmentation technology" refers to the technology of overlaying virtual information onto the real world, and is a means of improving the user's experience of the real world.
[0494] A "virtual environment" refers to a virtual space or situation created using computer technology, allowing users to experience a space that does not exist in reality.
[0495] "Interactive tourist information" refers to a form of information gathering where travelers obtain information through two-way communication with the system, and it is a dynamic guidance method that changes its response according to the user's input.
[0496] This invention is a system that provides personalized travel guidance for travelers, and mainly consists of a server, a terminal, and user interaction.
[0497] The server uses a generative model to generate travel plans based on the traveler's destination information and interests. The generative model learns from a large amount of tourist destination data and guide expertise, and provides plans that consider the optimal order of visits and means of transportation. Based on the generated travel plan, the server uses speech synthesis to create talk about each tourist destination in the plan, converts it into the voice of an AI character, and sends it to the terminal.
[0498] The device plays back received audio data and provides travelers with real-time tourist information. It also enables travelers to experience visual information within the virtual environment through visual augmentation technology. Smart glasses and head-mounted displays are used in this process.
[0499] Users can interact with AI characters through these devices. Through this interaction, users can request feedback or additional information, which is then sent from the device to the server. The server analyzes the received feedback and reuses the generative model to generate an appropriate response.
[0500] For example, if a user requests to "look for summer clothes," the server creates a virtual tour plan of stores that carry the latest summer clothing. Based on this, the device will guide the user by saying, "This is a store showcasing the latest summer clothing collection."
[0501] An example of a prompt in a generative AI model is: "The user wants to buy summer clothes. Introduce stores that feature the latest trends and generate an audio guide."
[0502] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0503] Step 1:
[0504] The user uses a terminal to input travel destination information and interests. This information is sent to a server. The server analyzes the data and uses a generative AI model to create a customized travel plan. The user's past travel history and preferences are also considered to determine the optimal order of visits and modes of transportation. The output is an optimized travel plan.
[0505] Step 2:
[0506] The server creates guided tours for each tourist spot based on the generated travel plan. Using speech synthesis technology, this talk is converted into the voice of an AI character. The input is the text data of the generated travel plan, and the output is an audio file. This audio data is then sent to the terminal.
[0507] Step 3:
[0508] The terminal plays audio data received from the server, providing real-time guidance to travelers. Simultaneously, users experience information within the virtual environment visually using visual augmentation technology. Specifically, this involves displaying information using smart glasses or head-mounted displays.
[0509] Step 4:
[0510] The user inputs questions and requests for additional information during the tour via voice. The input is sent from the terminal to the server. The server analyzes this information and uses a generated AI model to produce appropriate responses to the user's questions and requests. The output is voice data for the AI character to provide answers.
[0511] Step 5:
[0512] The terminal plays audio data generated by the server and provides responses to the user. Travelers can enjoy an interactive sightseeing experience through this process. The use of prompts in this sequence improves the accuracy of the generated AI model and user satisfaction.
[0513] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0514] This invention is a travel guidance system using an AI character that incorporates an emotion engine to recognize the traveler's emotions and improve the travel experience. The system mainly consists of a server, a user terminal, an AI character, and an emotion engine, and functions as follows:
[0515] 1. Server-based plan generation and sentiment analysis
[0516] Users input their travel destinations and interests into a device, and based on this, the server uses a generative model to create a customized travel plan. Furthermore, an emotion engine is used to analyze data obtained from the user's voice and facial expressions to determine the user's emotional state. The server uses this information to dynamically adjust the travel plan and guidance content, personalizing the user's travel experience.
[0517] 2. Speech synthesis and emotion-based real-time guidance
[0518] Based on the generated plan, the server creates guidance dialogue that reflects the output of the emotion engine and converts it into a character's voice using speech synthesis. The terminal receives this audio data and provides the user with real-time, emotion-responsive voice guidance. The voice guidance is delivered in a tone and speed that matches the user's emotions, resulting in a more natural and engaging guided experience.
[0519] 3. Emotional interaction with the user
[0520] Users interact with an AI character via their device while traveling. The server utilizes an emotion engine to analyze the user's voice and facial expressions, generating real-time responses based on their current emotional state. For example, if the user appears tired, it suggests rest stops; if they seem to be enjoying themselves, it recommends further activities.
[0521] Specific example
[0522] As a concrete example, imagine a user is a traveler interested in history who wants to tour historical landmarks in New York City. The user enters "I want to visit historic buildings in New York" into their device. The server uses a generative model and an emotion engine to create a plan that includes landmarks such as Central Park and the Flatiron Building. During the tour, the character will announce in real time, "Next is the Flatiron Building," and if the emotion engine recognizes that the user's expression shows surprise, it will add a more detailed explanation, such as, "This building was constructed in 1902 and is characterized by its triangular exterior." In this way, the guidance content is flexibly adjusted according to the user's emotional state, enhancing the quality of the trip.
[0523] This system aims to provide travelers with innovative and engaging guidance services by deeply understanding their emotions and individually optimizing their travel experience.
[0524] The following describes the processing flow.
[0525] Step 1:
[0526] The user uses their device to input their travel destination and topics of interest. The device then sends the entered information to the server.
[0527] Step 2:
[0528] The server receives user input and uses a generative model to create a travel plan tailored to the user's preferences. The server also references a database to collect tourist destination information.
[0529] Step 3:
[0530] The server starts the emotion engine and prepares to analyze the user's voice and facial expression data. The user grants permission for the device to access the camera and microphone.
[0531] Step 4:
[0532] The device captures the user's voice and facial expressions in real time and sends them to a server as emotion data. The data is processed while protecting the user's privacy.
[0533] Step 5:
[0534] The server's emotion engine analyzes the user's emotional state and dynamically adjusts the generated travel plan and information based on the results.
[0535] Step 6:
[0536] The server prepares a pre-arranged guided speech and converts it into a character's voice using speech synthesis technology. The generated audio data is then sent to the terminal.
[0537] Step 7:
[0538] The device receives audio data from the server and provides real-time, emotion-responsive voice guidance to the user. The guidance content is tailored to the user's emotions.
[0539] Step 8:
[0540] The user enters additional questions or requests into the device during the guidance process. The device then sends the interaction data to the server.
[0541] Step 9:
[0542] The server processes additional user interactions and generates an appropriate response, taking into account the results of the emotion engine. The response is then converted into speech and sent to the terminal.
[0543] Step 10:
[0544] The device provides the user with generated responses and offers continuous guidance to enrich the travel experience.
[0545] (Example 2)
[0546] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0547] Traditional travel guidance systems have struggled to individually optimize the travel experience based on travelers' emotions and real-time interests. In particular, the one-way nature of voice guidance is problematic, as it lacks the flexibility to respond to users' immediate emotional changes. Furthermore, personalization using information specific to individual travelers, such as past travel history, is insufficient.
[0548] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0549] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for providing a speech synthesis device for a virtual character to provide voice guidance, and means for using an emotion analysis device to analyze the user's emotional state. This makes it possible to dynamically adjust guidance according to the traveler's emotions and provide a personalized travel experience.
[0550] A "travel plan" is a series of guided information that suggests tourist attractions and activities based on the traveler's destination and interests.
[0551] A "generative model" is an artificial intelligence technique used to generate new information from given inputs, based on a large dataset.
[0552] A "virtual character" is a computer-generated character, whether human or animal, that interacts with or guides the user.
[0553] A "speech synthesis device" is a device that has the function of converting text data into speech data and providing information to the user as speech.
[0554] "User's emotional state" refers to the user's psychological or emotional condition at a given time, analyzed based on data such as the user's voice and facial expressions.
[0555] An "emotion analysis device" is a device used to analyze a user's facial expressions and voice, and to estimate their emotions from that information.
[0556] A "dynamic adjustment mechanism" refers to a system that can change the content of guidance and the services provided in response to real-time changes in the situation.
[0557] A "personalized travel experience" refers to a personalized travel experience that is optimized for the individual interests and emotional state of the traveler.
[0558] This invention is an interactive travel guidance system that analyzes travelers' emotional states in real time and personalizes their travel experience. This system primarily consists of three components: a server, a terminal, and a user. The following describes how each component functions.
[0559] The server uses a generative AI model to generate a customized travel plan based on destination information and interests sent from the user via their device. The generative AI model used is based on a large dataset, analyzes text-based input, and provides a new travel plan in response to the prompt "Provide a prompt for the user to enter places they want to visit and generate a travel plan."
[0560] The device is equipped with a camera and microphone to record the user's voice and facial expressions. This data is transmitted to a server in real time, and the server estimates the user's emotional state through an emotion analysis device. This analysis uses software incorporating existing emotion analysis algorithms to evaluate the user's psychological state, such as whether they are enjoying themselves or feeling tired.
[0561] Users interact with an AI character using a smartphone or other mobile device. This character plays back guidance speech generated on a server using a speech synthesis device. The speech synthesis uses common software for converting text data into speech, and the tone and speed are adjusted according to the user's emotional state.
[0562] For example, if a user enters "I want to visit a quiet tourist spot" into their device, the server uses a generative AI model to generate a plan suggesting quiet tourist spots. If the user appears tired during their visit, the system will provide voice guidance such as, "There's a cafe nearby. Why don't you take a break?" In this way, by providing personalized travel guidance based on the user's emotional state in real time, the quality of the travel experience can be improved.
[0563] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0564] Step 1:
[0565] The user enters their travel destination and interests into the device. The device formats this data to send it to the server. This input includes specific keywords related to the places they want to visit and their interests. The device formats this text data and generates text packets to send to the server.
[0566] Step 2:
[0567] The server generates a travel plan using a generative AI model based on destination information and interests received from the terminal. The server inputs a prompt message into the AI model: "Provide a prompt that allows the user to input places they want to visit and generate a travel plan." The AI model analyzes this input and outputs a travel plan that meets the user's needs. This output includes information on tourist spots to visit and recommended routes.
[0568] Step 3:
[0569] When a user begins interacting with the device, the device uses its camera and microphone to collect facial and audio data from the user. This data is used to analyze the user's emotional state. The device collects this raw data, converts it into a format suitable for emotion analysis, and sends it to the server.
[0570] Step 4:
[0571] The server passes the received facial and audio data to an emotion analysis device to analyze the user's emotional state. The analysis engine evaluates whether the user is surprised, amused, or tired based on factors such as voice tone, volume, and changes in facial expression. Based on this, it outputs parameters indicating the user's psychological state.
[0572] Step 5:
[0573] The server dynamically adjusts the guidance content using the generated travel plan and the results of the emotion analysis. The server generates guidance talk tailored to the user's emotional state and converts it into audio data using a speech synthesizer. For example, if the analysis indicates that the user is surprised, it incorporates explanations with more detailed background information. Audio data based on this dynamic adjustment is output.
[0574] Step 6:
[0575] The device receives audio data from the server and plays audio guidance to the user in real time. The device applies settings to play the audio at a tone and speed that matches the user's emotional state, providing guidance that naturally appeals to both sight and hearing. The played audio data is delivered to the user, complementing their travel experience.
[0576] (Application Example 2)
[0577] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0578] Conventional travel guidance systems often provide uniform plans and information without considering the emotional state of travelers, making it difficult to enhance the experience according to the individual emotions and preferences of travelers. Furthermore, when using autonomous vehicles as a means of transportation, dynamic adjustment of guidance content in response to the emotions of travelers during operation is required, but there is a lack of technology to achieve this.
[0579] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0580] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests; means for a proxy to provide voice guidance based on the generated travel plan; means for playing the voice guidance on the user's device and providing the traveler with real-time tourist information; means for analyzing the passenger's emotional state and dynamically adjusting the travel plan and guidance content using that information; and means for receiving feedback and questions from the user and generating appropriate responses using the generative model. This makes it possible to provide a personalized travel experience that responds to the traveler's emotions.
[0581] A "travel plan" is a detailed itinerary that outlines the places to visit and activities to do, based on the traveler's destination information and interests.
[0582] A "generative model" refers to an algorithm or data processing system that calculates, selects, and proposes the optimal travel plan based on the traveler's input information.
[0583] "Speech synthesis means" refers to technologies and devices for converting text information into speech and outputting it.
[0584] "Emotional analysis methods" are technologies that use data such as a traveler's voice and video to determine their emotional state at a given time.
[0585] "Dynamic adjustment" refers to the process of changing the itinerary and guidance content in real time according to the situation.
[0586] "Personalization" refers to optimizing the experience according to each traveler's individual characteristics, preferences, and emotions.
[0587] To implement this invention, a system is needed that generates travel plans based on travelers' destination information and interests. The server employs a generation AI model to create a travel plan based on this information and automatically suggests optimal destinations and activities. This allows for the provision of personalized travel plans for each traveler.
[0588] The server generates voice guidance based on a plan created using speech synthesis technology and transmits it to the user's mobile device or the navigation system of the autonomous vehicle. The device plays the voice guidance received from the server as is, providing the traveler with real-time tourist information. This voice guidance is adjusted in real time, taking into account the traveler's current emotional state.
[0589] Furthermore, cameras and microphones installed inside the train cars are used to analyze passengers' emotions. Through these input devices, the emotion analysis system analyzes the passengers' emotional state and sends the data to a server. The server uses this information to dynamically adjust the travel plan and information provided.
[0590] For example, if a traveler's emotions are analyzed while visiting a tourist attraction, and they appear tired, the server might suggest a place to rest as their next destination. Conversely, if they appear excited, it could suggest more active activities.
[0591] Example of a prompt:
[0592] "Are the passengers excited? Please generate information about the next tourist attraction accordingly."
[0593] "If the passengers are tired, please suggest a route that allows them to take a break."
[0594] These features allow the system to provide an innovative travel experience that incorporates the traveler's emotions.
[0595] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0596] Step 1:
[0597] The server receives the traveler's destination information and interests. User information is sent via the terminal as input. The server uses this input data to generate prompts for the AI model and constructs a travel plan based on that information. Through this process, suggestions tailored to the traveler's interests are formed.
[0598] Step 2:
[0599] The terminal receives travel plan information sent from the server. The terminal passes this information to a speech synthesis system, which converts the text information into audio data to provide guidance to the traveler. The input is text-based guidance information from the server, and the output is generated audio guidance. In this step, acoustic data processing is performed with the aim of conveying the guide content in a natural tone and speaking speed.
[0600] Step 3:
[0601] The server analyzes the user's emotions in real time via cameras and microphones inside the vehicle. Input includes audio and video data from the emotion analysis system, and the server processes this data to infer the emotional state. The output is the inferred emotional state. This information is used in the next step to adjust the travel plan.
[0602] Step 4:
[0603] The server dynamically adjusts travel plans and guidance based on the emotion analysis results. The input is generated emotion state data, which is used to create new guidance and routes optimized for the traveler's emotions. The output is an updated travel plan, which is then sent back to the terminal for speech synthesis. This process personalizes the traveler's experience.
[0604] Step 5:
[0605] The user sends feedback and additional questions to the server via their device. Based on this input, the server uses a generative AI model to generate an appropriate response and sends it back to the device. The output here is a customized answer or suggestion to the user's inquiry, which is then provided to the traveler again as voice guidance.
[0606] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0607] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0608] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0609] [Fourth Embodiment]
[0610] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0611] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0612] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0613] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0614] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0615] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0616] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0617] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0618] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0619] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0620] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0621] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0622] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0623] This invention is a system that provides professional travel guidance using an AI character, generating personalized travel plans based on the traveler's destination and interests, and providing real-time voice guidance. The system mainly consists of a server, a user terminal, and an AI character, and functions as follows:
[0624] 1. Server-driven plan generation
[0625] The server uses a generative model to create customized travel plans based on destination and interest information received from the user. This generative model has learned from a large amount of tourist destination data and guide expertise, and makes suggestions that take into account the optimal order of visits and modes of transportation. The server also reflects the user's past travel history and preferences to personalize the plan.
[0626] 2. Speech synthesis and real-time guidance
[0627] Based on a travel plan generated on the server, talk segments about each tourist destination in the plan are created and converted into a character's voice using speech synthesis technology. The user terminal receives this voice data and provides real-time voice guidance during the trip. Speech synthesis technology enables natural-sounding guidance at visited locations, supporting a smooth traveler experience.
[0628] 3. User Interaction
[0629] Users can interact with an AI character through their device to get answers to questions that arise during their sightseeing and to request additional information. The server analyzes the user's questions and requests and returns appropriate responses in real time using a generative model. This allows the character to support travelers as if they were a real-life guide.
[0630] Specific example
[0631] For example, consider a user who is a traveler interested in Japanese culture and is visiting Kyoto. The user sends a request from their device saying, "I want to visit traditional temples in Kyoto." The server uses a generative model to generate a travel plan that includes representative locations such as Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. During the trip, the user's device provides voice guidance in a character's voice, such as, "Next is Kinkaku-ji. Enjoy the iconic architectural style of the Muromachi period." If the user then asks, "I want to know more about the history," the server provides detailed historical background information based on that request.
[0632] In this way, this system utilizes AI technology to provide travelers with high-value tourism experiences and aims to improve the quality of their travels.
[0633] The following describes the processing flow.
[0634] Step 1:
[0635] The user enters their travel destination and areas of interest into the device. The device then sends the entered information to the server.
[0636] Step 2:
[0637] The server receives information from the user and starts a generative model to generate a travel plan. The server consults a database to collect information about relevant tourist destinations.
[0638] Step 3:
[0639] The server uses collected tourist destination information and a generative model to create a travel plan optimized for the user. The plan includes the order of visits, modes of transportation, and time allocation.
[0640] Step 4:
[0641] Based on the travel plan created by the server, the content of the guided conversation is constructed. The server then converts the conversation content into a character's voice using speech synthesis technology.
[0642] Step 5:
[0643] The device receives audio data sent from the server. The device then uses this data to begin providing real-time audio guidance to the user.
[0644] Step 6:
[0645] The user interacts with the character while traveling. The device forwards user questions and requests to the server via voice or text.
[0646] Step 7:
[0647] The server receives interaction from the user and generates an appropriate response using a generative model. The server converts the response into audio data and sends it to the terminal.
[0648] Step 8:
[0649] The device relays responses sent from the server to the user. The device continuously provides updated plans and information in real time as needed.
[0650] (Example 1)
[0651] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0652] Modern travelers tend to demand personalized information and real-time guidance, but traditional guide systems are unable to adequately meet these demands. Specifically, it has been difficult to provide personalized itineraries based on travelers' preferences and past travel history, as well as effective support including immediate responses to questions and requests.
[0653] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0654] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for using a speech synthesis device for a character to provide voice guidance based on the generated travel plan, and means for personalizing the travel plan based on the traveler's past visit history and preferences. This enables more personalized and real-time information and guidance for the traveler.
[0655] A "traveler" refers to an individual who visits a tourist destination or place of interest and seeks information and guidance regarding their itinerary.
[0656] "Destination information" refers to information about specific tourist destinations or cities that travelers wish to visit.
[0657] A "generative model" refers to methods and algorithms, including artificial intelligence technology, used to create travel plans, which make suggestions based on tourist destination data and guide expertise.
[0658] A "speech synthesis device" refers to a technology that converts text data into speech data, enabling information to be transmitted using a human-like voice.
[0659] "Personalization" refers to the process of adjusting services and information based on the individual traveler's preferences and past visit history.
[0660] "Real-time" refers to a state where information and services can be provided immediately, meaning that travelers can resolve their questions and requests on the spot.
[0661] This invention is a system for providing travelers with personalized travel plans and real-time voice guidance. The main components of the invention are a server, a user terminal, and a generative AI model.
[0662] The server analyzes destination information and interest prompts provided by travelers and generates personalized travel plans based on this analysis. Natural language processing technology is used for this analysis, and a generating AI model utilizes tourist data and guide expertise to create the optimal plan. For example, if a user sends the prompt "I want to visit traditional temples in Kyoto," the server will generate a plan including Kinkaku-ji, Ginkaku-ji, and Kiyomizu-dera. In doing so, it also considers past travel history and preferences to provide the most fulfilling travel experience for the user.
[0663] The device provides travelers with audio data based on their travel plan, received from the server. Utilizing speech synthesis technology, an AI character delivers information about each tourist spot as if it were a real-life guide. For example, as a user is on their way to Kinkaku-ji Temple, the device will play an audio announcement saying, "Next stop is Kinkaku-ji Temple. Please enjoy this iconic architectural style from the Muromachi period."
[0664] Users can interact with an AI character via their device during their trip, getting their questions answered and seeking more in-depth information. The server analyzes the received requests in real time, uses a generative AI model to quickly generate appropriate responses, and provides them to the user. This process allows travelers to gain a more valuable experience than they could have planned on their own.
[0665] Thus, this invention uses AI technology to improve the quality of travel and ensure consistency and convenience in providing information to travelers.
[0666] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0667] Step 1:
[0668] The user uses the terminal to enter prompts about their desired destination or interests. Specifically, if the user enters "I want to visit traditional temples in Kyoto," a prompt will be generated. This is sent to the server and becomes input data for the next processing step.
[0669] Step 2:
[0670] The server parses the received prompt message. This parsing uses natural language processing to understand the user's intent and desired destinations. Based on the data analysis, the server utilizes a generative AI model to generate a personalized travel plan from tourist destination data and guide expertise. The generated travel plan becomes the output data used in the next step.
[0671] Step 3:
[0672] Based on the travel plan generated on the server, audio guides for each tourist destination are created. The server uses speech synthesis technology to convert explanations about the places visited in the plan into natural-sounding speech. This audio data is generated and becomes the output that is sent to the user's terminal.
[0673] Step 4:
[0674] The device plays audio data received from the server, providing real-time guidance to travelers. When a user visits Kinkaku-ji Temple, the device plays an audio announcement such as, "Next stop is Kinkaku-ji Temple. Please enjoy the iconic architectural style of the Muromachi period." This allows users to experience a realistic guided tour.
[0675] Step 5:
[0676] Users can request additional information or ask questions during their trip. These requests are sent to the server via their device. Specifically, the user might say something like, "I want to know more about the history of this temple," and that request becomes new input data.
[0677] Step 6:
[0678] The server analyzes user requests in real time and uses a generative AI model to create appropriate responses. Based on the analysis results, the newly created response is generated as audio data and sent to the terminal. This allows the user to instantly obtain the information they need.
[0679] (Application Example 1)
[0680] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0681] Traditional travel information systems have struggled to provide personalized travel plans for individual travelers, and the quality of real-time guidance and interaction has been limited. Furthermore, the lack of means to offer virtual sightseeing experiences before actual travel meant that users were not adequately provided with engaging experiences during the travel planning stage. This leaves challenges remaining, such as providing flexible guidance tailored to travelers' needs and increasing travel motivation through simulated experiences before the trip.
[0682] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0683] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for voice synthesis to have a character provide voice guidance based on the generated travel plan, and means for providing guidance in a virtual environment to the user via visual augmentation technology, and for presenting guidance related to each scene in audio and visual form. This enables the provision of personalized travel plans that meet the individual needs of travelers, as well as real-time and interactive sightseeing guidance in a virtual environment.
[0684] The term "traveler" refers to individuals or groups who visit a specific destination and travel for various purposes, such as tourism or business.
[0685] "Destination information" refers to information about the geographical location and related tourist resources that travelers intend to visit, and includes specific place names and tourist attractions.
[0686] A "generative model" refers to an algorithm or technology used to automatically create travel plans based on input information. It is a system that learns from large amounts of data and patterns to generate responses.
[0687] "Speech synthesis means" refers to technology that converts text data into speech signals, and includes devices and software for artificially generating and providing speech to users.
[0688] "Visual terminals" refer to devices that allow users to receive visual information, and include smart glasses and head-mounted displays.
[0689] "Feedback" refers to the opinions, evaluations, and questions that users provide to a system, and is used to improve the system and generate responses.
[0690] "Visual augmentation technology" refers to the technology of overlaying virtual information onto the real world, and is a means of improving the user's experience of the real world.
[0691] A "virtual environment" refers to a virtual space or situation created using computer technology, allowing users to experience a space that does not exist in reality.
[0692] "Interactive tourist information" refers to a form of information gathering where travelers obtain information through two-way communication with the system, and it is a dynamic guidance method that changes its response according to the user's input.
[0693] This invention is a system that provides personalized travel guidance for travelers, and mainly consists of a server, a terminal, and user interaction.
[0694] The server uses a generative model to generate travel plans based on the traveler's destination information and interests. The generative model learns from a large amount of tourist destination data and guide expertise, and provides plans that consider the optimal order of visits and means of transportation. Based on the generated travel plan, the server uses speech synthesis to create talk about each tourist destination in the plan, converts it into the voice of an AI character, and sends it to the terminal.
[0695] The device plays back received audio data and provides travelers with real-time tourist information. It also enables travelers to experience visual information within the virtual environment through visual augmentation technology. Smart glasses and head-mounted displays are used in this process.
[0696] Users can interact with AI characters through these devices. Through this interaction, users can request feedback or additional information, which is then sent from the device to the server. The server analyzes the received feedback and reuses the generative model to generate an appropriate response.
[0697] For example, if a user requests to "look for summer clothes," the server creates a virtual tour plan of stores that carry the latest summer clothing. Based on this, the device will guide the user by saying, "This is a store showcasing the latest summer clothing collection."
[0698] An example of a prompt in a generative AI model is: "The user wants to buy summer clothes. Introduce stores that feature the latest trends and generate an audio guide."
[0699] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0700] Step 1:
[0701] The user uses a terminal to input travel destination information and interests. This information is sent to a server. The server analyzes the data and uses a generative AI model to create a customized travel plan. The user's past travel history and preferences are also considered to determine the optimal order of visits and modes of transportation. The output is an optimized travel plan.
[0702] Step 2:
[0703] The server creates guided tours for each tourist spot based on the generated travel plan. Using speech synthesis technology, this talk is converted into the voice of an AI character. The input is the text data of the generated travel plan, and the output is an audio file. This audio data is then sent to the terminal.
[0704] Step 3:
[0705] The terminal plays audio data received from the server, providing real-time guidance to travelers. Simultaneously, users experience information within the virtual environment visually using visual augmentation technology. Specifically, this involves displaying information using smart glasses or head-mounted displays.
[0706] Step 4:
[0707] The user inputs questions and requests for additional information during the tour via voice. The input is sent from the terminal to the server. The server analyzes this information and uses a generated AI model to produce appropriate responses to the user's questions and requests. The output is voice data for the AI character to provide answers.
[0708] Step 5:
[0709] The terminal plays audio data generated by the server and provides responses to the user. Travelers can enjoy an interactive sightseeing experience through this process. The use of prompts in this sequence improves the accuracy of the generated AI model and user satisfaction.
[0710] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0711] This invention is a travel guidance system using an AI character that incorporates an emotion engine to recognize the traveler's emotions and improve the travel experience. The system mainly consists of a server, a user terminal, an AI character, and an emotion engine, and functions as follows:
[0712] 1. Server-based plan generation and sentiment analysis
[0713] Users input their travel destinations and interests into a device, and based on this, the server uses a generative model to create a customized travel plan. Furthermore, an emotion engine is used to analyze data obtained from the user's voice and facial expressions to determine the user's emotional state. The server uses this information to dynamically adjust the travel plan and guidance content, personalizing the user's travel experience.
[0714] 2. Speech synthesis and emotion-based real-time guidance
[0715] Based on the generated plan, the server creates guidance dialogue that reflects the output of the emotion engine and converts it into a character's voice using speech synthesis. The terminal receives this audio data and provides the user with real-time, emotion-responsive voice guidance. The voice guidance is delivered in a tone and speed that matches the user's emotions, resulting in a more natural and engaging guided experience.
[0716] 3. Emotional interaction with the user
[0717] Users interact with an AI character via their device while traveling. The server utilizes an emotion engine to analyze the user's voice and facial expressions, generating real-time responses based on their current emotional state. For example, if the user appears tired, it suggests rest stops; if they seem to be enjoying themselves, it recommends further activities.
[0718] Specific example
[0719] As a concrete example, imagine a user is a traveler interested in history who wants to tour historical landmarks in New York City. The user enters "I want to visit historic buildings in New York" into their device. The server uses a generative model and an emotion engine to create a plan that includes landmarks such as Central Park and the Flatiron Building. During the tour, the character will announce in real time, "Next is the Flatiron Building," and if the emotion engine recognizes that the user's expression shows surprise, it will add a more detailed explanation, such as, "This building was constructed in 1902 and is characterized by its triangular exterior." In this way, the guidance content is flexibly adjusted according to the user's emotional state, enhancing the quality of the trip.
[0720] This system aims to provide travelers with innovative and engaging guidance services by deeply understanding their emotions and individually optimizing their travel experience.
[0721] The following describes the processing flow.
[0722] Step 1:
[0723] The user uses their device to input their travel destination and topics of interest. The device then sends the entered information to the server.
[0724] Step 2:
[0725] The server receives user input and uses a generative model to create a travel plan tailored to the user's preferences. The server also references a database to collect tourist destination information.
[0726] Step 3:
[0727] The server starts the emotion engine and prepares to analyze the user's voice and facial expression data. The user grants permission for the device to access the camera and microphone.
[0728] Step 4:
[0729] The device captures the user's voice and facial expressions in real time and sends them to a server as emotion data. The data is processed while protecting the user's privacy.
[0730] Step 5:
[0731] The server's emotion engine analyzes the user's emotional state and dynamically adjusts the generated travel plan and information based on the results.
[0732] Step 6:
[0733] The server prepares a pre-arranged guided speech and converts it into a character's voice using speech synthesis technology. The generated audio data is then sent to the terminal.
[0734] Step 7:
[0735] The device receives audio data from the server and provides real-time, emotion-responsive voice guidance to the user. The guidance content is tailored to the user's emotions.
[0736] Step 8:
[0737] The user enters additional questions or requests into the device during the guidance process. The device then sends the interaction data to the server.
[0738] Step 9:
[0739] The server processes additional user interactions and generates an appropriate response, taking into account the results of the emotion engine. The response is then converted into speech and sent to the terminal.
[0740] Step 10:
[0741] The device provides the user with generated responses and offers continuous guidance to enrich the travel experience.
[0742] (Example 2)
[0743] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] Traditional travel guidance systems have struggled to individually optimize the travel experience based on travelers' emotions and real-time interests. In particular, the one-way nature of voice guidance is problematic, as it lacks the flexibility to respond to users' immediate emotional changes. Furthermore, personalization using information specific to individual travelers, such as past travel history, is insufficient.
[0745] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0746] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests, means for providing a speech synthesis device for a virtual character to provide voice guidance, and means for using an emotion analysis device to analyze the user's emotional state. This makes it possible to dynamically adjust guidance according to the traveler's emotions and provide a personalized travel experience.
[0747] A "travel plan" is a series of guided information that suggests tourist attractions and activities based on the traveler's destination and interests.
[0748] A "generative model" is an artificial intelligence technique used to generate new information from given inputs, based on a large dataset.
[0749] A "virtual character" is a computer-generated character, whether human or animal, that interacts with or guides the user.
[0750] A "speech synthesis device" is a device that has the function of converting text data into speech data and providing information to the user as speech.
[0751] "User's emotional state" refers to the user's psychological or emotional condition at a given time, analyzed based on data such as the user's voice and facial expressions.
[0752] An "emotion analysis device" is a device used to analyze a user's facial expressions and voice, and to estimate their emotions from that information.
[0753] A "dynamic adjustment mechanism" refers to a system that can change the content of guidance and the services provided in response to real-time changes in the situation.
[0754] A "personalized travel experience" refers to a personalized travel experience that is optimized for the individual interests and emotional state of the traveler.
[0755] This invention is an interactive travel guidance system that analyzes travelers' emotional states in real time and personalizes their travel experience. This system primarily consists of three components: a server, a terminal, and a user. The following describes how each component functions.
[0756] The server uses a generative AI model to generate a customized travel plan based on destination information and interests sent from the user via their device. The generative AI model used is based on a large dataset, analyzes text-based input, and provides a new travel plan in response to the prompt "Provide a prompt for the user to enter places they want to visit and generate a travel plan."
[0757] The device is equipped with a camera and microphone to record the user's voice and facial expressions. This data is transmitted to a server in real time, and the server estimates the user's emotional state through an emotion analysis device. This analysis uses software incorporating existing emotion analysis algorithms to evaluate the user's psychological state, such as whether they are enjoying themselves or feeling tired.
[0758] Users interact with an AI character using a smartphone or other mobile device. This character plays back guidance speech generated on a server using a speech synthesis device. The speech synthesis uses common software for converting text data into speech, and the tone and speed are adjusted according to the user's emotional state.
[0759] For example, if a user enters "I want to visit a quiet tourist spot" into their device, the server uses a generative AI model to generate a plan suggesting quiet tourist spots. If the user appears tired during their visit, the system will provide voice guidance such as, "There's a cafe nearby. Why don't you take a break?" In this way, by providing personalized travel guidance based on the user's emotional state in real time, the quality of the travel experience can be improved.
[0760] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0761] Step 1:
[0762] The user enters their travel destination and interests into the device. The device formats this data to send it to the server. This input includes specific keywords related to the places they want to visit and their interests. The device formats this text data and generates text packets to send to the server.
[0763] Step 2:
[0764] The server generates a travel plan using a generative AI model based on destination information and interests received from the terminal. The server inputs a prompt message into the AI model: "Provide a prompt that allows the user to input places they want to visit and generate a travel plan." The AI model analyzes this input and outputs a travel plan that meets the user's needs. This output includes information on tourist spots to visit and recommended routes.
[0765] Step 3:
[0766] When a user begins interacting with the device, the device uses its camera and microphone to collect facial and audio data from the user. This data is used to analyze the user's emotional state. The device collects this raw data, converts it into a format suitable for emotion analysis, and sends it to the server.
[0767] Step 4:
[0768] The server passes the received facial and audio data to an emotion analysis device to analyze the user's emotional state. The analysis engine evaluates whether the user is surprised, amused, or tired based on factors such as voice tone, volume, and changes in facial expression. Based on this, it outputs parameters indicating the user's psychological state.
[0769] Step 5:
[0770] The server dynamically adjusts the guidance content using the generated travel plan and the results of the emotion analysis. The server generates guidance talk tailored to the user's emotional state and converts it into audio data using a speech synthesizer. For example, if the analysis indicates that the user is surprised, it incorporates explanations with more detailed background information. Audio data based on this dynamic adjustment is output.
[0771] Step 6:
[0772] The device receives audio data from the server and plays audio guidance to the user in real time. The device applies settings to play the audio at a tone and speed that matches the user's emotional state, providing guidance that naturally appeals to both sight and hearing. The played audio data is delivered to the user, complementing their travel experience.
[0773] (Application Example 2)
[0774] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0775] Conventional travel guidance systems often provide uniform plans and information without considering the emotional state of travelers, making it difficult to enhance the experience according to the individual emotions and preferences of travelers. Furthermore, when using autonomous vehicles as a means of transportation, dynamic adjustment of guidance content in response to the emotions of travelers during operation is required, but there is a lack of technology to achieve this.
[0776] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0777] In this invention, the server includes means for using a generative model to generate a travel plan based on the traveler's destination information and interests; means for a proxy to provide voice guidance based on the generated travel plan; means for playing the voice guidance on the user's device and providing the traveler with real-time tourist information; means for analyzing the passenger's emotional state and dynamically adjusting the travel plan and guidance content using that information; and means for receiving feedback and questions from the user and generating appropriate responses using the generative model. This makes it possible to provide a personalized travel experience that responds to the traveler's emotions.
[0778] A "travel plan" is a detailed itinerary that outlines the places to visit and activities to do, based on the traveler's destination information and interests.
[0779] A "generative model" refers to an algorithm or data processing system that calculates, selects, and proposes the optimal travel plan based on the traveler's input information.
[0780] "Speech synthesis means" refers to technologies and devices for converting text information into speech and outputting it.
[0781] "Emotional analysis methods" are technologies that use data such as a traveler's voice and video to determine their emotional state at a given time.
[0782] "Dynamic adjustment" refers to the process of changing the itinerary and guidance content in real time according to the situation.
[0783] "Personalization" refers to optimizing the experience according to each traveler's individual characteristics, preferences, and emotions.
[0784] To implement this invention, a system is needed that generates travel plans based on travelers' destination information and interests. The server employs a generation AI model to create a travel plan based on this information and automatically suggests optimal destinations and activities. This allows for the provision of personalized travel plans for each traveler.
[0785] The server generates voice guidance based on a plan created using speech synthesis technology and transmits it to the user's mobile device or the navigation system of the autonomous vehicle. The device plays the voice guidance received from the server as is, providing the traveler with real-time tourist information. This voice guidance is adjusted in real time, taking into account the traveler's current emotional state.
[0786] Furthermore, cameras and microphones installed inside the train cars are used to analyze passengers' emotions. Through these input devices, the emotion analysis system analyzes the passengers' emotional state and sends the data to a server. The server uses this information to dynamically adjust the travel plan and information provided.
[0787] For example, if a traveler's emotions are analyzed while visiting a tourist attraction, and they appear tired, the server might suggest a place to rest as their next destination. Conversely, if they appear excited, it could suggest more active activities.
[0788] Example of a prompt:
[0789] "Are the passengers excited? Please generate information about the next tourist attraction accordingly."
[0790] "If the passengers are tired, please suggest a route that allows them to take a break."
[0791] These features allow the system to provide an innovative travel experience that incorporates the traveler's emotions.
[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0793] Step 1:
[0794] The server receives the traveler's destination information and interests. User information is sent via the terminal as input. The server uses this input data to generate prompts for the AI model and constructs a travel plan based on that information. Through this process, suggestions tailored to the traveler's interests are formed.
[0795] Step 2:
[0796] The terminal receives travel plan information sent from the server. The terminal passes this information to a speech synthesis system, which converts the text information into audio data to provide guidance to the traveler. The input is text-based guidance information from the server, and the output is generated audio guidance. In this step, acoustic data processing is performed with the aim of conveying the guide content in a natural tone and speaking speed.
[0797] Step 3:
[0798] The server analyzes the user's emotions in real time via cameras and microphones inside the vehicle. Input includes audio and video data from the emotion analysis system, and the server processes this data to infer the emotional state. The output is the inferred emotional state. This information is used in the next step to adjust the travel plan.
[0799] Step 4:
[0800] The server dynamically adjusts travel plans and guidance based on the emotion analysis results. The input is generated emotion state data, which is used to create new guidance and routes optimized for the traveler's emotions. The output is an updated travel plan, which is then sent back to the terminal for speech synthesis. This process personalizes the traveler's experience.
[0801] Step 5:
[0802] The user sends feedback and additional questions to the server via their device. Based on this input, the server uses a generative AI model to generate an appropriate response and sends it back to the device. The output here is a customized answer or suggestion to the user's inquiry, which is then provided to the traveler again as voice guidance.
[0803] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0804] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0805] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0806] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0807] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0808] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0809] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0810] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0811] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0812] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0813] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0814] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0815] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0816] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0817] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0818] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0819] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0820] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0821] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0822] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0823] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0824] The following is further disclosed regarding the embodiments described above.
[0825] (Claim 1)
[0826] A means of using a generative model to generate travel plans based on travelers' destination information and interests,
[0827] A voice synthesis means for a character to provide voice guidance based on a generated travel plan,
[0828] A means of providing real-time tourist information to travelers by playing audio guidance on the user's device,
[0829] A means of receiving user feedback and questions and generating appropriate responses using a generative model,
[0830] A system that includes this.
[0831] (Claim 2)
[0832] The system according to claim 1, comprising means for dynamically adjusting the guidance content based on the travel plan according to the traveler's location information.
[0833] (Claim 3)
[0834] The system according to claim 1, further comprising means for personalizing a traveler's travel plan based on their past travel history and preferences.
[0835] "Example 1"
[0836] (Claim 1)
[0837] A device that uses a generative model to generate travel plans based on travelers' destination information and interests,
[0838] A voice synthesis device for a character to provide an audio guide based on the generated travel plan,
[0839] A device that plays audio guides on the user's information processing device and provides travelers with real-time information about tourist destinations,
[0840] A device that receives user feedback and questions and generates appropriate responses using a generative model,
[0841] A device that personalizes travel plans based on travelers' past visit history and preferences,
[0842] A system that includes this.
[0843] (Claim 2)
[0844] The system according to claim 1, comprising a device that dynamically adjusts the guide content based on the travel plan according to the traveler's current location information.
[0845] (Claim 3)
[0846] The system according to claim 1, further comprising a device for providing real-time information based on interaction between a user and a character.
[0847] "Application Example 1"
[0848] (Claim 1)
[0849] A means of using a generative model to generate travel plans based on travelers' destination information and interests,
[0850] A voice synthesis means for a character to provide voice guidance based on a generated travel plan,
[0851] A means of providing real-time tourist information to travelers by playing audio guidance on the user's visual device,
[0852] A means of receiving user feedback and questions and generating appropriate responses using a generative model,
[0853] A means of providing users with guidance in a virtual environment through visual augmentation technology, and presenting guidance related to each scene in audio and visual form,
[0854] A system that includes this.
[0855] (Claim 2)
[0856] The system according to claim 1, comprising means for dynamically adjusting the guidance content based on the travel plan according to the traveler's location information.
[0857] (Claim 3)
[0858] The system according to claim 1, further comprising means for personalizing a traveler's travel plan based on their past travel history and preferences.
[0859] "Example 2 of combining an emotion engine"
[0860] (Claim 1)
[0861] A means of using a generative model to generate travel plans based on travelers' destination information and interests,
[0862] A means comprising a voice synthesis device for a virtual character to provide voice guidance based on a generated travel plan,
[0863] A means of providing real-time tourist information to travelers by playing audio guidance on the user's device,
[0864] A means of using an emotion analysis device to analyze the emotional state of a user,
[0865] A dynamic adjustment mechanism for adjusting the guidance content based on emotional state,
[0866] A system that includes this.
[0867] (Claim 2)
[0868] The system according to claim 1, further comprising means for dynamically adjusting the guidance content according to the traveler's emotional state based on user emotion analysis.
[0869] (Claim 3)
[0870] The system according to claim 1, further comprising means for personalizing travel plans based on a traveler's past visit history and preferences.
[0871] "Application example 2 of combining emotional engines"
[0872] (Claim 1)
[0873] A means of using a generative model to generate travel plans based on travelers' destination information and interests,
[0874] A speech synthesis means for an agent to provide voice guidance based on a generated travel plan,
[0875] A means of providing real-time tourist information to travelers by playing audio guidance on the user's device,
[0876] An emotion analysis tool that analyzes the emotional state of passengers and uses that information to dynamically adjust travel plans and information content,
[0877] A means of receiving feedback and questions from users and generating appropriate responses using a generative model,
[0878] A system that includes this.
[0879] (Claim 2)
[0880] The system according to claim 1, comprising means for dynamically adjusting the content of the travel plan based on the traveler's geographical information.
[0881] (Claim 3)
[0882] The system according to claim 1, further comprising means for personalizing travel plans based on a traveler's past travel history and preferences. [Explanation of Symbols]
[0883] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of using a generative model to generate travel plans based on travelers' destination information and interests, A voice synthesis means for a character to provide voice guidance based on a generated travel plan, A means of providing real-time tourist information to travelers by playing audio guidance on the user's device, A means of receiving user feedback and questions and generating appropriate responses using a generative model, A system that includes this.
2. The system according to claim 1, comprising means for dynamically adjusting the guidance content based on the travel plan according to the traveler's location information.
3. The system according to claim 1, further comprising means for personalizing a traveler's travel plan based on their past travel history and preferences.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A