system

The system addresses language and cultural barriers for foreign travelers in Japan by offering personalized travel plans, multilingual translation, and emergency assistance, enhancing their experience and safety.

JP2026104471APending Publication Date: 2026-06-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-13
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Foreign travelers visiting Japan face challenges due to language barriers, cultural differences, and insufficient information about local customs, which hinder their ability to fully enjoy their trips and respond effectively to emergencies.

Method used

A system that supports travelers by creating personalized travel plans, providing multilingual translation, cultural information, and emergency assistance through a terminal, server, and user interaction, utilizing speech recognition, emotion analysis, and AI to enhance communication and safety.

Benefits of technology

Enables travelers to have a safe, fulfilling, and culturally enriching experience by overcoming language barriers and providing timely emergency support, ensuring smooth communication and personalized travel recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026104471000001_ABST
    Figure 2026104471000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of entering user information and creating a travel itinerary, Means for performing multilingual translation, Means of providing information on local rules and customs, A means of providing information on the nearest support organization in an emergency, A means of dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests, A system that includes a means of providing a quiz function about local culture and customs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] When foreign travelers visit Japan, there are problems that they cannot fully enjoy their trips due to language barriers, cultural differences, and insufficient information about local customs. Also, it is difficult to receive appropriate support for emergencies. The purpose is to solve such situations so that travelers can enjoy a safe and fulfilling travel experience.

Means for Solving the Problems

[0005] This invention provides a system that comprehensively supports the travel experience of foreign tourists within Japan. Specifically, it includes a means for creating travel plans based on the user's personal information and travel purpose, and supports smooth communication through a multilingual translation function. It also helps travelers avoid cultural misunderstandings by providing information on local rules and customs. Furthermore, in the event of an emergency, it ensures the safety of travelers by providing a means for quickly providing information on nearby support organizations.

[0006] "Foreign tourists" refers to individuals of other nationalities who visit and travel within Japan.

[0007] "Travel experience" refers to all the experiences that travelers gain at their destination through sightseeing, dining, cultural experiences, transportation, accommodation, and so on.

[0008] A "system" refers to a cooperative unit in which multiple functions or devices work together to achieve a specific purpose.

[0009] "User information" refers to personal data such as the traveler's name, age, nationality, language settings, and purpose of travel.

[0010] A "travel plan" refers to a plan that outlines destinations, planned activities, and other actions taken during a trip, based on the travel itinerary.

[0011] "Multilingual translation" refers to a language conversion function that enables communication between different languages.

[0012] "Local rules and customs" refers collectively to information concerning laws, social customs, and cultural understanding within Japan.

[0013] An "emergency situation" refers to a situation in which a traveler unexpectedly requires urgent assistance, and includes accidents, illnesses, and disasters.

[0014] "Support agencies" refer to police, hospitals, and other public assistance agencies that travelers can access in case of an emergency.

Brief Description of the Drawings

[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Modes for Carrying Out the Invention

[0016] Next, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention provides a system to support the travel experience of foreign tourists visiting Japan. This system consists of three main components: a terminal, a server, and a user.

[0037] The terminal is the traveler's smartphone or other device, through which the traveler interacts with the system. The user uses the terminal to register personal information before the start of their trip, including information such as name, nationality, language spoken, and areas of interest.

[0038] The server receives information sent by users and stores it in a database. This lays the foundation for providing personalized travel plans tailored to travelers. The server has the function of suggesting tourist spots and activities based on the traveler's interests and information about the area they are staying in.

[0039] Furthermore, the device provides a multilingual translation function to help users communicate more smoothly in the local area. Users can use the translation function through the device to convert voice and text into other languages. In this process, speech recognition technology is used to process the input, and the server generates the translated data and sends it back to the device.

[0040] Regarding explanations of local rules, when a user asks a question about local regulations or customs, the server looks up information in the database and provides the device with the appropriate information. This helps avoid cultural misunderstandings and ensures that local manners are observed.

[0041] Furthermore, in an emergency, the user can activate the "emergency assistance" function, and the device immediately sends its location information to a server. Based on that location, the server quickly provides information on the nearest police station, hospital, or other appropriate assistance organization. This information is then presented to the user through the device.

[0042] As a concrete example, consider a scenario where a user visits a major Japanese city and creates a plan to maximize their local cultural experience. Before starting their trip, the user inputs their areas of interest and activities through the app. The server analyzes this input and provides recommended plans, such as temple tours in Kyoto or local cuisine experiences in Tokyo. This allows the user to make the most of their time during their stay.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user launches the application on their device and opens the "New Registration" screen. They enter personal information such as their name, nationality, language settings, and areas of interest, and then press the submit button.

[0046] Step 2:

[0047] The device collects data entered by the user, encrypts it, and then sends it to the server.

[0048] Step 3:

[0049] The server analyzes the received user data and stores it in a database. This prepares the server to create travel plans based on the user's individual needs.

[0050] Step 4:

[0051] The user selects the "Create a travel plan" option and enters their interests and desired tourist destinations.

[0052] Step 5:

[0053] The device sends the user's requested information to the server.

[0054] Step 6:

[0055] The server uses AI to suggest optimal tourist spots and activities based on the user's interests and revised plan. These suggestions are then sent to the user's device.

[0056] Step 7:

[0057] The terminal displays suggestions sent from the server to the user, allowing the user to review the details and make modifications as needed.

[0058] Step 8:

[0059] If a user needs to translate a conversation while on location, they can use the app's translation function.

[0060] Step 9:

[0061] The device receives voice input and converts the speech into text.

[0062] Step 10:

[0063] Text data is sent to the server and translated into the specified language.

[0064] Step 11:

[0065] The translated text is sent back to the terminal from the server and displayed or outputted to the user as audio.

[0066] Step 12:

[0067] When a user seeks information about local rules and customs, they enter relevant questions into the app.

[0068] Step 13:

[0069] The terminal sends the entered question to the server.

[0070] Step 14:

[0071] The server searches the database and sends instructions for the appropriate local rules to the terminal.

[0072] Step 15:

[0073] The terminal displays guidance information sent from the server to the user.

[0074] Step 16:

[0075] In an emergency, the user presses the "Emergency Assistance" button.

[0076] Step 17:

[0077] The device sends data, including its current location, to the server.

[0078] Step 18:

[0079] The server analyzes the location information to identify the nearest police station or hospital.

[0080] Step 19:

[0081] Information about support organizations is sent to the terminal and presented to the user.

[0082] (Example 1)

[0083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0084] When foreign tourists visit Japan, their travel experiences may be limited due to individual interests or language differences. Furthermore, they may not have sufficient information about local culture and regulations, leading to misunderstandings. In addition, they may face difficulties in responding quickly in emergencies. The challenge is to resolve these issues and enable tourists to have a more fulfilling experience.

[0085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0086] In this invention, the server includes means for receiving user information and generating a dynamic travel plan, means for converting voice input into text using speech recognition technology and performing multilingual translation, and means for retrieving local rules and cultural information from a database and providing it to the user. This allows travelers to receive individually customized travel plans and communicate smoothly across language barriers. Furthermore, by obtaining accurate information about local culture and rules, cultural misunderstandings can be avoided. In addition, based on location information, information on support organizations can be quickly provided in emergencies, supporting a safe and secure trip.

[0087] "User information" refers to basic personal data entered by travelers, including name, nationality, language spoken, and areas of interest.

[0088] A "dynamic travel plan" is a personalized travel plan generated by a server based on the user's interests and planned destinations.

[0089] "Voice recognition technology" is a technology that allows a device to understand the user's voice input and convert it into text data.

[0090] "Multilingual translation" is the process of converting text or audio input in a specific language into another specified language.

[0091] "Local rules and cultural information" refers to information about the culture, laws, and customs of the area that travelers are visiting.

[0092] "Location information" refers to data about the user's current location acquired by the device, and is useful information in emergencies.

[0093] A "generative AI model" is an artificial intelligence model designed to analyze user input data and suggest optimal tourist destinations and activities.

[0094] "Suggestions for tourist destinations and activities" refers to tourist spots and activities that the server recommends based on the user's interests and planned destinations.

[0095] In order to implement this invention, the following system configuration is necessary. It mainly consists of three elements: a terminal, a server, and a user.

[0096] The terminal is a device such as a smartphone or tablet owned by the user. A dedicated application is installed on this terminal, providing an interface for the user to input information and interact with the server. The user enters their personal information (name, nationality, language spoken, areas of interest, etc.) through the terminal and registers it with the system. The terminal can also use speech recognition technology to convert the user's voice into text, thereby enabling multilingual translation functionality.

[0097] The server acts as a central processing unit, receiving information sent by the user and storing it in a database. Using a generative AI model, the server generates dynamic travel plans based on the user's interests and input information. The generated travel plans are delivered to the user's device in real time, offering suggestions for tourist destinations and activities. The server also retrieves local regulations and cultural information from the database and provides it to the user as needed. Furthermore, in emergencies, it identifies the nearest support organization based on location information sent from the device and notifies the user of that information.

[0098] For example, if a user wishes to experience culture in Tokyo during a visit to Japan, they would input their preference into their device. The server would then analyze this data and, using a generative AI model, suggest relevant tourist destinations and activities, such as festivals or traditional theatrical performances in Asakusa. The user would receive this information on their device and be able to plan their trip. Furthermore, the multilingual translation function would enable smoother communication during their visit.

[0099] Examples of prompt statements include the following:

[0100] "Could you recommend a temple-hopping itinerary in Kyoto?"

[0101] "Please tell me about regional cuisine that I can experience in Tokyo."

[0102] These configurations and functions provide support to foreign travelers to enhance their travel experience within Japan, ensuring it is safe and efficient.

[0103] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0104] Step 1:

[0105] Users activate a device such as a smartphone, access a dedicated application, and enter their personal information. This information includes name, nationality, language spoken, and areas of interest. This data is necessary for customizing travel plans, and the entered data is transmitted from the device to a server.

[0106] Step 2:

[0107] The server stores user information received from the terminal in a database. At this point, a user profile is generated based on the entered name, nationality, and areas of interest. The stored data forms the basis for generating user-specific travel plans.

[0108] Step 3:

[0109] The server uses a generative AI model to generate dynamic travel plans based on stored user information. Input includes the user's interests and planned destinations. Based on this information, the server references tourist spots and activities in a database to construct the optimal plan. The generated travel plan is sent to the user's device and presented to them.

[0110] Step 4:

[0111] The device provides multilingual translation capabilities to support communication in the local area. Users input either voice or text through the application. In the case of voice input, the device uses speech recognition technology to convert it into text. The converted text is sent to a server and translated into the specified language. The translation result is returned to the device, and the user uses it in conversations in the local area.

[0112] Step 5:

[0113] When a user needs information about local rules and culture, they can query the server from their device. The server retrieves the relevant information from its database and sends it to the device. This information helps users avoid cultural misunderstandings in their local area.

[0114] Step 6:

[0115] In an emergency, the user activates the "Emergency Assistance" function on their device. The device immediately sends the user's location information to the server. Based on this location information, the server identifies the most appropriate support organization (police station, hospital, etc.) and provides that information to the device. This allows the user to receive the necessary assistance quickly.

[0116] (Application Example 1)

[0117] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0118] When foreign tourists visit Japan, their travel experience is often limited due to language barriers and a lack of cultural understanding. Furthermore, emergencies during their trip and a lack of up-to-date tourist information are also challenges. A system is needed that comprehensively addresses these issues and provides tourists with a better experience of staying in Japan.

[0119] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0120] In this invention, the server includes means for inputting user information and creating a travel itinerary, means for performing multilingual translation, and means for dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests. This enables travelers to obtain appropriate sightseeing information in real time and enjoy their trip with peace of mind, overcoming language and cultural barriers.

[0121] "User information" refers to data including personal information about travelers, their interests, and their accommodation plans, and is used to personalize their travel experience.

[0122] A "travel itinerary" is a plan of sightseeing and activities suggested based on the user's input information, and is a schedule that dynamically changes according to the traveler's preferences and current location.

[0123] "Multilingual translation" is a function that supports communication between users who speak different languages. It is a technology that converts speech and text into a specified language to facilitate the transmission of information.

[0124] A "tour route" is a recommended travel path to tourist spots based on the user's interests and current location, serving as a guide for travelers to efficiently enjoy sightseeing.

[0125] A "local attraction" is a nearby place of cultural or historical significance that travelers should visit, offering a unique local experience.

[0126] "Dynamic suggestions" refers to updating and providing content based on real-time data and conditions, a method of always providing travelers with the most up-to-date information and options.

[0127] The "quiz function" is a feature that encourages users to learn about the culture and customs of their destination in a fun way, and it is a tool that provides information in an interactive format.

[0128] In an embodiment of this invention, the system includes a server, a terminal, and a user as its main components. The server receives information entered by the user through the terminal and performs travel planning, multilingual translation, and dynamic updates of sightseeing suggestions. Specifically, the server processes user information and suggests sightseeing routes and events in real time based on interests and current location. It also provides multilingual translation using voice and text to support communication on-site.

[0129] The hardware primarily consists of the user's smartphone, utilizing GPS and microphone functions for voice input and location information acquisition. Furthermore, a powerful cloud service with a robust database and processing capabilities is used on the server side. For software, Google® Cloud Speech-to-Text is used for speech recognition technology, and Microsoft® Translator is used for translation services, enabling real-time information processing.

[0130] For example, when a user visits a smart city, they can launch the app on their device, use the translation function, and learn about the local culture and customs while sightseeing. For instance, if a traveler asks, "What are some recommended tourist attractions nearby?", the system will suggest the best route and locations.

[0131] An example of a prompt for a generative AI model is, "Please recommend some sightseeing routes during my trip. I'd also like to know about local events."

[0132] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0133] Step 1:

[0134] The user launches the application using a terminal and enters basic user information (e.g., name, nationality, language, areas of interest). The server receives this information and stores it in a database. This process builds the foundational data needed to prepare the most suitable travel plan for the user. The input is the user's basic information, and the output is the database record in which that information is stored.

[0135] Step 2:

[0136] The server generates sightseeing routes and suggested activities based on stored user information and local information about the travel destination. Using a generation AI model, it selects tourist destinations and activities that match the user's interests and proposes an optimized route. The input is user information and local data, and the output is a suggested sightseeing plan. By using prompts, users can instruct the AI ​​with questions such as, "Please tell me your recommended sightseeing route during my trip," to generate a specific plan.

[0137] Step 3:

[0138] The user utilizes GPS functionality through their device to obtain real-time location information. The server uses this location information to list nearby tourist attractions and event information, and sends dynamically updated information to the device. The input is the current GPS data, and the output is information about nearby tourist attractions. Specifically, the system periodically updates nearby attractions in response to changes in the user's location.

[0139] Step 4:

[0140] The terminal converts voice input to text, and the server uses this text to perform multilingual translation. Voice data from the terminal's microphone is processed by Google Cloud Speech-to-Text, converted to text, and then translated by Microsoft Translator before being provided to the user in multiple languages. Input is voice data, and output is translated text. Users can use the translation function to communicate smoothly with local people.

[0141] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0142] This invention relates to a system that supports the travel experience of foreign tourists within Japan, incorporating an emotion engine that recognizes the user's emotions and reflects the recognition results in travel plans and other services. This system mainly consists of terminals, servers, and users, and aims to improve the traveler's experience through the emotion engine.

[0143] The terminal is a device such as a traveler's smartphone, which interacts with the user through a multi-functional application. The user launches the app, registers personal information, and then interacts with the system during subsequent use. The terminal is equipped with an interface for receiving voice and text input, through which the user's emotions are captured.

[0144] The server is responsible for receiving data transmitted from the terminal and storing it in a database. Furthermore, it uses an emotion engine to analyze the user's emotional state from the received audio data and text. Based on this emotion recognition, the server adjusts travel plan suggestions and selects activities that match the user's interests and emotions. In addition, during translation, it controls tone and word choice according to the user's emotions, enabling more personalized communication.

[0145] For example, if a user expresses feelings of anxiety or confusion, the server might suggest relaxing places or activities that provide a sense of security. For instance, it might enhance recommendations for strolling through a Japanese garden or visiting a hot spring facility. Conversely, if a user is excited, the server might plan activities that lead them to more active tourist destinations or activities.

[0146] By combining these emotional engines, it becomes possible to personalize the user's travel experience and increase their satisfaction. This system will be an extremely useful tool in helping travelers understand and enjoy Japanese culture more deeply.

[0147] The following describes the processing flow.

[0148] Step 1:

[0149] Users launch the app on their smartphones, enter the required personal information, and register. After registration, they answer questions to create a travel plan.

[0150] Step 2:

[0151] The terminal processes the user's input data, encrypts it, and sends it to the server. This data includes travel itinerary, destination, and activities of interest.

[0152] Step 3:

[0153] The server stores the received user information in a database and creates a user profile. This lays the foundation for providing personalized services.

[0154] Step 4:

[0155] Users use voice input to communicate emotions to their device or provide some kind of feedback within the app.

[0156] Step 5:

[0157] The device acquires audio data and converts it to text using speech recognition technology. The acquired data is then sent to a server for sentiment analysis.

[0158] Step 6:

[0159] The server uses an emotion engine to analyze the user's emotions from the received voice or text data. This analysis determines, for example, whether the user is excited or seeking relaxation.

[0160] Step 7:

[0161] The server uses the results of sentiment analysis to adjust the travel plan. For example, if the user is feeling anxious, it recommends relaxing activities; if they are excited, it recommends active activities.

[0162] Step 8:

[0163] The terminal displays a pre-arranged travel plan from the server to the user and provides detailed information about selected activities and tourist destinations.

[0164] Step 9:

[0165] If a user needs multilingual translation while on-site, they can input voice or text into their device.

[0166] Step 10:

[0167] The terminal transmits the input voice or text to the server, which then generates a translation based on the user's sentiment.

[0168] Step 11:

[0169] The translation results, which take emotions into account, are sent back to the device, and the device displays or outputs the results to the user as audio. This enables communication in an appropriate tone.

[0170] Step 12:

[0171] In an emergency, the user activates the "emergency assistance" function. At this time, the device sends the user's location information to the server.

[0172] Step 13:

[0173] The server analyzes location information to identify the nearest support organization.

[0174] Step 14:

[0175] Information about support organizations is transmitted to the device and displayed. This allows users to receive appropriate support quickly.

[0176] (Example 2)

[0177] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0178] Currently, many travel support systems provide travel plans without considering the individual feelings of users, and therefore do not adequately address the alleviation of anxiety or the stimulation of interest in cross-cultural environments. Furthermore, translation services also lack the ability to adjust expressions to reflect emotions, which hinders the improvement of communication quality.

[0179] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0180] In this invention, the server includes means for inputting user information and creating a travel plan based on sentiment analysis, means for adjusting the translation expression according to the user's emotions when performing multilingual translation, and means for analyzing the user's emotional state using voice input or text input. This enables the provision of a personalized travel plan that responds to the user's emotions and facilitates more natural and approachable communication.

[0181] "User information" refers to personal data that travelers provide to the system, which is used for creating travel plans and sentiment analysis.

[0182] "Emotional analysis" is the process of identifying a user's emotional state based on voice and text data they input.

[0183] A "travel plan" is a recommendation of tourist destinations and activities suggested based on the user's interests, emotional state, and accommodation plans.

[0184] "Multilingual translation" is the process of converting text or audio into other languages ​​while preserving meaning between them.

[0185] "Adjusting the expression" refers to modifying the translated content to use appropriate tone and expression based on the user's emotional state.

[0186] "Voice input" is an input method in which the system receives and analyzes the voice that the user speaks into the device.

[0187] "Text input" refers to a method in which a user enters text using an input device such as a keyboard, and the system receives that text.

[0188] "Local rules and customs" refers to information about the laws, culture, and social manners of the travel destination.

[0189] An "emergency support agency" is a service provider or organization that travelers can contact if they encounter difficulties.

[0190] This invention is a system that provides a personalized travel experience tailored to the traveler's emotions through interaction between terminals, servers, and users.

[0191] Terminal:

[0192] Users utilize a multi-functional application installed on devices such as smartphones and tablets. First, users launch the application and register their personal information. Input methods include voice input using speech recognition software and text input via a text box, through which users communicate their mood and emotions to the device. Specifically, a user can communicate their current feelings to the system by voice-inputting, "I'm a little tired today."

[0193] server:

[0194] Data sent from the device is received by the server. The server uses a generative AI model capable of advanced natural language processing to analyze the user's emotions from the data. This analysis involves converting the audio data into text and identifying emotions using an emotion engine. Based on the analysis results, the server compares the travel database with the user's current situation to provide the optimal travel plan. It also provides translation assistance, adjusting multilingual translations to match the user's emotions. For example, for a user who needs to relax, it might provide a softer translation such as, "This activity is perfect for refreshing yourself."

[0195] User:

[0196] Users can review the provided travel plan and provide feedback via their device. This feedback allows the system to provide even more personalized services.

[0197] Example of a prompt:

[0198] The prompt message would be something like, "The user said 'I feel a little anxious.' What is the emotional state based on this?" This would lay the foundation for obtaining an appropriate sentiment analysis result.

[0199] Through this invention, travelers can quickly receive travel plans that suit their own feelings, thereby increasing their sense of security and comfort while traveling.

[0200] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0201] Step 1:

[0202] The user launches the application on their device and registers their personal information. They then input their mood or feelings for the day using voice input or a text box. For example, if the user voice-inputs "I'm a little tired today," the device uses voice recognition software to convert this voice into text data. The input is voice data, and the output is text data.

[0203] Step 2:

[0204] The terminal sends the converted text data to the server. During this process, the data is transmitted via a secure communication protocol, ensuring data integrity and privacy. The input is text data, and the output is a data packet sent to the server.

[0205] Step 3:

[0206] The server inputs the received text data into a generative AI model to analyze the user's emotional state. The data is processed by the emotion engine, and the generative AI model performs emotion analysis. During this process, a prompt is issued, such as "The user said 'I'm a little tired.' What is the emotional state based on this?", and the output is a label of the user's emotional state.

[0207] Step 4:

[0208] The server generates an optimal travel plan for the user based on the analysis results. It refers to a travel database and selects tourist destinations and activities that match the user's emotional state. For example, if the server determines the user is tired, it might suggest visiting a relaxing hot spring facility or a quiet garden. The input is an emotional state label, and the output is a customized travel plan.

[0209] Step 5:

[0210] The server sends the generated travel plan to the terminal and simultaneously performs multilingual translation, adjusting the translation expression according to the user's emotions. Translation assistance is used to provide explanations using language that matches the emotions. The input is the travel plan, and the output is plan information including the adjusted translated text.

[0211] Step 6:

[0212] Users review the suggested travel plans on their devices and provide feedback as needed. This feedback is sent back to the server and used as data for future plan generation. The input is user feedback information, and the output is user profile information updated by the server.

[0213] In this way, a series of steps provides the user with the optimal travel experience.

[0214] (Application Example 2)

[0215] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0216] When foreign tourists travel within Japan, they often face anxieties and communication difficulties stemming from language barriers and cultural differences. While conventional travel support systems offer travel plan suggestions and multilingual translation, they have struggled to provide personalized experiences that respond to travelers' immediate emotions and interests. Therefore, there is a need to accurately recognize travelers' emotions and reflect them in travel plans in real time.

[0217] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0218] In this invention, the server includes means for inputting user information and creating a travel plan, means for performing multilingual translation, and means for analyzing the user's emotional state and adjusting the travel plan based on the analysis results. This makes it possible to suggest sightseeing activities that are in line with the traveler's emotions and to use appropriate language expressions that correspond to those emotions.

[0219] "Foreign tourists" refers to travelers who come from outside Japan and whose purpose is to experience Japanese tourist attractions and culture.

[0220] An "information processing system to support travel experiences" is a system that includes information processing means to provide information and support so that travelers can have a smooth and fulfilling sightseeing experience.

[0221] "Means for inputting user information and creating travel plans" refers to data input and processing means for generating appropriate sightseeing plans based on the traveler's interests and schedule.

[0222] "Means of performing multilingual translation" refers to means of translating between the languages ​​of different countries, enabling travelers to communicate across language barriers.

[0223] "Means for analyzing a user's emotional state and adjusting travel plans based on the analysis results" refers to processing methods that analyze a traveler's emotions and, accordingly, suggest the most suitable travel activities and plans.

[0224] "A means of suggesting tourism activities based on the user's current emotions in real time" refers to a means that captures changes in a traveler's emotions in real time and immediately presents suitable tourism activities.

[0225] "Means for generating language expressions that respond to the user's emotional state" refers to means for adjusting translated content to more appropriate and natural language expressions that match the traveler's emotions.

[0226] The system implementing this invention includes a personal information terminal carried by the traveler, a cloud-based information processing server, and a user interface to facilitate user interaction. Primarily, the personal information terminal receives voice and text input from the traveler and relays it to the server. The traveler uses a travel support application to register personal information and confirm travel plans.

[0227] When the server receives data transmitted from a mobile device, it first accesses a database to manage the traveler's basic information. For sentiment analysis, it uses the Google Cloud Natural Language API to analyze the traveler's emotional state from voice and text data. Based on this analysis, the server adjusts travel plan suggestions and selects activities and tourist destinations that match the traveler's emotional state.

[0228] Furthermore, the server overcomes language barriers using a multilingual translation AI engine to improve the naturalness of communication. Specifically, it adaptively adjusts tone and wording by referencing the traveler's emotions during translation, enabling more personalized communication.

[0229] As a concrete example, when a traveler is visiting multiple tourist spots in Tokyo, if the system analyzes the user's emotions suggesting they are "tired," the server will suggest quiet gardens or relaxing cafes. Conversely, if the user indicates they are feeling energetic, the system will guide them to lively shopping malls or events. As an example of a prompt to the generating AI model, inputting "Please suggest tourist spots recommended when the user's emotion is 'excited'" will cause the AI ​​to generate appropriate suggestions.

[0230] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0231] Step 1:

[0232] Users launch the application using their mobile devices and enter personal information. This input data, including the traveler's preferences and basic travel requests, is sent from the device to the server. The server receives this data and stores it in a database to manage information for each traveler.

[0233] Step 2:

[0234] When a user inputs voice or text through the application, the device sends this data to the server in real time. The server uses the Google Cloud Natural Language API to analyze this data and identify the traveler's emotional state. The emotional state is categorized into states such as excited, relaxed, or anxious, and the analysis results are stored on the server.

[0235] Step 3:

[0236] The server utilizes a generative AI model based on the results of emotion analysis to adjust travel plans according to the user's interests. For example, if an excited state is detected, the server uses the AI ​​model to generate a plan that suggests active tourist destinations and events. This generated plan is sent to the device and presented to the user within the application.

[0237] Step 4:

[0238] The server works in conjunction with a multilingual translation engine to translate the output text from the device according to the user's emotional state. The text is converted to an appropriate tone and phrasing, and the result is displayed to the user. Specifically, it provides communication optimized for travelers by increasing politeness or adjusting to more casual expressions.

[0239] Step 5:

[0240] Based on the presented travel plan and translated information, the user selects their next action. By inputting new voice or text data again, sentiment analysis is performed again, and new prompt sentences can be generated. For example, if the user inputs "Please provide activities that are best suited to my current mood," appropriate suggestions will be generated.

[0241] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0242] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0243] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0244] [Second Embodiment]

[0245] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0246] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0247] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0248] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0249] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0250] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0251] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0252] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0253] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0254] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0255] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0256] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0257] This invention provides a system to support the travel experience of foreign tourists visiting Japan. This system consists of three main components: a terminal, a server, and a user.

[0258] The terminal is the traveler's smartphone or other device, through which the traveler interacts with the system. The user uses the terminal to register personal information before the start of their trip, including information such as name, nationality, language spoken, and areas of interest.

[0259] The server receives information sent by users and stores it in a database. This lays the foundation for providing personalized travel plans tailored to travelers. The server has the function of suggesting tourist spots and activities based on the traveler's interests and information about the area they are staying in.

[0260] Furthermore, the device provides a multilingual translation function to help users communicate more smoothly in the local area. Users can use the translation function through the device to convert voice and text into other languages. In this process, speech recognition technology is used to process the input, and the server generates the translated data and sends it back to the device.

[0261] Regarding explanations of local rules, when a user asks a question about local regulations or customs, the server looks up information in the database and provides the device with the appropriate information. This helps avoid cultural misunderstandings and ensures that local manners are observed.

[0262] Furthermore, in an emergency, the user can activate the "emergency assistance" function, and the device immediately sends its location information to a server. Based on that location, the server quickly provides information on the nearest police station, hospital, or other appropriate assistance organization. This information is then presented to the user through the device.

[0263] As a concrete example, consider a scenario where a user visits a major Japanese city and creates a plan to maximize their local cultural experience. Before starting their trip, the user inputs their areas of interest and activities through the app. The server analyzes this input and provides recommended plans, such as temple tours in Kyoto or local cuisine experiences in Tokyo. This allows the user to make the most of their time during their stay.

[0264] The following describes the processing flow.

[0265] Step 1:

[0266] The user launches the application on their device and opens the "New Registration" screen. They enter personal information such as their name, nationality, language settings, and areas of interest, and then press the submit button.

[0267] Step 2:

[0268] The device collects data entered by the user, encrypts it, and then sends it to the server.

[0269] Step 3:

[0270] The server analyzes the received user data and stores it in a database. This prepares the server to create travel plans based on the user's individual needs.

[0271] Step 4:

[0272] The user selects the "Create a travel plan" option and enters their interests and desired tourist destinations.

[0273] Step 5:

[0274] The device sends the user's requested information to the server.

[0275] Step 6:

[0276] The server uses AI to suggest optimal tourist spots and activities based on the user's interests and revised plan. These suggestions are then sent to the user's device.

[0277] Step 7:

[0278] The terminal displays suggestions sent from the server to the user, allowing the user to review the details and make modifications as needed.

[0279] Step 8:

[0280] If a user needs to translate a conversation while on location, they can use the app's translation function.

[0281] Step 9:

[0282] The terminal receives the voice input and converts the voice into text.

[0283] Step 10:

[0284] The text data is sent to the server and translated into the specified language.

[0285] Step 11:

[0286] The translated text from the server is returned to the terminal and output to the user in the form of display or voice.

[0287] Step 12:

[0288] When the user requests information about local rules and customs, the user inputs relevant questions into the application.

[0289] Step 13:

[0290] The terminal sends the input questions to the server.

[0291] Step 14:

[0292] The server searches the database and sends appropriate local rule guidance to the terminal.

[0293] Step 15:

[0294] The terminal displays the guidance information sent from the server to the user.

[0295] Step 16:

[0296] In case of emergency, the user presses the "Emergency Support" button.

[0297] Step 17:

[0298] The device sends data, including its current location, to the server.

[0299] Step 18:

[0300] The server analyzes the location information to identify the nearest police station or hospital.

[0301] Step 19:

[0302] Information about support organizations is sent to the terminal and presented to the user.

[0303] (Example 1)

[0304] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0305] When foreign tourists visit Japan, their travel experiences may be limited due to individual interests or language differences. Furthermore, they may not have sufficient information about local culture and regulations, leading to misunderstandings. In addition, they may face difficulties in responding quickly in emergencies. The challenge is to resolve these issues and enable tourists to have a more fulfilling experience.

[0306] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0307] In this invention, the server includes means for receiving user information and generating a dynamic travel plan, means for converting voice input into text using speech recognition technology and performing multilingual translation, and means for retrieving local rules and cultural information from a database and providing it to the user. This allows travelers to receive individually customized travel plans and communicate smoothly across language barriers. Furthermore, by obtaining accurate information about local culture and rules, cultural misunderstandings can be avoided. In addition, based on location information, information on support organizations can be quickly provided in emergencies, supporting a safe and secure trip.

[0308] "User information" refers to the basic personal data input by travelers and includes name, nationality, language used, fields of interest, etc.

[0309] "Dynamic travel plan" refers to a personalized travel plan generated by the server based on the user's interests and planned destinations.

[0310] "Voice recognition technology" is the technology for the terminal to understand the user's voice input and convert it into text data.

[0311] "Multi-language translation" is the process of converting text or voice input in a specific language into another specified language.

[0312] "Local rules and cultural information" refers to information about the culture, laws, and customs of the region visited by travelers.

[0313] "Location information" is the data related to the user's current location obtained by the terminal and is information useful in case of emergency.

[0314] "Generative AI model" is an artificial intelligence model designed to analyze input data from users and propose optimal tourist destinations and activities.

[0315] "Recommendation of tourist destinations and activities" refers to the tourist spots and activities that can be participated in recommended by the server based on the user's interests and planned destinations.

[0316] To implement this invention, the following system configuration is required. It is mainly composed of three elements: a terminal, a server, and a user.

[0317] The terminal is a device such as a smartphone or tablet owned by the user. A dedicated application is installed on this terminal, providing an interface for the user to input information and interact with the server. The user enters their personal information (name, nationality, language spoken, areas of interest, etc.) through the terminal and registers it with the system. The terminal can also use speech recognition technology to convert the user's voice into text, thereby enabling multilingual translation functionality.

[0318] The server acts as a central processing unit, receiving information sent by the user and storing it in a database. Using a generative AI model, the server generates dynamic travel plans based on the user's interests and input information. The generated travel plans are delivered to the user's device in real time, offering suggestions for tourist destinations and activities. The server also retrieves local regulations and cultural information from the database and provides it to the user as needed. Furthermore, in emergencies, it identifies the nearest support organization based on location information sent from the device and notifies the user of that information.

[0319] For example, if a user wishes to experience culture in Tokyo during a visit to Japan, they would input their preference into their device. The server would then analyze this data and, using a generative AI model, suggest relevant tourist destinations and activities, such as festivals or traditional theatrical performances in Asakusa. The user would receive this information on their device and be able to plan their trip. Furthermore, the multilingual translation function would enable smoother communication during their visit.

[0320] Examples of prompt statements include the following:

[0321] "Could you recommend a temple-hopping itinerary in Kyoto?"

[0322] "Please tell me about regional cuisine that I can experience in Tokyo."

[0323] These configurations and functions provide support to foreign travelers to enhance their travel experience within Japan, ensuring it is safe and efficient.

[0324] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0325] Step 1:

[0326] Users activate a device such as a smartphone, access a dedicated application, and enter their personal information. This information includes name, nationality, language spoken, and areas of interest. This data is necessary for customizing travel plans, and the entered data is transmitted from the device to a server.

[0327] Step 2:

[0328] The server stores user information received from the terminal in a database. At this point, a user profile is generated based on the entered name, nationality, and areas of interest. The stored data forms the basis for generating user-specific travel plans.

[0329] Step 3:

[0330] The server uses a generative AI model to generate dynamic travel plans based on stored user information. Input includes the user's interests and planned destinations. Based on this information, the server references tourist spots and activities in a database to construct the optimal plan. The generated travel plan is sent to the user's device and presented to them.

[0331] Step 4:

[0332] The device provides multilingual translation capabilities to support communication in the local area. Users input either voice or text through the application. In the case of voice input, the device uses speech recognition technology to convert it into text. The converted text is sent to a server and translated into the specified language. The translation result is returned to the device, and the user uses it in conversations in the local area.

[0333] Step 5:

[0334] When a user needs information about local rules and culture, they can query the server from their device. The server retrieves the relevant information from its database and sends it to the device. This information helps users avoid cultural misunderstandings in their local area.

[0335] Step 6:

[0336] In an emergency, the user activates the "Emergency Assistance" function on their device. The device immediately sends the user's location information to the server. Based on this location information, the server identifies the most appropriate support organization (police station, hospital, etc.) and provides that information to the device. This allows the user to receive the necessary assistance quickly.

[0337] (Application Example 1)

[0338] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0339] When foreign tourists visit Japan, their travel experience is often limited due to language barriers and a lack of cultural understanding. Furthermore, emergencies during their trip and a lack of up-to-date tourist information are also challenges. A system is needed that comprehensively addresses these issues and provides tourists with a better experience of staying in Japan.

[0340] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0341] In this invention, the server includes means for inputting user information and creating a travel itinerary, means for performing multilingual translation, and means for dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests. This enables travelers to obtain appropriate sightseeing information in real time and enjoy their trip with peace of mind, overcoming language and cultural barriers.

[0342] "User information" refers to data including personal information about travelers, their interests, and their accommodation plans, and is used to personalize their travel experience.

[0343] A "travel itinerary" is a plan of sightseeing and activities suggested based on the user's input information, and is a schedule that dynamically changes according to the traveler's preferences and current location.

[0344] "Multilingual translation" is a function that supports communication between users who speak different languages. It is a technology that converts speech and text into a specified language to facilitate the transmission of information.

[0345] A "tour route" is a recommended travel path to tourist spots based on the user's interests and current location, serving as a guide for travelers to efficiently enjoy sightseeing.

[0346] A "local attraction" is a nearby place of cultural or historical significance that travelers should visit, offering a unique local experience.

[0347] "Dynamic suggestion" refers to updating and providing content based on real-time data and conditions, a method of always providing travelers with the most up-to-date information and options.

[0348] The "quiz function" is a feature that encourages users to learn about the culture and customs of their destination in a fun way, and it is a tool that provides information in an interactive format.

[0349] In an embodiment of this invention, the system includes a server, a terminal, and a user as its main components. The server receives information entered by the user through the terminal and performs travel planning, multilingual translation, and dynamic updates of sightseeing suggestions. Specifically, the server processes user information and suggests sightseeing routes and events in real time based on interests and current location. It also provides multilingual translation using voice and text to support communication on-site.

[0350] The hardware primarily consists of the user's smartphone, utilizing GPS and microphone functions for voice input and location information acquisition. Furthermore, a powerful cloud service with a robust database and processing capabilities is used on the server side. For software, Google Cloud Speech-to-Text is used for speech recognition technology, and Microsoft Translator for translation services, enabling real-time information processing.

[0351] For example, when a user visits a smart city, they can launch the app on their device, use the translation function, and learn about the local culture and customs while sightseeing. For instance, if a traveler asks, "What are some recommended tourist attractions nearby?", the system will suggest the best route and locations.

[0352] An example of a prompt for a generative AI model is, "Please recommend some sightseeing routes during my trip. I'd also like to know about local events."

[0353] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0354] Step 1:

[0355] The user launches the application using a terminal and enters basic user information (e.g., name, nationality, language, areas of interest). The server receives this information and stores it in a database. This process builds the foundational data needed to prepare the most suitable travel plan for the user. The input is the user's basic information, and the output is the database record in which that information is stored.

[0356] Step 2:

[0357] The server generates sightseeing routes and suggested activities based on stored user information and regional information of the travel destination. Using a generation AI model, it selects tourist destinations and activities that match the user's interests and proposes an optimized route. The input is user information and regional data, and the output is a suggested sightseeing plan. By using prompts, users can instruct the AI ​​with questions such as, "Please tell me your recommended sightseeing route during my trip," to generate a specific plan.

[0358] Step 3:

[0359] The user utilizes GPS functionality through their device to obtain real-time location information. The server uses this location information to list nearby tourist attractions and event information, and sends dynamically updated information to the device. The input is the current GPS data, and the output is information about nearby tourist attractions. Specifically, the system periodically updates nearby attractions in response to changes in the user's location.

[0360] Step 4:

[0361] The terminal converts voice input to text, and the server uses this text to perform multilingual translation. Voice data from the terminal's microphone is processed by Google Cloud Speech-to-Text, converted to text, and then translated by Microsoft Translator before being provided to the user in multiple languages. Input is voice data, and output is translated text. Users can use the translation function to communicate smoothly with local people.

[0362] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0363] This invention relates to a system that supports the travel experience of foreign tourists within Japan, incorporating an emotion engine that recognizes the user's emotions and reflects the recognition results in travel plans and other services. This system mainly consists of terminals, servers, and users, and aims to improve the traveler's experience through the emotion engine.

[0364] The terminal is a device such as a traveler's smartphone, which interacts with the user through a multi-functional application. The user launches the app, registers personal information, and then interacts with the system during subsequent use. The terminal is equipped with an interface for receiving voice and text input, through which the user's emotions are captured.

[0365] The server is responsible for receiving data transmitted from the terminal and storing it in a database. Furthermore, it uses an emotion engine to analyze the user's emotional state from the received audio data and text. Based on this emotion recognition, the server adjusts travel plan suggestions and selects activities that match the user's interests and emotions. In addition, during translation, it controls tone and word choice according to the user's emotions, enabling more personalized communication.

[0366] For example, if a user expresses feelings of anxiety or confusion, the server might suggest relaxing places or activities that provide a sense of security. For instance, it might enhance recommendations for strolling through a Japanese garden or visiting a hot spring facility. Conversely, if a user is excited, the server might plan activities that lead them to more active tourist destinations or activities.

[0367] By combining these emotional engines, it becomes possible to personalize the user's travel experience and increase their satisfaction. This system will be an extremely useful tool in helping travelers understand and enjoy Japanese culture more deeply.

[0368] The following describes the processing flow.

[0369] Step 1:

[0370] Users launch the app on their smartphones, enter the required personal information, and register. After registration, they answer questions to create a travel plan.

[0371] Step 2:

[0372] The terminal processes the user's input data, encrypts it, and sends it to the server. This data includes travel itinerary, destination, and activities of interest.

[0373] Step 3:

[0374] The server stores the received user information in a database and creates a user profile. This lays the foundation for providing personalized services.

[0375] Step 4:

[0376] Users use voice input to communicate emotions to their device or provide some kind of feedback within the app.

[0377] Step 5:

[0378] The device acquires audio data and converts it to text using speech recognition technology. The acquired data is then sent to a server for sentiment analysis.

[0379] Step 6:

[0380] The server uses an emotion engine to analyze the user's emotions from the received voice or text data. This analysis determines, for example, whether the user is excited or seeking relaxation.

[0381] Step 7:

[0382] The server uses the results of sentiment analysis to adjust the travel plan. For example, if the user is feeling anxious, it recommends relaxing activities; if they are excited, it recommends active activities.

[0383] Step 8:

[0384] The terminal displays a pre-arranged travel plan from the server to the user and provides detailed information about selected activities and tourist destinations.

[0385] Step 9:

[0386] If a user needs multilingual translation while on-site, they can input voice or text into their device.

[0387] Step 10:

[0388] The terminal transmits the input voice or text to the server, which then generates a translation based on the user's sentiment.

[0389] Step 11:

[0390] The translation results, which take emotions into account, are sent back to the device, and the device displays or outputs the results to the user as audio. This enables communication in an appropriate tone.

[0391] Step 12:

[0392] In an emergency, the user activates the "emergency assistance" function. At this time, the device sends the user's location information to the server.

[0393] Step 13:

[0394] The server analyzes location information to identify the nearest support organization.

[0395] Step 14:

[0396] Information about support organizations is transmitted to the device and displayed. This allows users to receive appropriate support quickly.

[0397] (Example 2)

[0398] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0399] Currently, many travel support systems provide travel plans without considering the individual feelings of users, and therefore do not adequately address the alleviation of anxiety or the stimulation of interest in cross-cultural environments. Furthermore, translation services also lack the ability to adjust expressions to reflect emotions, which hinders the improvement of communication quality.

[0400] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0401] In this invention, the server includes means for inputting user information and creating a travel plan based on sentiment analysis, means for adjusting the translation expression according to the user's emotions when performing multilingual translation, and means for analyzing the user's emotional state using voice input or text input. This enables the provision of a personalized travel plan that responds to the user's emotions and facilitates more natural and approachable communication.

[0402] "User information" refers to personal data that travelers provide to the system, which is used for creating travel plans and sentiment analysis.

[0403] "Emotional analysis" is the process of identifying a user's emotional state based on voice and text data they input.

[0404] A "travel plan" is a recommendation of tourist destinations and activities suggested based on the user's interests, emotional state, and accommodation plans.

[0405] "Multilingual translation" is the process of converting text or audio into other languages ​​while preserving meaning between them.

[0406] "Adjusting the expression" refers to modifying the translated content to use appropriate tone and expression based on the user's emotional state.

[0407] "Voice input" is an input method in which the system receives and analyzes the voice that the user speaks into the device.

[0408] "Text input" refers to a method in which a user enters text using an input device such as a keyboard, and the system receives that text.

[0409] "Local rules and customs" refers to information about the laws, culture, and social manners of the travel destination.

[0410] An "emergency support agency" is a service provider or organization that travelers can contact if they encounter difficulties.

[0411] This invention is a system that provides a personalized travel experience tailored to the traveler's emotions through interaction between terminals, servers, and users.

[0412] Terminal:

[0413] Users utilize a multi-functional application installed on devices such as smartphones and tablets. First, users launch the application and register their personal information. Input methods include voice input using speech recognition software and text input via a text box, through which users communicate their mood and emotions to the device. Specifically, a user can communicate their current feelings to the system by voice-inputting, "I'm a little tired today."

[0414] server:

[0415] Data sent from the device is received by the server. The server uses a generative AI model capable of advanced natural language processing to analyze the user's emotions from the data. This analysis involves converting the audio data into text and identifying emotions using an emotion engine. Based on the analysis results, the server compares the travel database with the user's current situation to provide the optimal travel plan. It also provides translation assistance, adjusting multilingual translations to match the user's emotions. For example, for a user who needs to relax, it might provide a softer translation such as, "This activity is perfect for refreshing yourself."

[0416] User:

[0417] Users can review the provided travel plan and provide feedback via their device. This feedback allows the system to provide even more personalized services.

[0418] Example of a prompt:

[0419] The prompt message would be something like, "The user said, 'I feel a little anxious.' What is the emotional state based on this?" This would lay the foundation for obtaining an appropriate sentiment analysis result.

[0420] Through this invention, travelers can quickly receive travel plans that suit their own feelings, thereby increasing their sense of security and comfort while traveling.

[0421] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0422] Step 1:

[0423] The user launches the application on their device and registers their personal information. They then input their mood or feelings for the day using voice input or a text box. For example, if the user voice-inputs "I'm a little tired today," the device uses voice recognition software to convert this voice into text data. The input is voice data, and the output is text data.

[0424] Step 2:

[0425] The terminal sends the converted text data to the server. During this process, the data is transmitted via a secure communication protocol, ensuring data integrity and privacy. The input is text data, and the output is a data packet sent to the server.

[0426] Step 3:

[0427] The server inputs the received text data into a generative AI model to analyze the user's emotional state. The data is processed by the emotion engine, and the generative AI model performs emotion analysis. During this process, a prompt is issued, such as "The user said 'I'm a little tired.' What is the emotional state based on this?", and the output is a label of the user's emotional state.

[0428] Step 4:

[0429] The server generates an optimal travel plan for the user based on the analysis results. It refers to a travel database and selects tourist destinations and activities that match the user's emotional state. For example, if the server determines the user is tired, it might suggest visiting a relaxing hot spring facility or a quiet garden. The input is an emotional state label, and the output is a customized travel plan.

[0430] Step 5:

[0431] The server sends the generated travel plan to the terminal and simultaneously performs multilingual translation, adjusting the translation expression according to the user's emotions. Translation assistance is used to provide explanations using language that matches the emotions. The input is the travel plan, and the output is plan information including the adjusted translated text.

[0432] Step 6:

[0433] Users review the suggested travel plans on their devices and provide feedback as needed. This feedback is sent back to the server and used as data for future plan generation. The input is user feedback information, and the output is user profile information updated by the server.

[0434] In this way, a series of steps provides the user with the optimal travel experience.

[0435] (Application Example 2)

[0436] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0437] When foreign tourists travel within Japan, they often face anxieties and communication difficulties stemming from language barriers and cultural differences. While conventional travel support systems offer travel plan suggestions and multilingual translation, they have struggled to provide personalized experiences that respond to travelers' immediate emotions and interests. Therefore, there is a need to accurately recognize travelers' emotions and reflect them in travel plans in real time.

[0438] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0439] In this invention, the server includes means for inputting user information and creating a travel plan, means for performing multilingual translation, and means for analyzing the user's emotional state and adjusting the travel plan based on the analysis results. This makes it possible to suggest sightseeing activities that are in line with the traveler's emotions and to use appropriate language expressions that correspond to those emotions.

[0440] "Foreign tourists" refers to travelers who come from outside Japan and whose purpose is to experience Japanese tourist attractions and culture.

[0441] An "information processing system to support travel experiences" is a system that includes information processing means to provide information and support so that travelers can have a smooth and fulfilling sightseeing experience.

[0442] "Means for inputting user information and creating travel plans" refers to data input and processing means for generating appropriate sightseeing plans based on the traveler's interests and schedule.

[0443] "Means of performing multilingual translation" refers to means of translating between the languages ​​of different countries, enabling travelers to communicate across language barriers.

[0444] "Means for analyzing a user's emotional state and adjusting travel plans based on the analysis results" refers to processing methods that analyze a traveler's emotions and, accordingly, propose the most suitable travel activities and plans.

[0445] "A means of suggesting tourism activities based on the user's current emotions in real time" refers to a means that captures changes in a traveler's emotions in real time and immediately presents suitable tourism activities.

[0446] "Means for generating language expressions that respond to the user's emotional state" refers to means for adjusting translated content to more appropriate and natural language expressions that match the traveler's emotions.

[0447] The system implementing this invention includes a personal information terminal carried by the traveler, a cloud-based information processing server, and a user interface to facilitate user interaction. Primarily, the personal information terminal receives voice and text input from the traveler and relays it to the server. The traveler uses a travel support application to register personal information and confirm travel plans.

[0448] When the server receives data transmitted from a mobile device, it first accesses a database to manage the traveler's basic information. For sentiment analysis, it uses the Google Cloud Natural Language API to analyze the traveler's emotional state from voice and text data. Based on this analysis, the server adjusts travel plan suggestions and selects activities and tourist destinations that match the traveler's emotional state.

[0449] Furthermore, the server overcomes language barriers using a multilingual translation AI engine to improve the naturalness of communication. Specifically, it adaptively adjusts tone and wording by referencing the traveler's emotions during translation, enabling more personalized communication.

[0450] As a concrete example, when a traveler is visiting multiple tourist spots in Tokyo, if the system analyzes the user's emotions suggesting they are "tired," the server will suggest quiet gardens or relaxing cafes. Conversely, if the user indicates they are feeling energetic, the system will guide them to lively shopping malls or events. As an example of a prompt to the generating AI model, inputting "Please suggest tourist spots recommended when the user's emotion is 'excited'" will cause the AI ​​to generate appropriate suggestions.

[0451] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0452] Step 1:

[0453] Users launch the application using their mobile devices and enter personal information. This input data, including the traveler's preferences and basic travel requests, is sent from the device to the server. The server receives this data and stores it in a database to manage information for each traveler.

[0454] Step 2:

[0455] When a user inputs voice or text through the application, the device sends this data to the server in real time. The server uses the Google Cloud Natural Language API to analyze this data and identify the traveler's emotional state. The emotional state is categorized into states such as excited, relaxed, or anxious, and the analysis results are stored on the server.

[0456] Step 3:

[0457] The server utilizes a generative AI model based on the results of emotion analysis to adjust travel plans according to the user's interests. For example, if an excited state is detected, the server uses the AI ​​model to generate a plan that suggests active tourist destinations and events. This generated plan is sent to the device and presented to the user within the application.

[0458] Step 4:

[0459] The server works in conjunction with a multilingual translation engine to translate the output text from the device according to the user's emotional state. The text is converted to an appropriate tone and phrasing, and the result is displayed to the user. Specifically, it provides communication optimized for travelers by increasing politeness or adjusting to more casual expressions.

[0460] Step 5:

[0461] Based on the presented travel plan and translated information, the user selects their next action. By inputting new voice or text data again, sentiment analysis is performed again, and new prompt sentences can be generated. For example, if the user inputs "Please provide activities that are best suited to my current mood," appropriate suggestions will be generated.

[0462] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0463] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0464] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0465] [Third Embodiment]

[0466] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0467] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0468] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0469] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0470] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0471] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0472] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0473] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0474] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0475] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0476] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0477] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0478] This invention provides a system to support the travel experience of foreign tourists visiting Japan. This system consists of three main components: a terminal, a server, and a user.

[0479] The terminal is the traveler's smartphone or other device, through which the traveler interacts with the system. The user uses the terminal to register personal information before the start of their trip, including information such as name, nationality, language spoken, and areas of interest.

[0480] The server receives information sent by users and stores it in a database. This lays the foundation for providing personalized travel plans tailored to travelers. The server has the function of suggesting tourist spots and activities based on the traveler's interests and information about the area they are staying in.

[0481] Furthermore, the device provides a multilingual translation function to help users communicate more smoothly in the local area. Users can use the translation function through the device to convert voice and text into other languages. In this process, speech recognition technology is used to process the input, and the server generates the translated data and sends it back to the device.

[0482] Regarding explanations of local rules, when a user asks a question about local regulations or customs, the server looks up information in the database and provides the device with the appropriate information. This helps avoid cultural misunderstandings and ensures that local manners are observed.

[0483] Furthermore, in an emergency, the user can activate the "emergency assistance" function, and the device immediately sends its location information to a server. Based on that location, the server quickly provides information on the nearest police station, hospital, or other appropriate assistance organization. This information is then presented to the user through the device.

[0484] As a concrete example, consider a scenario where a user visits a major Japanese city and creates a plan to maximize their local cultural experience. Before starting their trip, the user inputs their areas of interest and activities through the app. The server analyzes this input and provides recommended plans, such as temple tours in Kyoto or local cuisine experiences in Tokyo. This allows the user to make the most of their time during their stay.

[0485] The following describes the processing flow.

[0486] Step 1:

[0487] The user launches the application on their device and opens the "New Registration" screen. They enter personal information such as their name, nationality, language settings, and areas of interest, and then press the submit button.

[0488] Step 2:

[0489] The device collects data entered by the user, encrypts it, and then sends it to the server.

[0490] Step 3:

[0491] The server analyzes the received user data and stores it in a database. This prepares the server to create travel plans based on the user's individual needs.

[0492] Step 4:

[0493] The user selects the "Create a travel plan" option and enters their interests and desired tourist destinations.

[0494] Step 5:

[0495] The device sends the user's requested information to the server.

[0496] Step 6:

[0497] The server uses AI to suggest optimal tourist spots and activities based on the user's interests and revised plan. These suggestions are then sent to the user's device.

[0498] Step 7:

[0499] The terminal displays suggestions sent from the server to the user, allowing the user to review the details and make modifications as needed.

[0500] Step 8:

[0501] If a user needs to translate a conversation while on location, they can use the app's translation function.

[0502] Step 9:

[0503] The device receives voice input and converts the speech into text.

[0504] Step 10:

[0505] Text data is sent to the server and translated into the specified language.

[0506] Step 11:

[0507] The translated text is sent back to the terminal from the server and displayed or outputted to the user as audio.

[0508] Step 12:

[0509] When a user seeks information about local rules and customs, they enter relevant questions into the app.

[0510] Step 13:

[0511] The terminal sends the entered question to the server.

[0512] Step 14:

[0513] The server searches the database and sends instructions for the appropriate local rules to the terminal.

[0514] Step 15:

[0515] The terminal displays guidance information sent from the server to the user.

[0516] Step 16:

[0517] In an emergency, the user presses the "Emergency Assistance" button.

[0518] Step 17:

[0519] The device sends data, including its current location, to the server.

[0520] Step 18:

[0521] The server analyzes the location information to identify the nearest police station or hospital.

[0522] Step 19:

[0523] Information about support organizations is sent to the terminal and presented to the user.

[0524] (Example 1)

[0525] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0526] When foreign tourists visit Japan, their travel experiences may be limited due to individual interests or language differences. Furthermore, they may not have sufficient information about local culture and regulations, leading to misunderstandings. In addition, they may face difficulties in responding quickly in emergencies. The challenge is to resolve these issues and enable tourists to have a more fulfilling experience.

[0527] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0528] In this invention, the server includes means for receiving user information and generating a dynamic travel plan, means for converting voice input into text using speech recognition technology and performing multilingual translation, and means for retrieving local rules and cultural information from a database and providing it to the user. This allows travelers to receive individually customized travel plans and communicate smoothly across language barriers. Furthermore, by obtaining accurate information about local culture and rules, cultural misunderstandings can be avoided. In addition, based on location information, information on support organizations can be quickly provided in emergencies, supporting a safe and secure trip.

[0529] "User information" refers to basic personal data entered by travelers, including name, nationality, language spoken, and areas of interest.

[0530] A "dynamic travel plan" is a personalized travel plan generated by a server based on the user's interests and planned destinations.

[0531] "Voice recognition technology" is a technology that allows a device to understand the user's voice input and convert it into text data.

[0532] "Multilingual translation" is the process of converting text or audio input in a specific language into another specified language.

[0533] "Local rules and cultural information" refers to information about the culture, laws, and customs of the area that travelers are visiting.

[0534] "Location information" refers to data about the user's current location acquired by the device, and is useful information in emergencies.

[0535] A "generative AI model" is an artificial intelligence model designed to analyze user input data and suggest optimal tourist destinations and activities.

[0536] "Suggestions for tourist destinations and activities" refers to tourist spots and activities that the server recommends based on the user's interests and planned destinations.

[0537] In order to implement this invention, the following system configuration is necessary. It mainly consists of three elements: a terminal, a server, and a user.

[0538] The terminal is a device such as a smartphone or tablet owned by the user. A dedicated application is installed on this terminal, providing an interface for the user to input information and interact with the server. The user enters their personal information (name, nationality, language spoken, areas of interest, etc.) through the terminal and registers it with the system. The terminal can also use speech recognition technology to convert the user's voice into text, thereby enabling multilingual translation functionality.

[0539] The server acts as a central processing unit, receiving information sent by the user and storing it in a database. Using a generative AI model, the server generates dynamic travel plans based on the user's interests and input information. The generated travel plans are delivered to the user's device in real time, offering suggestions for tourist destinations and activities. The server also retrieves local regulations and cultural information from the database and provides it to the user as needed. Furthermore, in emergencies, it identifies the nearest support organization based on location information sent from the device and notifies the user of that information.

[0540] For example, if a user wishes to experience culture in Tokyo during a visit to Japan, they would input their preference into their device. The server would then analyze this data and, using a generative AI model, suggest relevant tourist destinations and activities, such as festivals or traditional theatrical performances in Asakusa. The user would receive this information on their device and be able to plan their trip. Furthermore, the multilingual translation function would enable smoother communication during their visit.

[0541] Examples of prompt statements include the following:

[0542] "Could you recommend a temple-hopping itinerary in Kyoto?"

[0543] "Please tell me about regional cuisine that I can experience in Tokyo."

[0544] These configurations and functions provide support to foreign travelers to enhance their travel experience within Japan, ensuring it is safe and efficient.

[0545] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0546] Step 1:

[0547] Users activate a device such as a smartphone, access a dedicated application, and enter their personal information. This information includes name, nationality, language spoken, and areas of interest. This data is necessary for customizing travel plans, and the entered data is transmitted from the device to a server.

[0548] Step 2:

[0549] The server stores user information received from the terminal in a database. At this point, a user profile is generated based on the entered name, nationality, and areas of interest. The stored data forms the basis for generating user-specific travel plans.

[0550] Step 3:

[0551] The server uses a generative AI model to generate dynamic travel plans based on stored user information. Input includes the user's interests and planned destinations. Based on this information, the server references tourist spots and activities in a database to construct the optimal plan. The generated travel plan is sent to the user's device and presented to them.

[0552] Step 4:

[0553] The device provides multilingual translation capabilities to support communication in the local area. Users input either voice or text through the application. In the case of voice input, the device uses speech recognition technology to convert it into text. The converted text is sent to a server and translated into the specified language. The translation result is returned to the device, and the user uses it in conversations in the local area.

[0554] Step 5:

[0555] When a user needs information about local rules and culture, they can query the server from their device. The server retrieves the relevant information from its database and sends it to the device. This information helps users avoid cultural misunderstandings in their local area.

[0556] Step 6:

[0557] In an emergency, the user activates the "Emergency Assistance" function on their device. The device immediately sends the user's location information to the server. Based on this location information, the server identifies the most appropriate support organization (police station, hospital, etc.) and provides that information to the device. This allows the user to receive the necessary assistance quickly.

[0558] (Application Example 1)

[0559] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0560] When foreign tourists visit Japan, their travel experience is often limited due to language barriers and a lack of cultural understanding. Furthermore, emergencies during their trip and a lack of up-to-date tourist information are also challenges. A system is needed that comprehensively addresses these issues and provides tourists with a better experience of staying in Japan.

[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0562] In this invention, the server includes means for inputting user information and creating a travel itinerary, means for performing multilingual translation, and means for dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests. This enables travelers to obtain appropriate sightseeing information in real time and enjoy their trip with peace of mind, overcoming language and cultural barriers.

[0563] "User information" refers to data including personal information about travelers, their interests, and their accommodation plans, and is used to personalize their travel experience.

[0564] A "travel itinerary" is a plan of sightseeing and activities suggested based on the user's input information, and is a schedule that dynamically changes according to the traveler's preferences and current location.

[0565] "Multilingual translation" is a function that supports communication between users who speak different languages. It is a technology that converts speech and text into a specified language to facilitate the transmission of information.

[0566] A "tour route" is a recommended travel path to tourist spots based on the user's interests and current location, serving as a guide for travelers to efficiently enjoy sightseeing.

[0567] A "local attraction" is a nearby place of cultural or historical significance that travelers should visit, offering a unique local experience.

[0568] "Dynamic suggestion" refers to updating and providing content based on real-time data and conditions, a method of always providing travelers with the most up-to-date information and options.

[0569] The "quiz function" is a feature that encourages users to learn about the culture and customs of their destination in a fun way, and it is a tool that provides information in an interactive format.

[0570] In an embodiment of this invention, the system includes a server, a terminal, and a user as its main components. The server receives information entered by the user through the terminal and performs travel planning, multilingual translation, and dynamic updates of sightseeing suggestions. Specifically, the server processes user information and suggests sightseeing routes and events in real time based on interests and current location. It also provides multilingual translation using voice and text to support communication on-site.

[0571] The hardware primarily consists of the user's smartphone, utilizing GPS and microphone functions for voice input and location information acquisition. Furthermore, a powerful cloud service with a robust database and processing capabilities is used on the server side. For software, Google Cloud Speech-to-Text is used for speech recognition technology, and Microsoft Translator for translation services, enabling real-time information processing.

[0572] For example, when a user visits a smart city, they can launch the app on their device, use the translation function, and learn about the local culture and customs while sightseeing. For instance, if a traveler asks, "What are some recommended tourist attractions nearby?", the system will suggest the best route and locations.

[0573] An example of a prompt for a generative AI model is, "Please recommend some sightseeing routes during my trip. I'd also like to know about local events."

[0574] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0575] Step 1:

[0576] The user launches the application using a terminal and enters basic user information (e.g., name, nationality, language, areas of interest). The server receives this information and stores it in a database. This process builds the foundational data needed to prepare the most suitable travel plan for the user. The input is the user's basic information, and the output is the database record in which that information is stored.

[0577] Step 2:

[0578] The server generates sightseeing routes and suggested activities based on stored user information and regional information of the travel destination. Using a generation AI model, it selects tourist destinations and activities that match the user's interests and proposes an optimized route. The input is user information and regional data, and the output is a suggested sightseeing plan. By using prompts, users can instruct the AI ​​with questions such as, "Please tell me your recommended sightseeing route during my trip," to generate a specific plan.

[0579] Step 3:

[0580] The user utilizes GPS functionality through their device to obtain real-time location information. The server uses this location information to list nearby tourist attractions and event information, and sends dynamically updated information to the device. The input is the current GPS data, and the output is information about nearby tourist attractions. Specifically, the system periodically updates nearby attractions in response to changes in the user's location.

[0581] Step 4:

[0582] The terminal converts voice input to text, and the server uses this text to perform multilingual translation. Voice data from the terminal's microphone is processed by Google Cloud Speech-to-Text, converted to text, and then translated by Microsoft Translator before being provided to the user in multiple languages. Input is voice data, and output is translated text. Users can use the translation function to communicate smoothly with local people.

[0583] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0584] This invention relates to a system that supports the travel experience of foreign tourists within Japan, incorporating an emotion engine that recognizes the user's emotions and reflects the recognition results in travel plans and other services. This system mainly consists of terminals, servers, and users, and aims to improve the traveler's experience through the emotion engine.

[0585] The terminal is a device such as a traveler's smartphone, which interacts with the user through a multi-functional application. The user launches the app, registers personal information, and then interacts with the system during subsequent use. The terminal is equipped with an interface for receiving voice and text input, through which the user's emotions are captured.

[0586] The server is responsible for receiving data transmitted from the terminal and storing it in a database. Furthermore, it uses an emotion engine to analyze the user's emotional state from the received audio data and text. Based on this emotion recognition, the server adjusts travel plan suggestions and selects activities that match the user's interests and emotions. In addition, during translation, it controls tone and word choice according to the user's emotions, enabling more personalized communication.

[0587] For example, if a user expresses feelings of anxiety or confusion, the server might suggest relaxing places or activities that provide a sense of security. For instance, it might enhance recommendations for strolling through a Japanese garden or visiting a hot spring facility. Conversely, if a user is excited, the server might plan activities that lead them to more active tourist destinations or activities.

[0588] By combining these emotional engines, it becomes possible to personalize the user's travel experience and increase their satisfaction. This system will be an extremely useful tool in helping travelers understand and enjoy Japanese culture more deeply.

[0589] The following describes the processing flow.

[0590] Step 1:

[0591] Users launch the app on their smartphones, enter the required personal information, and register. After registration, they answer questions to create a travel plan.

[0592] Step 2:

[0593] The terminal processes the user's input data, encrypts it, and sends it to the server. This data includes travel itinerary, destination, and activities of interest.

[0594] Step 3:

[0595] The server stores the received user information in a database and creates a user profile. This lays the foundation for providing personalized services.

[0596] Step 4:

[0597] Users use voice input to communicate emotions to their device or provide some kind of feedback within the app.

[0598] Step 5:

[0599] The device acquires audio data and converts it to text using speech recognition technology. The acquired data is then sent to a server for sentiment analysis.

[0600] Step 6:

[0601] The server uses an emotion engine to analyze the user's emotions from the received voice or text data. This analysis determines, for example, whether the user is excited or seeking relaxation.

[0602] Step 7:

[0603] The server uses the results of sentiment analysis to adjust the travel plan. For example, if the user is feeling anxious, it recommends relaxing activities; if they are excited, it recommends active activities.

[0604] Step 8:

[0605] The terminal displays a pre-arranged travel plan from the server to the user and provides detailed information about selected activities and tourist destinations.

[0606] Step 9:

[0607] If a user needs multilingual translation while on-site, they can input voice or text into their device.

[0608] Step 10:

[0609] The terminal transmits the input voice or text to the server, which then generates a translation based on the user's sentiment.

[0610] Step 11:

[0611] The translation results, which take emotions into account, are sent back to the device, and the device displays or outputs the results to the user as audio. This enables communication in an appropriate tone.

[0612] Step 12:

[0613] In an emergency, the user activates the "emergency assistance" function. At this time, the device sends the user's location information to the server.

[0614] Step 13:

[0615] The server analyzes location information to identify the nearest support organization.

[0616] Step 14:

[0617] Information about support organizations is transmitted to the device and displayed. This allows users to receive appropriate support quickly.

[0618] (Example 2)

[0619] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0620] Currently, many travel support systems provide travel plans without considering the individual feelings of users, and therefore do not adequately address the alleviation of anxiety or the stimulation of interest in cross-cultural environments. Furthermore, translation services also lack the ability to adjust expressions to reflect emotions, which hinders the improvement of communication quality.

[0621] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0622] In this invention, the server includes means for inputting user information and creating a travel plan based on sentiment analysis, means for adjusting the translation expression according to the user's emotions when performing multilingual translation, and means for analyzing the user's emotional state using voice input or text input. This enables the provision of a personalized travel plan that responds to the user's emotions and facilitates more natural and approachable communication.

[0623] "User information" refers to personal data that travelers provide to the system, which is used for creating travel plans and sentiment analysis.

[0624] "Emotional analysis" is the process of identifying a user's emotional state based on voice and text data they input.

[0625] A "travel plan" is a recommendation of tourist destinations and activities suggested based on the user's interests, emotional state, and accommodation plans.

[0626] "Multilingual translation" is the process of converting text or audio into other languages ​​while preserving meaning between them.

[0627] "Adjusting the expression" refers to modifying the translated content to use appropriate tone and expression based on the user's emotional state.

[0628] "Voice input" is an input method in which the system receives and analyzes the voice that the user speaks into the device.

[0629] "Text input" refers to a method in which a user enters text using an input device such as a keyboard, and the system receives that text.

[0630] "Local rules and customs" refers to information about the laws, culture, and social manners of the travel destination.

[0631] An "emergency support agency" is a service provider or organization that travelers can contact if they encounter difficulties.

[0632] This invention is a system that provides a personalized travel experience tailored to the traveler's emotions through interaction between terminals, servers, and users.

[0633] Terminal:

[0634] Users utilize a multi-functional application installed on devices such as smartphones and tablets. First, users launch the application and register their personal information. Input methods include voice input using speech recognition software and text input via a text box, through which users communicate their mood and emotions to the device. Specifically, a user can communicate their current feelings to the system by voice-inputting, "I'm a little tired today."

[0635] server:

[0636] Data sent from the device is received by the server. The server uses a generative AI model capable of advanced natural language processing to analyze the user's emotions from the data. This analysis involves converting the audio data into text and identifying emotions using an emotion engine. Based on the analysis results, the server compares the travel database with the user's current situation to provide the optimal travel plan. It also provides translation assistance, adjusting multilingual translations to match the user's emotions. For example, for a user who needs to relax, it might provide a softer translation such as, "This activity is perfect for refreshing yourself."

[0637] User:

[0638] Users can review the provided travel plan and provide feedback via their device. This feedback allows the system to provide even more personalized services.

[0639] Example of a prompt:

[0640] The prompt message would be something like, "The user said, 'I feel a little anxious.' What is the emotional state based on this?" This would lay the foundation for obtaining an appropriate sentiment analysis result.

[0641] Through this invention, travelers can quickly receive travel plans that suit their own feelings, thereby increasing their sense of security and comfort while traveling.

[0642] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0643] Step 1:

[0644] The user launches the application on their device and registers their personal information. They then input their mood or feelings for the day using voice input or a text box. For example, if the user voice-inputs "I'm a little tired today," the device uses voice recognition software to convert this voice into text data. The input is voice data, and the output is text data.

[0645] Step 2:

[0646] The terminal sends the converted text data to the server. During this process, the data is transmitted via a secure communication protocol, ensuring data integrity and privacy. The input is text data, and the output is a data packet sent to the server.

[0647] Step 3:

[0648] The server inputs the received text data into a generative AI model to analyze the user's emotional state. The data is processed by the emotion engine, and the generative AI model performs emotion analysis. During this process, a prompt is issued, such as "The user said 'I'm a little tired.' What is the emotional state based on this?", and the output is a label of the user's emotional state.

[0649] Step 4:

[0650] The server generates an optimal travel plan for the user based on the analysis results. It refers to a travel database and selects tourist destinations and activities that match the user's emotional state. For example, if the server determines the user is tired, it might suggest visiting a relaxing hot spring facility or a quiet garden. The input is an emotional state label, and the output is a customized travel plan.

[0651] Step 5:

[0652] The server sends the generated travel plan to the terminal and simultaneously performs multilingual translation, adjusting the translation expression according to the user's emotions. Translation assistance is used to provide explanations using language that matches the emotions. The input is the travel plan, and the output is plan information including the adjusted translated text.

[0653] Step 6:

[0654] Users review the suggested travel plans on their devices and provide feedback as needed. This feedback is sent back to the server and used as data for future plan generation. The input is user feedback information, and the output is user profile information updated by the server.

[0655] In this way, a series of steps provides the user with the optimal travel experience.

[0656] (Application Example 2)

[0657] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0658] When foreign tourists travel within Japan, they often face anxieties and communication difficulties stemming from language barriers and cultural differences. While conventional travel support systems offer travel plan suggestions and multilingual translation, they have struggled to provide personalized experiences that respond to travelers' immediate emotions and interests. Therefore, there is a need to accurately recognize travelers' emotions and reflect them in travel plans in real time.

[0659] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0660] In this invention, the server includes means for inputting user information and creating a travel plan, means for performing multilingual translation, and means for analyzing the user's emotional state and adjusting the travel plan based on the analysis results. This makes it possible to suggest sightseeing activities that are in line with the traveler's emotions and to use appropriate language expressions that correspond to those emotions.

[0661] "Foreign tourists" refers to travelers who come from outside Japan and whose purpose is to experience Japanese tourist attractions and culture.

[0662] An "information processing system to support travel experiences" is a system that includes information processing means to provide information and support so that travelers can have a smooth and fulfilling sightseeing experience.

[0663] "Means for inputting user information and creating travel plans" refers to data input and processing means for generating appropriate sightseeing plans based on the traveler's interests and schedule.

[0664] "Means of performing multilingual translation" refers to means of translating between the languages ​​of different countries, enabling travelers to communicate across language barriers.

[0665] "Means for analyzing a user's emotional state and adjusting travel plans based on the analysis results" refers to processing methods that analyze a traveler's emotions and, accordingly, propose the most suitable travel activities and plans.

[0666] "A means of suggesting tourism activities based on the user's current emotions in real time" refers to a means that captures changes in a traveler's emotions in real time and immediately presents suitable tourism activities.

[0667] "Means for generating language expressions that respond to the user's emotional state" refers to means for adjusting translated content to more appropriate and natural language expressions that match the traveler's emotions.

[0668] The system implementing this invention includes a personal information terminal carried by the traveler, a cloud-based information processing server, and a user interface to facilitate user interaction. Primarily, the personal information terminal receives voice and text input from the traveler and relays it to the server. The traveler uses a travel support application to register personal information and confirm travel plans.

[0669] When the server receives data transmitted from a mobile device, it first accesses a database to manage the traveler's basic information. For sentiment analysis, it uses the Google Cloud Natural Language API to analyze the traveler's emotional state from voice and text data. Based on this analysis, the server adjusts travel plan suggestions and selects activities and tourist destinations that match the traveler's emotional state.

[0670] Furthermore, the server overcomes language barriers using a multilingual translation AI engine to improve the naturalness of communication. Specifically, it adaptively adjusts tone and wording by referencing the traveler's emotions during translation, enabling more personalized communication.

[0671] As a concrete example, when a traveler is visiting multiple tourist spots in Tokyo, if the system analyzes the user's emotions suggesting they are "tired," the server will suggest quiet gardens or relaxing cafes. Conversely, if the user indicates they are feeling energetic, the system will guide them to lively shopping malls or events. As an example of a prompt to the generating AI model, inputting "Please suggest tourist spots recommended when the user's emotion is 'excited'" will cause the AI ​​to generate appropriate suggestions.

[0672] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0673] Step 1:

[0674] Users launch the application using their mobile devices and enter personal information. This input data, including the traveler's preferences and basic travel requests, is sent from the device to the server. The server receives this data and stores it in a database to manage information for each traveler.

[0675] Step 2:

[0676] When a user inputs voice or text through the application, the device sends this data to the server in real time. The server uses the Google Cloud Natural Language API to analyze this data and identify the traveler's emotional state. The emotional state is categorized into states such as excited, relaxed, or anxious, and the analysis results are stored on the server.

[0677] Step 3:

[0678] The server utilizes a generative AI model based on the results of emotion analysis to adjust travel plans according to the user's interests. For example, if an excited state is detected, the server uses the AI ​​model to generate a plan that suggests active tourist destinations and events. This generated plan is sent to the device and presented to the user within the application.

[0679] Step 4:

[0680] The server works in conjunction with a multilingual translation engine to translate the output text from the device according to the user's emotional state. The text is converted to an appropriate tone and phrasing, and the result is displayed to the user. Specifically, it provides communication optimized for travelers by increasing politeness or adjusting to more casual expressions.

[0681] Step 5:

[0682] Based on the presented travel plan and translated information, the user selects their next action. By inputting new voice or text data again, sentiment analysis is performed again, and new prompt sentences can be generated. For example, if the user inputs "Please provide activities that are best suited to my current mood," appropriate suggestions will be generated.

[0683] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0684] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0685] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0686] [Fourth Embodiment]

[0687] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0688] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0689] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0690] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0691] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0692] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0693] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0694] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0695] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0696] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0697] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0698] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0699] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0700] This invention provides a system to support the travel experience of foreign tourists visiting Japan. This system consists of three main components: a terminal, a server, and a user.

[0701] The terminal is the traveler's smartphone or other device, through which the traveler interacts with the system. The user uses the terminal to register personal information before the start of their trip, including information such as name, nationality, language spoken, and areas of interest.

[0702] The server receives information sent by users and stores it in a database. This lays the foundation for providing personalized travel plans tailored to travelers. The server has the function of suggesting tourist spots and activities based on the traveler's interests and information about the area they are staying in.

[0703] Furthermore, the device provides a multilingual translation function to help users communicate more smoothly in the local area. Users can use the translation function through the device to convert voice and text into other languages. In this process, speech recognition technology is used to process the input, and the server generates the translated data and sends it back to the device.

[0704] Regarding explanations of local rules, when a user asks a question about local regulations or customs, the server looks up information in the database and provides the device with the appropriate information. This helps avoid cultural misunderstandings and ensures that local manners are observed.

[0705] Furthermore, in an emergency, the user can activate the "emergency assistance" function, and the device immediately sends its location information to a server. Based on that location, the server quickly provides information on the nearest police station, hospital, or other appropriate assistance organization. This information is then presented to the user through the device.

[0706] As a concrete example, consider a scenario where a user visits a major Japanese city and creates a plan to maximize their local cultural experience. Before starting their trip, the user inputs their areas of interest and activities through the app. The server analyzes this input and provides recommended plans, such as temple tours in Kyoto or local cuisine experiences in Tokyo. This allows the user to make the most of their time during their stay.

[0707] The following describes the processing flow.

[0708] Step 1:

[0709] The user launches the application on their device and opens the "New Registration" screen. They enter personal information such as their name, nationality, language settings, and areas of interest, and then press the submit button.

[0710] Step 2:

[0711] The device collects data entered by the user, encrypts it, and then sends it to the server.

[0712] Step 3:

[0713] The server analyzes the received user data and stores it in a database. This prepares the server to create travel plans based on the user's individual needs.

[0714] Step 4:

[0715] The user selects the "Create a travel plan" option and enters their interests and desired tourist destinations.

[0716] Step 5:

[0717] The device sends the user's requested information to the server.

[0718] Step 6:

[0719] The server uses AI to suggest optimal tourist spots and activities based on the user's interests and revised plan. These suggestions are then sent to the user's device.

[0720] Step 7:

[0721] The terminal displays suggestions sent from the server to the user, allowing the user to review the details and make modifications as needed.

[0722] Step 8:

[0723] If a user needs to translate a conversation while on location, they can use the app's translation function.

[0724] Step 9:

[0725] The device receives voice input and converts the speech into text.

[0726] Step 10:

[0727] Text data is sent to the server and translated into the specified language.

[0728] Step 11:

[0729] The translated text is sent back to the terminal from the server and displayed or outputted to the user as audio.

[0730] Step 12:

[0731] When a user seeks information about local rules and customs, they enter relevant questions into the app.

[0732] Step 13:

[0733] The terminal sends the entered question to the server.

[0734] Step 14:

[0735] The server searches the database and sends instructions for the appropriate local rules to the terminal.

[0736] Step 15:

[0737] The terminal displays guidance information sent from the server to the user.

[0738] Step 16:

[0739] In an emergency, the user presses the "Emergency Assistance" button.

[0740] Step 17:

[0741] The device sends data, including its current location, to the server.

[0742] Step 18:

[0743] The server analyzes the location information to identify the nearest police station or hospital.

[0744] Step 19:

[0745] Information about support organizations is sent to the terminal and presented to the user.

[0746] (Example 1)

[0747] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0748] When foreign tourists visit Japan, their travel experiences may be limited due to individual interests or language differences. Furthermore, they may not have sufficient information about local culture and regulations, leading to misunderstandings. In addition, they may face difficulties in responding quickly in emergencies. The challenge is to resolve these issues and enable tourists to have a more fulfilling experience.

[0749] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0750] In this invention, the server includes means for receiving user information and generating a dynamic travel plan, means for converting voice input into text using speech recognition technology and performing multilingual translation, and means for retrieving local rules and cultural information from a database and providing it to the user. This allows travelers to receive individually customized travel plans and communicate smoothly across language barriers. Furthermore, by obtaining accurate information about local culture and rules, cultural misunderstandings can be avoided. In addition, based on location information, information on support organizations can be quickly provided in emergencies, supporting a safe and secure trip.

[0751] "User information" refers to basic personal data entered by travelers, including name, nationality, language spoken, and areas of interest.

[0752] A "dynamic travel plan" is a personalized travel plan generated by a server based on the user's interests and planned destinations.

[0753] "Voice recognition technology" is a technology that allows a device to understand the user's voice input and convert it into text data.

[0754] "Multilingual translation" is the process of converting text or audio input in a specific language into another specified language.

[0755] "Local rules and cultural information" refers to information about the culture, laws, and customs of the area that travelers are visiting.

[0756] "Location information" refers to data about the user's current location acquired by the device, and is useful information in emergencies.

[0757] A "generative AI model" is an artificial intelligence model designed to analyze user input data and suggest optimal tourist destinations and activities.

[0758] "Suggestions for tourist destinations and activities" refers to tourist spots and activities that the server recommends based on the user's interests and planned destinations.

[0759] In order to implement this invention, the following system configuration is necessary. It mainly consists of three elements: a terminal, a server, and a user.

[0760] The terminal is a device such as a smartphone or tablet owned by the user. A dedicated application is installed on this terminal, providing an interface for the user to input information and interact with the server. The user enters their personal information (name, nationality, language spoken, areas of interest, etc.) through the terminal and registers it with the system. The terminal can also use speech recognition technology to convert the user's voice into text, thereby enabling multilingual translation functionality.

[0761] The server acts as a central processing unit, receiving information sent by the user and storing it in a database. Using a generative AI model, the server generates dynamic travel plans based on the user's interests and input information. The generated travel plans are delivered to the user's device in real time, offering suggestions for tourist destinations and activities. The server also retrieves local regulations and cultural information from the database and provides it to the user as needed. Furthermore, in emergencies, it identifies the nearest support organization based on location information sent from the device and notifies the user of that information.

[0762] For example, if a user wishes to experience culture in Tokyo during a visit to Japan, they would input their preference into their device. The server would then analyze this data and, using a generative AI model, suggest relevant tourist destinations and activities, such as festivals or traditional theatrical performances in Asakusa. The user would receive this information on their device and be able to plan their trip. Furthermore, the multilingual translation function would enable smoother communication during their visit.

[0763] Examples of prompt statements include the following:

[0764] "Could you recommend a temple-hopping itinerary in Kyoto?"

[0765] "Please tell me about regional cuisine that I can experience in Tokyo."

[0766] These configurations and functions provide support to foreign travelers to enhance their travel experience within Japan, ensuring it is safe and efficient.

[0767] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0768] Step 1:

[0769] Users activate a device such as a smartphone, access a dedicated application, and enter their personal information. This information includes name, nationality, language spoken, and areas of interest. This data is necessary for customizing travel plans, and the entered data is transmitted from the device to a server.

[0770] Step 2:

[0771] The server stores user information received from the terminal in a database. At this point, a user profile is generated based on the entered name, nationality, and areas of interest. The stored data forms the basis for generating user-specific travel plans.

[0772] Step 3:

[0773] The server uses a generative AI model to generate dynamic travel plans based on stored user information. Input includes the user's interests and planned destinations. Based on this information, the server references tourist spots and activities in a database to construct the optimal plan. The generated travel plan is sent to the user's device and presented to them.

[0774] Step 4:

[0775] The device provides multilingual translation capabilities to support communication in the local area. Users input either voice or text through the application. In the case of voice input, the device uses speech recognition technology to convert it into text. The converted text is sent to a server and translated into the specified language. The translation result is returned to the device, and the user uses it in conversations in the local area.

[0776] Step 5:

[0777] When a user needs information about local rules and culture, they can query the server from their device. The server retrieves the relevant information from its database and sends it to the device. This information helps users avoid cultural misunderstandings in their local area.

[0778] Step 6:

[0779] In an emergency, the user activates the "Emergency Assistance" function on their device. The device immediately sends the user's location information to the server. Based on this location information, the server identifies the most appropriate support organization (police station, hospital, etc.) and provides that information to the device. This allows the user to receive the necessary assistance quickly.

[0780] (Application Example 1)

[0781] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0782] When foreign tourists visit Japan, their travel experience is often limited due to language barriers and a lack of cultural understanding. Furthermore, emergencies during their trip and a lack of up-to-date tourist information are also challenges. A system is needed that comprehensively addresses these issues and provides tourists with a better experience of staying in Japan.

[0783] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0784] In this invention, the server includes means for inputting user information and creating a travel itinerary, means for performing multilingual translation, and means for dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests. This enables travelers to obtain appropriate sightseeing information in real time and enjoy their trip with peace of mind, overcoming language and cultural barriers.

[0785] "User information" refers to data including personal information about travelers, their interests, and their accommodation plans, and is used to personalize their travel experience.

[0786] A "travel itinerary" is a plan of sightseeing and activities suggested based on the user's input information, and is a schedule that dynamically changes according to the traveler's preferences and current location.

[0787] "Multilingual translation" is a function that supports communication between users who speak different languages. It is a technology that converts speech and text into a specified language to facilitate the transmission of information.

[0788] A "tour route" is a recommended travel path to tourist spots based on the user's interests and current location, serving as a guide for travelers to efficiently enjoy sightseeing.

[0789] A "local attraction" is a nearby place of cultural or historical significance that travelers should visit, offering a unique local experience.

[0790] "Dynamic suggestion" refers to updating and providing content based on real-time data and conditions, a method of always providing travelers with the most up-to-date information and options.

[0791] The "quiz function" is a feature that encourages users to learn about the culture and customs of their destination in a fun way, and it is a tool that provides information in an interactive format.

[0792] In an embodiment of this invention, the system includes a server, a terminal, and a user as its main components. The server receives information entered by the user through the terminal and performs travel planning, multilingual translation, and dynamic updates of sightseeing suggestions. Specifically, the server processes user information and suggests sightseeing routes and events in real time based on interests and current location. It also provides multilingual translation using voice and text to support communication on-site.

[0793] The hardware primarily consists of the user's smartphone, utilizing GPS and microphone functions for voice input and location information acquisition. Furthermore, a powerful cloud service with a robust database and processing capabilities is used on the server side. For software, Google Cloud Speech-to-Text is used for speech recognition technology, and Microsoft Translator for translation services, enabling real-time information processing.

[0794] For example, when a user visits a smart city, they can launch the app on their device, use the translation function, and learn about the local culture and customs while sightseeing. For instance, if a traveler asks, "What are some recommended tourist attractions nearby?", the system will suggest the best route and locations.

[0795] An example of a prompt for a generative AI model is, "Please recommend some sightseeing routes during my trip. I'd also like to know about local events."

[0796] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0797] Step 1:

[0798] The user launches the application using a terminal and enters basic user information (e.g., name, nationality, language, areas of interest). The server receives this information and stores it in a database. This process builds the foundational data needed to prepare the most suitable travel plan for the user. The input is the user's basic information, and the output is the database record in which that information is stored.

[0799] Step 2:

[0800] The server generates sightseeing routes and suggested activities based on stored user information and regional information of the travel destination. Using a generation AI model, it selects tourist destinations and activities that match the user's interests and proposes an optimized route. The input is user information and regional data, and the output is a suggested sightseeing plan. By using prompts, users can instruct the AI ​​with questions such as, "Please tell me your recommended sightseeing route during my trip," to generate a specific plan.

[0801] Step 3:

[0802] The user utilizes GPS functionality through their device to obtain real-time location information. The server uses this location information to list nearby tourist attractions and event information, and sends dynamically updated information to the device. The input is the current GPS data, and the output is information about nearby tourist attractions. Specifically, the system periodically updates nearby attractions in response to changes in the user's location.

[0803] Step 4:

[0804] The terminal converts voice input to text, and the server uses this text to perform multilingual translation. Voice data from the terminal's microphone is processed by Google Cloud Speech-to-Text, converted to text, and then translated by Microsoft Translator before being provided to the user in multiple languages. Input is voice data, and output is translated text. Users can use the translation function to communicate smoothly with local people.

[0805] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0806] This invention relates to a system that supports the travel experience of foreign tourists within Japan, incorporating an emotion engine that recognizes the user's emotions and reflects the recognition results in travel plans and other services. This system mainly consists of terminals, servers, and users, and aims to improve the traveler's experience through the emotion engine.

[0807] The terminal is a device such as a traveler's smartphone, which interacts with the user through a multi-functional application. The user launches the app, registers personal information, and then interacts with the system during subsequent use. The terminal is equipped with an interface for receiving voice and text input, through which the user's emotions are captured.

[0808] The server is responsible for receiving data transmitted from the terminal and storing it in a database. Furthermore, it uses an emotion engine to analyze the user's emotional state from the received audio data and text. Based on this emotion recognition, the server adjusts travel plan suggestions and selects activities that match the user's interests and emotions. In addition, during translation, it controls tone and word choice according to the user's emotions, enabling more personalized communication.

[0809] For example, if a user expresses feelings of anxiety or confusion, the server might suggest relaxing places or activities that provide a sense of security. For instance, it might enhance recommendations for strolling through a Japanese garden or visiting a hot spring facility. Conversely, if a user is excited, the server might plan activities that lead them to more active tourist destinations or activities.

[0810] By combining these emotional engines, it becomes possible to personalize the user's travel experience and increase their satisfaction. This system will be an extremely useful tool in helping travelers understand and enjoy Japanese culture more deeply.

[0811] The following describes the processing flow.

[0812] Step 1:

[0813] Users launch the app on their smartphones, enter the required personal information, and register. After registration, they answer questions to create a travel plan.

[0814] Step 2:

[0815] The terminal processes the user's input data, encrypts it, and sends it to the server. This data includes travel itinerary, destination, and activities of interest.

[0816] Step 3:

[0817] The server stores the received user information in a database and creates a user profile. This lays the foundation for providing personalized services.

[0818] Step 4:

[0819] Users use voice input to communicate emotions to their device or provide some kind of feedback within the app.

[0820] Step 5:

[0821] The device acquires audio data and converts it to text using speech recognition technology. The acquired data is then sent to a server for sentiment analysis.

[0822] Step 6:

[0823] The server uses an emotion engine to analyze the user's emotions from the received voice or text data. This analysis determines, for example, whether the user is excited or seeking relaxation.

[0824] Step 7:

[0825] The server uses the results of sentiment analysis to adjust the travel plan. For example, if the user is feeling anxious, it recommends relaxing activities; if they are excited, it recommends active activities.

[0826] Step 8:

[0827] The terminal displays a pre-arranged travel plan from the server to the user and provides detailed information about selected activities and tourist destinations.

[0828] Step 9:

[0829] If a user needs multilingual translation while on-site, they can input voice or text into their device.

[0830] Step 10:

[0831] The terminal transmits the input voice or text to the server, which then generates a translation based on the user's sentiment.

[0832] Step 11:

[0833] The translation results, which take emotions into account, are sent back to the device, and the device displays or outputs the results to the user as audio. This enables communication in an appropriate tone.

[0834] Step 12:

[0835] In an emergency, the user activates the "emergency assistance" function. At this time, the device sends the user's location information to the server.

[0836] Step 13:

[0837] The server analyzes location information to identify the nearest support organization.

[0838] Step 14:

[0839] Information about support organizations is transmitted to the device and displayed. This allows users to receive appropriate support quickly.

[0840] (Example 2)

[0841] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0842] Currently, many travel support systems provide travel plans without considering the individual feelings of users, and therefore do not adequately address the alleviation of anxiety or the stimulation of interest in cross-cultural environments. Furthermore, translation services also lack the ability to adjust expressions to reflect emotions, which hinders the improvement of communication quality.

[0843] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0844] In this invention, the server includes means for inputting user information and creating a travel plan based on sentiment analysis, means for adjusting the translation expression according to the user's emotions when performing multilingual translation, and means for analyzing the user's emotional state using voice input or text input. This enables the provision of a personalized travel plan that responds to the user's emotions and facilitates more natural and approachable communication.

[0845] "User information" refers to personal data that travelers provide to the system, which is used for creating travel plans and sentiment analysis.

[0846] "Emotional analysis" is the process of identifying a user's emotional state based on voice and text data they input.

[0847] A "travel plan" is a recommendation of tourist destinations and activities suggested based on the user's interests, emotional state, and accommodation plans.

[0848] "Multilingual translation" is the process of converting text or audio into other languages ​​while preserving meaning between them.

[0849] "Adjusting the expression" refers to modifying the translated content to use appropriate tone and expression based on the user's emotional state.

[0850] "Voice input" is an input method in which the system receives and analyzes the voice that the user speaks into the device.

[0851] "Text input" refers to a method in which a user enters text using an input device such as a keyboard, and the system receives that text.

[0852] "Local rules and customs" refers to information about the laws, culture, and social manners of the travel destination.

[0853] An "emergency support agency" is a service provider or organization that travelers can contact if they encounter difficulties.

[0854] This invention is a system that provides a personalized travel experience tailored to the traveler's emotions through interaction between terminals, servers, and users.

[0855] Terminal:

[0856] Users utilize a multi-functional application installed on devices such as smartphones and tablets. First, users launch the application and register their personal information. Input methods include voice input using speech recognition software and text input via a text box, through which users communicate their mood and emotions to the device. Specifically, a user can communicate their current feelings to the system by voice-inputting, "I'm a little tired today."

[0857] server:

[0858] Data sent from the device is received by the server. The server uses a generative AI model capable of advanced natural language processing to analyze the user's emotions from the data. This analysis involves converting the audio data into text and identifying emotions using an emotion engine. Based on the analysis results, the server compares the travel database with the user's current situation to provide the optimal travel plan. It also provides translation assistance, adjusting multilingual translations to match the user's emotions. For example, for a user who needs to relax, it might provide a softer translation such as, "This activity is perfect for refreshing yourself."

[0859] User:

[0860] Users can review the provided travel plan and provide feedback via their device. This feedback allows the system to provide even more personalized services.

[0861] Example of a prompt:

[0862] The prompt message would be something like, "The user said, 'I feel a little anxious.' What is the emotional state based on this?" This would lay the foundation for obtaining an appropriate sentiment analysis result.

[0863] Through this invention, travelers can quickly receive travel plans that suit their own feelings, thereby increasing their sense of security and comfort while traveling.

[0864] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0865] Step 1:

[0866] The user launches the application on their device and registers their personal information. They then input their mood or feelings for the day using voice input or a text box. For example, if the user voice-inputs "I'm a little tired today," the device uses voice recognition software to convert this voice into text data. The input is voice data, and the output is text data.

[0867] Step 2:

[0868] The terminal sends the converted text data to the server. During this process, the data is transmitted via a secure communication protocol, ensuring data integrity and privacy. The input is text data, and the output is a data packet sent to the server.

[0869] Step 3:

[0870] The server inputs the received text data into a generative AI model to analyze the user's emotional state. The data is processed by the emotion engine, and the generative AI model performs emotion analysis. During this process, a prompt is issued, such as "The user said 'I'm a little tired.' What is the emotional state based on this?", and the output is a label of the user's emotional state.

[0871] Step 4:

[0872] The server generates an optimal travel plan for the user based on the analysis results. It refers to a travel database and selects tourist destinations and activities that match the user's emotional state. For example, if the server determines the user is tired, it might suggest visiting a relaxing hot spring facility or a quiet garden. The input is an emotional state label, and the output is a customized travel plan.

[0873] Step 5:

[0874] The server sends the generated travel plan to the terminal and simultaneously performs multilingual translation, adjusting the translation expression according to the user's emotions. Translation assistance is used to provide explanations using language that matches the emotions. The input is the travel plan, and the output is plan information including the adjusted translated text.

[0875] Step 6:

[0876] Users review the suggested travel plans on their devices and provide feedback as needed. This feedback is sent back to the server and used as data for future plan generation. The input is user feedback information, and the output is user profile information updated by the server.

[0877] In this way, a series of steps provides the user with the optimal travel experience.

[0878] (Application Example 2)

[0879] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0880] When foreign tourists travel within Japan, they often face anxieties and communication difficulties stemming from language barriers and cultural differences. While conventional travel support systems offer travel plan suggestions and multilingual translation, they have struggled to provide personalized experiences that respond to travelers' immediate emotions and interests. Therefore, there is a need to accurately recognize travelers' emotions and reflect them in travel plans in real time.

[0881] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0882] In this invention, the server includes means for inputting user information and creating a travel plan, means for performing multilingual translation, and means for analyzing the user's emotional state and adjusting the travel plan based on the analysis results. This makes it possible to suggest sightseeing activities that are in line with the traveler's emotions and to use appropriate language expressions that correspond to those emotions.

[0883] "Foreign tourists" refers to travelers who come from outside Japan and whose purpose is to experience Japanese tourist attractions and culture.

[0884] An "information processing system to support travel experiences" is a system that includes information processing means to provide information and support so that travelers can have a smooth and fulfilling sightseeing experience.

[0885] "Means for inputting user information and creating travel plans" refers to data input and processing means for generating appropriate sightseeing plans based on the traveler's interests and schedule.

[0886] "Means of performing multilingual translation" refers to means of translating between the languages ​​of different countries, enabling travelers to communicate across language barriers.

[0887] "Means for analyzing a user's emotional state and adjusting travel plans based on the analysis results" refers to processing methods that analyze a traveler's emotions and, accordingly, propose the most suitable travel activities and plans.

[0888] "A means of suggesting tourism activities based on the user's current emotions in real time" refers to a means that captures changes in a traveler's emotions in real time and immediately presents suitable tourism activities.

[0889] "Means for generating language expressions that respond to the user's emotional state" refers to means for adjusting translated content to more appropriate and natural language expressions that match the traveler's emotions.

[0890] The system implementing this invention includes a personal information terminal carried by the traveler, a cloud-based information processing server, and a user interface to facilitate user interaction. Primarily, the personal information terminal receives voice and text input from the traveler and relays it to the server. The traveler uses a travel support application to register personal information and confirm travel plans.

[0891] When the server receives data transmitted from a mobile device, it first accesses a database to manage the traveler's basic information. For sentiment analysis, it uses the Google Cloud Natural Language API to analyze the traveler's emotional state from voice and text data. Based on this analysis, the server adjusts travel plan suggestions and selects activities and tourist destinations that match the traveler's emotional state.

[0892] Furthermore, the server overcomes language barriers using a multilingual translation AI engine to improve the naturalness of communication. Specifically, it adaptively adjusts tone and wording by referencing the traveler's emotions during translation, enabling more personalized communication.

[0893] As a concrete example, when a traveler is visiting multiple tourist spots in Tokyo, if the system analyzes the user's emotions suggesting they are "tired," the server will suggest quiet gardens or relaxing cafes. Conversely, if the user indicates they are feeling energetic, the system will guide them to lively shopping malls or events. As an example of a prompt to the generating AI model, inputting "Please suggest tourist spots recommended when the user's emotion is 'excited'" will cause the AI ​​to generate appropriate suggestions.

[0894] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0895] Step 1:

[0896] Users launch the application using their mobile devices and enter personal information. This input data, including the traveler's preferences and basic travel requests, is sent from the device to the server. The server receives this data and stores it in a database to manage information for each traveler.

[0897] Step 2:

[0898] When a user inputs voice or text through the application, the device sends this data to the server in real time. The server uses the Google Cloud Natural Language API to analyze this data and identify the traveler's emotional state. The emotional state is categorized into states such as excited, relaxed, or anxious, and the analysis results are stored on the server.

[0899] Step 3:

[0900] The server utilizes a generative AI model based on the results of emotion analysis to adjust travel plans according to the user's interests. For example, if an excited state is detected, the server uses the AI ​​model to generate a plan that suggests active tourist destinations and events. This generated plan is sent to the device and presented to the user within the application.

[0901] Step 4:

[0902] The server works in conjunction with a multilingual translation engine to translate the output text from the device according to the user's emotional state. The text is converted to an appropriate tone and phrasing, and the result is displayed to the user. Specifically, it provides communication optimized for travelers by increasing politeness or adjusting to more casual expressions.

[0903] Step 5:

[0904] Based on the presented travel plan and translated information, the user selects their next action. By inputting new voice or text data again, sentiment analysis is performed again, and new prompt sentences can be generated. For example, if the user inputs "Please provide activities that are best suited to my current mood," appropriate suggestions will be generated.

[0905] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0906] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0907] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0908] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0909] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0910] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0911] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0912] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0913] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0914] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0915] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0916] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0917] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0918] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0919] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0920] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0921] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0922] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0923] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0924] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0925] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0926] The following is further disclosed regarding the embodiments described above.

[0927] (Claim 1)

[0928] This is a system designed to support the travel experience within Japan for foreign tourists.

[0929] A means of entering user information and creating a travel plan,

[0930] Means for performing multilingual translation,

[0931] Means of providing information on local rules and customs,

[0932] A system that includes means of providing information on the nearest support organization in times of emergency.

[0933] (Claim 2)

[0934] The means for creating the aforementioned travel plan is,

[0935] The system according to claim 1, which suggests tourist destinations and activities based on the user's interests and accommodation plans.

[0936] (Claim 3)

[0937] The means for performing the aforementioned multilingual translation is

[0938] The system according to claim 1, which receives voice input, converts it to text, and translates it into a specified language.

[0939] "Example 1"

[0940] (Claim 1)

[0941] A means of receiving user information and generating a dynamic travel plan,

[0942] A means of converting voice input into text using speech recognition technology and performing multilingual translation,

[0943] A means of retrieving local rules and cultural information from a database and providing it to the user,

[0944] A means of identifying the nearest support organization based on location information and providing information in emergencies,

[0945] A system that includes means of analyzing user interests using generative AI models and suggesting optimal tourist destinations and activities.

[0946] (Claim 2)

[0947] The system according to claim 1, which uses a generative AI model to suggest tourist destinations and activities based on the user's interests and place of stay.

[0948] (Claim 3)

[0949] The system according to claim 1, which uses speech recognition technology to convert speech input into text and translates it into a specified language.

[0950] "Application Example 1"

[0951] (Claim 1)

[0952] A means of entering user information and creating a travel itinerary,

[0953] Means for performing multilingual translation,

[0954] Means of providing information on local rules and customs,

[0955] A means of providing information on the nearest support organization in an emergency,

[0956] A means of dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests,

[0957] A system that includes a means of providing a quiz function about local culture and customs.

[0958] (Claim 2)

[0959] The system according to claim 1, wherein the means for creating the travel itinerary suggests tourist destinations and activities based on the user's interests and accommodation plans, and dynamically updates them in association with location information.

[0960] (Claim 3)

[0961] The system according to claim 1, wherein the means for performing the multilingual translation receives voice input, converts it to text, translates it into a specified language, and provides tourist information and event information in real time.

[0962] "Example 2 of combining an emotion engine"

[0963] (Claim 1)

[0964] A means of inputting user information and creating a travel plan based on sentiment analysis,

[0965] When performing multilingual translation, a means of adjusting the expression of the translation according to the user's emotions,

[0966] Methods for analyzing a user's emotional state using voice input or text input,

[0967] Means of providing information on local rules and customs,

[0968] A system that includes means of providing information on the nearest support organization in times of emergency.

[0969] (Claim 2)

[0970] The system according to claim 1, wherein the means for creating the travel plan proposes tourist destinations and activities based on the user's interests, emotional state and accommodation plans.

[0971] (Claim 3)

[0972] The system according to claim 1, wherein the means for performing the multilingual translation includes receiving voice input, converting it to text, and translating it into a specified language, including sentiment analysis.

[0973] "Application example 2 when combining with an emotional engine"

[0974] (Claim 1)

[0975] An information processing system to support the travel experience of foreign tourists,

[0976] A means of entering user information and creating a travel plan,

[0977] Means for performing multilingual translation,

[0978] Means of providing information on local rules and customs,

[0979] A means of providing information on the nearest support organization in an emergency,

[0980] A means of analyzing the user's emotional state and adjusting the travel plan based on the analysis results,

[0981] A means of suggesting tourism activities based on the user's current emotions in real time,

[0982] A means of generating linguistic expressions that correspond to the user's emotional state.

[0983] A system that includes this.

[0984] (Claim 2)

[0985] The means for creating the aforementioned travel plan is,

[0986] The system according to claim 1, which suggests tourist destinations and activities based on the user's interests and accommodation plans, and further based on the user's emotional state.

[0987] (Claim 3)

[0988] The means for performing the aforementioned multilingual translation is

[0989] The system according to claim 1, which receives voice input, converts it to text, translates it into a specified language, and further adjusts the tone and wording based on the user's emotional state. [Explanation of symbols]

[0990] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of entering user information and creating a travel itinerary, Means for performing multilingual translation, Means of providing information on local rules and customs, A means of providing information on the nearest support organization in an emergency, A means of dynamically suggesting sightseeing routes, local attractions, and events based on the user's current location and interests, A system that includes a means of providing a quiz function about local culture and customs.

2. The system according to claim 1, wherein the means for creating the travel itinerary suggests tourist destinations and activities based on the user's interests and accommodation plans, and dynamically updates them in association with location information.

3. The system according to claim 1, wherein the means for performing the multilingual translation receives voice input, converts it to text, translates it into a specified language, and provides tourist information and event information in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A