system
A system using location data and generative AI provides personalized travel guidance, improving satisfaction by tailoring content to individual interests and learning from user interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-10
- Publication Date
- 2026-06-22
AI Technical Summary
Modern travelers face challenges in accessing personalized travel experiences that cater to their individual interests and past experiences, with limited real-time interaction and insufficient knowledge about local history and culture, leading to suboptimal satisfaction.
A system that utilizes location information, interest and history data, and generative artificial intelligence to provide customized guidance content in real-time, incorporating voice recognition and recording user responses for continuous improvement.
Enables personalized and interactive travel experiences by providing tailored information and learning from user feedback, enhancing traveler satisfaction through continuous refinement.
Smart Images

Figure 2026101304000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Modern travelers can access a variety of information sources, but many of them remain general in content and are difficult to provide a deep travel experience based on individual interests and past experiences. Also, due to limited opportunities to obtain knowledge about local detailed history and culture, a more personalized travel proposal is necessary to improve traveler satisfaction. Furthermore, there is a need for a system that enables real-time interaction during travel and allows users to actively enjoy their trips.
Means for Solving the Problems
[0005] This invention provides a means for acquiring a user's location information and individually filtering information on tourist spots using interest and history data. This makes it possible to generate customized guidance content tailored to the user in real time using a generative artificial intelligence model. Furthermore, by providing this guidance in audio format, an integrated experience utilizing both sight and hearing is realized. In addition, the invention provides a means for providing additional information based on the user's questions and interests through speech recognition, and for recording the user's responses to be used as learning data to improve future guidance. This makes it possible to provide travelers with deeper knowledge and richer experiences, thereby improving their satisfaction.
[0006] A "user" refers to an individual who uses the system to receive travel guide services.
[0007] "Location information" refers to data that indicates the user's current geographical location.
[0008] "Interest and history data" refers to information about the user's preferred categories and past travel history that they have registered in advance.
[0009] A "tourist spot" refers to a geographical or cultural place that travelers visit.
[0010] "Filtering means" refers to a function that performs a process of extracting only the data that matches specific conditions from the acquired information.
[0011] A "generative artificial intelligence model" refers to a computer system that uses machine learning, deep learning, and other techniques to generate information that meets the specific needs of a user.
[0012] "Customized guidance content" refers to information and suggestions that are individually tailored to the user's interests and circumstances.
[0013] "Means of providing information via voice" refers to methods and devices for transmitting generated information to users in voice format.
[0014] "Voice recognition" refers to the technology that analyzes the user's voice input and processes it as text information.
[0015] "Learning data" refers to data based on the user's reactions and history, which is used to improve the quality of the content that the system will generate in the future.
Brief Explanation of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Embodiments for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disk (e.g., hard disk), or magnetic tape, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention relates to a system that provides personalized tourist information to travelers, and the specific form of its implementation is described below. This system consists of a user's terminal, a server, and application software that connects them.
[0038] The user's device obtains location information using GPS or Wi-Fi. The device sends individual data to the server, including interest and history data that the user has registered in advance. This includes, for example, interest categories such as history and food culture, and past visit history.
[0039] Based on the received location information and interest / history data, the server searches for tourist spot information in the database and creates a filtered set of information. Based on this information, the server uses a generative artificial intelligence model to generate user-specific, customized guidance content.
[0040] As a concrete example, if the user is in an ancient Japanese capital, the server will prioritize extracting information on nearby historical temples and gardens, and use artificial intelligence to generate guidance content that includes their background and related anecdotes. This guidance content is then sent to the user's device and provided via voice output.
[0041] Through voice guidance, users can learn details about specific locations, and the device can even take additional questions from the user via voice recognition. This information is continuously transmitted to a server, providing further details tailored to the user's interests, such as the founding date of the temple or information about particularly famous abbots.
[0042] User responses and behavioral history are recorded on the server and used as learning data to create more refined suggestions in future guidance. This allows the system to continuously provide tourist information that is more tailored to each user's individual preferences.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The device obtains location information using GPS or Wi-Fi to determine the user's current location. This location information is essential for subsequent data processing.
[0046] Step 2:
[0047] The device sends pre-set interest categories and past browsing history to the server. This provides the basic data needed to identify information that is highly relevant to the user.
[0048] Step 3:
[0049] The server searches for tourist spots in its database based on location information received and user interest / history data. Spots near the specified location are prioritized for extraction.
[0050] Step 4:
[0051] Based on filtered tourist spot information, the server uses a generative artificial intelligence model to generate customized guide content. This includes the history, culture, and interesting anecdotes of the place.
[0052] Step 5:
[0053] The server generates the guidance content and sends it to the terminal, which then provides it to the user via voice. Through voice guidance, the user can receive information not only visually but also aurally.
[0054] Step 6:
[0055] When a user asks a question or requests additional information in response to voice guidance, the device processes it using speech recognition. The question is converted into text data and sent to the server.
[0056] Step 7:
[0057] The server retrieves relevant information from the database based on the user's question and generates additional guidance. This new information provides detailed answers tailored to the user's interests.
[0058] Step 8:
[0059] The device provides newly generated information to the user via voice, continuing the guidance interaction. User feedback and behavioral data are recorded on the server and used for personalization in the future.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] Traditional tourist information systems provided users with uniform information, lacking sufficient customization to suit individual hobbies and interests. Furthermore, they struggled to offer real-time information based on the user's current location. Additionally, the lack of mechanisms to utilize user feedback and behavioral history for future guidance hindered service improvement based on experience.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes a device for acquiring the user's location, a device for selecting geographical spot data based on the location and the user's interests and history information, and a device for generating personalized guidance information based on the selected data using an artificial intelligence model. This makes it possible to provide timely and personalized tourist information based on each user's interests.
[0065] A "device for acquiring user location" is a device or technology for measuring and collecting location information to detect where a user is currently located.
[0066] "Interest and history information" refers to data based on individual areas of interest and past behavioral history that users have registered in advance.
[0067] A "device for selecting geographical spot data" is equipment or a process for selecting and filtering relevant tourist destination and spot information based on the user's current location and preferences.
[0068] A "device that generates information using an artificial intelligence model" is a system that uses artificial intelligence technology to generate guidance content optimized for a specific user based on the data it receives.
[0069] "Personalized information" refers to travel and sightseeing information that is customized according to the individual user's interests and current circumstances.
[0070] "Feedback and behavioral history" refers to a record of a user's reactions to the information they receive and their subsequent actions.
[0071] This invention is a system that provides users with personalized tourist information in real time. The system mainly consists of terminals, servers, and applications that connect them.
[0072] The device obtains the user's location information using GPS and Wi-Fi. In particular, using mobile devices such as smartphones and tablets allows for more precise location tracking. This enables location-based personalization.
[0073] The user's device sends data to the server, including pre-registered interest categories (e.g., history, art, food culture, etc.) and past browsing history. This data communication uses common internet protocols, and security is ensured by applying data encryption technology as needed.
[0074] The server searches a database based on received location information and user interest and history data. This database contains a large amount of tourist destination information, which the server filters to create the most suitable information set for the user. In this process, a generative AI model is used to generate user-specific guidance information based on the filtered data. The AI model generates text according to the user's interests, providing detailed background information and anecdotes.
[0075] As a concrete example, suppose a user is interested in history and is visiting an ancient Japanese capital. In this case, the server identifies nearby historical sites and cultural landmarks, and uses an AI model to generate guidance content that includes the historical significance and anecdotes behind them. The AI model might receive a prompt like this: "Please explain the history and important events of the historical building closest to your current location."
[0076] The generated guidance information is sent to the user's device and provided to the user via voice output. The device uses speech synthesis technology to convert the generated text into natural language speech, allowing the user to receive the information hands-free.
[0077] Furthermore, if the user provides additional questions via voice, the device can recognize this and resend the data to the server to provide even more detailed information. This allows for a more interactive and deeper understanding of the tourist experience.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The device uses GPS or Wi-Fi to obtain the user's current location.
[0081] Input: GPS sensor or Wi-Fi signal.
[0082] Processing: The terminal uses this data to perform triangulation and location algorithms to calculate the user's precise geographic coordinates.
[0083] Output: Geographic coordinate information such as latitude and longitude.
[0084] Specific operation: The device periodically updates its location information and tracks the user's movements in real time.
[0085] Step 2:
[0086] The device sends user interest and past history data to the server.
[0087] Input: Data on the user's interest categories and past visit history.
[0088] Processing: The terminal packages this data using a secure protocol and sends it to the server.
[0089] Output: Data about the user's interests and history received by the server.
[0090] Specific operation: Encrypt data using transport layer security to prevent unauthorized access by third parties.
[0091] Step 3:
[0092] The server searches the database based on location information and interest / history data received, and filters the information on tourist spots.
[0093] Input: Location information, interest data, history data.
[0094] Processing: The server uses SQL queries or similar database search methods to select relevant tourist destination information.
[0095] Output: A list of tourist spots that match the user's interests.
[0096] Specific operation: The server uses a scoring algorithm to prioritize selecting the most relevant spots.
[0097] Step 4:
[0098] The server uses an AI model to generate customized guidance content.
[0099] Input: Filtered tourist spot information.
[0100] Processing: Input prompt text into the generation AI model and generate user-specific detailed instructions.
[0101] Output: Customized guidance information.
[0102] Specific operation: The AI model utilizes natural language generation technology to create personalized explanatory text as information.
[0103] Step 5:
[0104] The server generates guidance information and sends it to the user's terminal.
[0105] Input: Guidance information.
[0106] Processing: The server generates a data packet and sends it to the terminal using a communication protocol.
[0107] Output: Guidance information received by the terminal.
[0108] Specific actions: To improve the reliability of data transmission, utilize error checking functions.
[0109] Step 6:
[0110] The terminal provides the user with the received guidance content using its voice output function.
[0111] Input: Guidance information.
[0112] Processing: The terminal's speech synthesis engine converts the text information into speech.
[0113] Output: User receives guidance via voice.
[0114] Specific actions: Adjust the tone and pronunciation of the synthesized speech to ensure smoothness and naturalness of the voice.
[0115] Step 7:
[0116] The device uses voice recognition to accept additional questions from the user.
[0117] Input: User's voice question.
[0118] Processing: The speech recognition engine converts the speech into text and sends the question to the server.
[0119] Output: Text data of the question.
[0120] Specific actions: Noise filtering and optimization of the speech model are performed to improve the accuracy of speech recognition.
[0121] Step 8:
[0122] The server will provide further information based on the user's additional questions.
[0123] Input: User question text.
[0124] Processing: The server retrieves relevant information from the database and generates information to answer the user's question.
[0125] Output: Additional information.
[0126] Specific operation: The server uses the AI model again to generate more detailed and relevant information.
[0127] (Application Example 1)
[0128] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0129] In modern urban tourism, travelers struggle to find places that match their interests amidst a vast amount of information. Furthermore, the information provided is often too generalized, resulting in a lack of personalized tourism experiences tailored to individual interests and preferences. Additionally, providing users with relevant information in real time while they are navigating urban areas presents a significant challenge.
[0130] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0131] In this invention, the server includes means for acquiring the user's location information, means for filtering information on tourist spots based on the location information and the user's interest and history data, and means for generating customized guidance content based on the filtered information using a generative artificial intelligence model. This enables the user to receive real-time, personalized tourist information tailored to their individual interests while they are in a city.
[0132] "Means for obtaining user location information" refers to devices or methods that have the function of identifying the user's current location and transmitting that information to a server.
[0133] "Means for filtering tourist spot information based on user interest and history data" refers to a function that selects relevant tourist destination information based on the user's past interests and visit history.
[0134] "Means for generating customized guidance content based on the filtered information using a generative artificial intelligence model" refers to a function that uses AI to process selected tourist information in a way that is suitable for individual users and to create original guidance content.
[0135] "Means of providing generated guidance content to users in audio format" refers to a function that converts text information into audio to convey guidance to users in an easy-to-understand manner.
[0136] "A means of recognizing user input through speech and providing additional information" refers to a function that understands voice-based questions and requests from users and provides information accordingly.
[0137] "Means for recording user responses and using them as learning data" refers to a function that accumulates data on users' actions and responses when using the system, and uses that data to improve future guidance.
[0138] "A means of acquiring and providing real-time tourist information within a city" refers to a function that instantly acquires and provides the most appropriate tourist information for a user as they move around the city.
[0139] "A means of providing additional details via voice based on user questions" refers to a function that can prepare detailed information in response to questions asked by the user and deliver it in voice format.
[0140] To realize this invention, the user's device first obtains its current location information using GPS or Wi-Fi. The device has the function to send this location information, along with the user's previously registered interests and history data, to a server. The server receives this data and searches a tourism database to filter for appropriate spot information.
[0141] Next, the server uses a generative AI model (specifically, OpenAI's ChatGPT API) based on the filtered information to generate customized tourist information tailored to the user. The generated information is then transmitted from the terminal to the user using speech synthesis technology.
[0142] Users can ask questions using voice recognition technology, and their input is sent to the server. The server analyzes the questions, generates corresponding additional information, and provides it to the user again via voice. This allows users to obtain detailed information about tourist destinations.
[0143] Furthermore, the user's device records their responses, and this data is stored on the server and used as learning data to further personalize future guidance.
[0144] Specific example:
[0145] For example, if a user is a tourist interested in Japanese culture and is visiting a specific area of Tokyo, the server will prioritize providing information about traditional events and historical buildings rooted in that area. If the user asks, "What is the history of this building?" during the audio guidance, the server will search for the appropriate information and provide it to the user verbally.
[0146] Examples of prompts for a generative AI model:
[0147] "The user is in the city center. Please briefly describe the nearby historical landmarks."
[0148] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0149] Step 1:
[0150] The user's device uses GPS and Wi-Fi to obtain its current location. This location information is temporarily stored on the device while waiting for the next data transmission. The input is GPS data, and the output is coordinate data represented as location information.
[0151] Step 2:
[0152] The device collects location information and interest / history data previously registered by the user, and sends this information to the server. The input consists of location information and interest / history data, and the output is a request packet sent to the server.
[0153] Step 3:
[0154] The server receives the request packet and analyzes the location information and interest / history data. Based on the analysis results, it filters the relevant spot information from the tourism database. The input is the request packet, and the output is the filtered tourist spot information. Here, as a data processing step, spot information that matches the specified criteria is selected.
[0155] Step 4:
[0156] The server uses a generative artificial intelligence model to generate customized guidance content based on filtered tourist spot information. The input is filtered information, and the output is user-specific guidance content. Specifically, prompts are generated for the AI model.
[0157] Step 5:
[0158] The generated guidance content is converted into an audio file using speech synthesis technology and sent to the terminal. The input is the generated guidance content, and the output is an audio file.
[0159] Step 6:
[0160] The device plays an audio file and provides instructions to the user. The user can listen to this and make further questions or requests by voice. The input is an audio file, and the output is audio data recorded as the user's response.
[0161] Step 7:
[0162] Questions and requests made by the user are converted into text by the device's speech recognition function and sent to the server. The input is the user's voice, and the output is sent to the server as a text message.
[0163] Step 8:
[0164] The server receives user questions transcribed into text via speech recognition, searches for additional information based on the content, and generates further detailed guides using a generative AI model. The input is the user's question, and the output is the additional guidance content.
[0165] Step 9:
[0166] The generated additional informational content is converted back into an audio file and sent to the device. The device then plays this for the user, completing the provision of detailed information. The input is the additional informational content, and the output is the audio guide delivered to the user.
[0167] Step 10:
[0168] User behavior history and responses are recorded on the device and sent to the server's database. This data is used as training data to improve future guidance content. The input is user behavior data, and the output is stored as training data.
[0169] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0170] This invention combines an emotion engine with a system that provides personalized tourist information to users, and a specific embodiment thereof will be described. This system consists of a user's terminal, a server, and an application including the emotion engine.
[0171] First, the device uses GPS or Wi-Fi to obtain the user's location information. Based on this information, it prepares to collect information about the user's current location and nearby tourist attractions.
[0172] Next, the device collects the user's interest and history data and sends it to the server. This data is set by the user in advance and includes areas of interest such as "nature," "history," and "art."
[0173] Based on the received data, the server uses a generative artificial intelligence model to create customized guidance tailored to the user. At this stage, the emotion engine analyzes the user's emotions from their voice and input, and uses this information to appropriately adjust the guidance. For example, if it detects that the user is tired, it optimizes the guidance to recommend relaxing places to visit.
[0174] The generated guidance content is provided to the user as audio via the device. The tone of the guidance is adjusted based on the analysis results from the emotion engine. Furthermore, the device can integrate with the camera function to overlay digital information onto the real-world scenery. This allows users to enjoy an interactive experience that fully utilizes both sight and hearing.
[0175] When a user asks a question or responds to the guidance, the device uses its voice recognition function to analyze the content and send it to the server. The server stores the user's feedback and uses it as training data to improve future guidance. Data acquired by the emotion engine is also used for training, improving the accuracy of personalized guidance.
[0176] For example, if a user is interested in art and is visiting a museum in Tokyo, the emotion engine will sense their joy and excitement and provide highly tailored information and experiences, such as detailed background information related to the exhibits and other recommended spots. This allows users to receive advanced, personalized travel guidance, improving their overall travel satisfaction.
[0177] The following describes the processing flow.
[0178] Step 1:
[0179] The device uses GPS or Wi-Fi to determine the user's current location in order to obtain their location information. This location information serves as basic data for providing tourist information.
[0180] Step 2:
[0181] The device checks the user's pre-set interests and past travel history and sends them to the server. These include areas of interest such as history, nature, and art.
[0182] Step 3:
[0183] The server receives location information and interest / history data, searches the database, and filters out tourist spot information that is most relevant to the user.
[0184] Step 4:
[0185] Based on the information filtered by the server, a generative artificial intelligence model is used to generate customized guidance content. During this process, the emotion engine analyzes the user's voice and input data to determine the user's current emotional state.
[0186] Step 5:
[0187] Based on the analysis results of the emotion engine, the server adjusts the guidance content. For example, if the user is seeking relaxation, it will recommend calm places.
[0188] Step 6:
[0189] The server generates and sends the adjusted guidance content to the terminal, which then provides it to the user as audio. The audio guidance is played back in a tone that matches the user's emotional state.
[0190] Step 7:
[0191] As the user listens to the instructions and asks questions or makes responses, the device uses speech recognition to analyze the content and sends it to the server. This data is recorded as feedback on the responses.
[0192] Step 8:
[0193] The server collects user feedback as training data to improve the personalization accuracy of future guidance. Additionally, the results of the emotion engine's analysis are also stored as data and used for future guidance.
[0194] (Example 2)
[0195] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0196] When travelers choose places to visit and the order in which to visit them, they are only provided with general information, making it difficult to provide guidance tailored to their individual interests and circumstances. Furthermore, it is challenging to optimize the guidance content in real time by reflecting the user's emotions and feedback during their trip.
[0197] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0198] In this invention, the server includes means for acquiring the user's spatial location information, means for extracting information on travel spots, and means for generating personalized guidance content using a generative information processing model. This makes it possible to provide dynamically personalized guidance according to the user's interests and current emotional state.
[0199] "Spatial location information" refers to data used to identify a user's geographical location, and is acquired using technologies such as GPS and Wi-Fi.
[0200] "Interest and history data" refers to information based on a user's past interests and behavioral history, and reflects the user's preferences and interests.
[0201] A "generative information processing model" is an algorithm or method that generates personalized information based on user data, and primarily utilizes AI technology.
[0202] "Personalized guidance" refers to travel and sightseeing recommendations tailored to the user's specific interests and current circumstances.
[0203] "Audio output" refers to the means and devices used to convey information to users as sound, and is primarily done through speech generation technology.
[0204] "Speech recognition" is a technology in which a computer analyzes the voice spoken by a user and interprets it as text data or instructions.
[0205] "Analyzing emotional state" is a process that estimates what emotions the user is currently experiencing, based on their voice and text input.
[0206] An "image acquisition device" is a device such as a camera or video camera used to record real-world scenes and situations in digital format.
[0207] "Information overlay" is a technology that adds digital information to the real-world field of view, helping users intuitively understand their physical environment.
[0208] The embodiments for carrying out this invention are shown below.
[0209] First, the device uses GPS and Wi-Fi to acquire the user's spatial location information. This information is used to determine the user's current geographical location and is part of the criteria for selecting the next tourist spot to visit. The device has a built-in location sensor, and this data is acquired through dedicated application software.
[0210] Next, the user inputs their interests and historical data into the device. This includes pre-set areas of interest and past activity history. Through the application's settings screen, the user can select interests such as "nature," "history," and "art."
[0211] Subsequently, the server receives the user's location information and interest data transmitted from the terminal. The server utilizes a generative information processing model to generate personalized guidance based on this data. The generated guidance is optimized by AI to best suit the user's specific needs. As a concrete example of a prompt, the instruction "List tourist spots that the user is of high interest" is input to the AI model.
[0212] Furthermore, the server uses an emotion analysis engine to analyze voice and text input from the user, adjusting the guidance content according to the user's emotional state. This includes a function that, based on the process of analyzing the emotional state, recommends tourist spots that are appropriate for a user who wants to relax.
[0213] The generated guidance content is delivered to the user through the terminal via natural sound output. This allows the user to intuitively plan their trip while receiving voice guidance. Furthermore, the terminal uses a video acquisition device to overlay information onto the real-world scenery, enabling the user to receive information more interactively.
[0214] This format provides a dynamic sightseeing guide experience tailored to each user's individual interests and emotional state, while also utilizing user feedback as learning data, ensuring that the accuracy of the guide content improves over time.
[0215] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0216] Step 1:
[0217] The device acquires the user's spatial location information via GPS or Wi-Fi. Specifically, it collects current geographic location data using a built-in location sensor. This input data is processed through the device application and output as the user's coordinate information. This location information is useful as basic data when selecting tourist spots.
[0218] Step 2:
[0219] The user enters their interests and history data into the device. In the application's settings screen, they select categories of interest (e.g., "Nature," "History," "Art"). This information is stored in the system as user preference data and transferred to the server as input data.
[0220] Step 3:
[0221] The server uses a generative AI model to generate customized guidance content based on the user's location information and interest data received from the terminal. The AI model processes the input data as prompts and extracts and generates information that best suits the user's needs. As output, a user-optimized list of tourist information is generated.
[0222] Step 4:
[0223] The server analyzes the user's input (voice and text) using an emotion analysis engine. Voice and text data are taken in as input, and the user's emotional state is determined from their content. The output of this analysis provides information about the user's current emotional state, and the guidance content is adjusted accordingly.
[0224] Step 5:
[0225] The terminal provides the user with guidance content generated on the server via audio output. Specifically, the content is generated using speech synthesis technology and then spoken to the user. This output is transmitted in a tone that is easy for the user to hear.
[0226] Step 6:
[0227] The device utilizes a camera and AR technology to overlay digital information onto real-world scenery. Users capture their real-world space through the camera, and digital information related to sightseeing is displayed on top of that image. This enables the creation of a visual travel guide.
[0228] Step 7:
[0229] The user provides feedback on the guidance content, and the device analyzes it using speech recognition technology. The input voice and text data is converted into text format and sent to the server as feedback data. This data is incorporated as training data and will be reflected in the guidance content for future visits.
[0230] (Application Example 2)
[0231] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0232] Modern tourism demands the provision of information that caters to the diverse interests and emotions of visitors. However, conventional information systems only provide uniform information to visitors, making it difficult to personalize the experience based on their current emotions and specific interests. Therefore, there is a need to provide flexible information tailored to the individual needs of each visitor.
[0233] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0234] In this invention, the server includes means for analyzing the user's emotions and adjusting the tone of the guidance, means for acquiring the user's location data, and means for filtering tourist destination information based on interest and history data. This makes it possible to provide highly personalized guidance that matches the visitor's emotions and interests.
[0235] A "user" is a visitor who uses the tourist information system.
[0236] "Location data" refers to information indicating the user's current location, and is obtained through GPS, Wi-Fi, etc.
[0237] "Interest and history data" refers to information related to the user's past interests and browsing history.
[0238] "Tourist destination" refers to a place that users visit.
[0239] "Information data" refers to detailed guide information about tourist destinations.
[0240] "Filtering" is the process of selecting only the information that matches the user's interests from the information that has been acquired.
[0241] A "generative artificial intelligence model" is an AI technology that generates user-specific guidance based on acquired data.
[0242] "Guidance content" refers to the collection of information provided to users as tourist information.
[0243] "Emotional analysis" refers to evaluating a user's emotional state based on voice and other inputs.
[0244] A "display device" is a visual output device used to display tourist information.
[0245] "Personalization accuracy" is an indicator that shows the degree to which guidance is adapted to each individual user.
[0246] "Adjusting the tone" refers to changing the tone or manner of speaking in the announcement.
[0247] This invention aims to build a system for personalizing tourist information. It primarily utilizes terminals, servers, an emotion engine, and a generative AI model.
[0248] The device utilizes GPS and Wi-Fi to obtain visitor location data. It also collects user interest and history data and sends it to a server. The server receives this data and uses a generative artificial intelligence model to generate personalized guidance for the visitor. The emotion engine analyzes the visitor's emotions from their voice and other inputs, and adjusts the tone of the guidance based on this analysis.
[0249] Furthermore, the terminal can display information data through its display device. For example, when visiting a traditional building in a tourist area, information about the building's history and background can be displayed visually, allowing visitors to gain a deeper understanding of the place.
[0250] User responses are recorded immediately and stored as training data. This data is used to improve the accuracy of personalized guidance for future visitors.
[0251] For example, if a visitor is interested in nature, the emotion engine analyzes their level of satisfaction and provides information about the natural environment and flora and fauna related to the visited location. Furthermore, if the system detects that a visitor is tired, it suggests nearby places to refresh, optimizing the guidance according to the visitor's state.
[0252] An example of a prompt to input into the generating AI model is: "The visitor is currently in a famous garden and is interested in nature. Please provide the visitor with detailed information about the plants and animals in the garden."
[0253] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0254] Step 1:
[0255] The device acquires the visitor's location data. Specifically, it obtains the current latitude and longitude data using a GPS module or Wi-Fi communication. This is output as location data and becomes input for the next step.
[0256] Step 2:
[0257] The server receives location data transmitted from the terminal, along with user interest and history data. The received data is input into a filtering algorithm, which then selects and extracts relevant tourist destination information. This results in the output of relevant tourist destination information, which is then used for subsequent processing.
[0258] Step 3:
[0259] The server inputs filtered tourist destination information into a generating artificial intelligence model. This model generates user-specific guidance content based on prompt messages. The generated guidance content is output and becomes the guidance information sent to the terminal.
[0260] Step 4:
[0261] The terminal outputs the generated guidance content as audio using speech synthesis technology. The guidance information is conveyed to the visitor via the audio output device. At this stage, the guidance content is provided to the user.
[0262] Step 5:
[0263] The terminal analyzes user input using speech recognition technology. The obtained input data is sent to a server, which then uses a regenerative AI model to generate additional information. Through this process, information is output that responds to user feedback and additional requests.
[0264] Step 6:
[0265] The device uses an emotion engine to analyze the user's emotions from their voice and facial expressions. Based on this analysis data, information is output to adjust the tone of the guidance, and this information is applied to the voice guidance.
[0266] Step 7:
[0267] The server records user responses and stores them in a learning database. This recorded data is analyzed to improve the accuracy of future guidance and is used as feedback information for guiding subsequent visitors.
[0268] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0269] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0270] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0271] [Second Embodiment]
[0272] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0273] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0274] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0275] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0276] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0277] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0278] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0279] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0280] The specific processing program 56 is an example of the "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by operating as the specific processing unit 290 according to the specific processing program 56 executed by the processor 28 on the RAM 30.
[0281] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the specific processing unit 290.
[0282] In the smart glasses 214, the processor 46 performs reception output processing. The storage 50 stores a reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by operating as the control unit 46A according to the reception output program 60 executed by the processor 46 on the RAM 48.
[0283] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0284] The present invention is a system that provides personalized tourism guidance to travelers, and the specific implementation form thereof will be described below. This system is composed of a user's terminal, a server, and application software that links them.
[0285] The user's terminal acquires location information using GPS or Wi-Fi. The terminal transmits individual data including the user's pre-registered interest / history data to the server. This includes, for example, interest categories such as history and food culture and past visit history.
[0286] Based on the received location information and interest / history data, the server searches for tourist spot information in the database and creates a set of filtered information. Based on this information, the server uses a generated artificial intelligence model to generate user-specific customized guidance content.
[0287] As a specific example, when the user is in an ancient capital of Japan, the server preferentially extracts information on neighboring historical temples and gardens, and uses a generated artificial intelligence to generate background and related anecdotes as guidance content. This guidance content is transmitted to the user's terminal and provided by voice output.
[0288] Through voice guidance, the user can learn details about a specific spot, and furthermore, the terminal accepts additional questions from the user through voice recognition. This information is sequentially transmitted to the server, and more detailed information according to the user's interests, such as the founding period of the temple and information about particularly famous abbots, is provided.
[0289] The history of the user's reactions and actions is recorded by the server, which is used as learning data for more refined proposals in subsequent guidance. Thereby, the system can continuously provide tourism guidance that better suits the individual preferences of the user.
[0290] The following describes the processing flow.
[0291] Step 1:
[0292] The terminal uses GPS or Wi-Fi to obtain location information in order to identify the user's current location. This location information is essential for subsequent data processing.
[0293] Step 2:
[0294] The terminal transmits the categories of interests and past visit history pre-set by the user to the server. Thereby, basic data for identifying information highly relevant to the user is provided.
[0295] Step 3:
[0296] The server searches for tourist spots in its database based on location information received and user interest / history data. Spots near the specified location are prioritized for extraction.
[0297] Step 4:
[0298] Based on filtered tourist spot information, the server uses a generative artificial intelligence model to generate customized guide content. This includes the history, culture, and interesting anecdotes of the place.
[0299] Step 5:
[0300] The server generates the guidance content and sends it to the terminal, which then provides it to the user via voice. Through voice guidance, the user can receive information not only visually but also aurally.
[0301] Step 6:
[0302] When a user asks a question or requests additional information in response to voice guidance, the device processes it using speech recognition. The question is converted into text data and sent to the server.
[0303] Step 7:
[0304] The server retrieves relevant information from the database based on the user's question and generates additional guidance. This new information provides detailed answers tailored to the user's interests.
[0305] Step 8:
[0306] The device provides newly generated information to the user via voice, continuing the guidance interaction. User feedback and behavioral data are recorded on the server and used for personalization in the future.
[0307] (Example 1)
[0308] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0309] In the conventional tourism guidance system, the information provided to users was uniform, and customization according to individual hobbies and interests was insufficient. Also, there was a problem that it was difficult to provide real-time information proposals based on the user's current location. Furthermore, since there was no mechanism to utilize the user's feedback and behavior history for the next guidance, the improvement of services that utilized experience was hindered.
[0310] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0311] In this invention, the server includes a device for acquiring the user's position, a device for selecting data of geographical spots based on the position and the user's interests and history information, and a device for generating personalized guidance information based on the selected data using an artificial intelligence model. Thereby, it becomes possible to provide timely and personalized tourism guidance based on the interests of each user.
[0312] The "device for acquiring the user's position" is a device or technology for measuring and collecting position information for detecting where the user is currently located.
[0313] The "interests and history information" refers to data based on the individual areas of interest registered in advance by the user and the past behavior history.
[0314] The "device for selecting data of geographical spots" is a facility or process for selecting and filtering relevant tourist destinations and spot information based on the user's current location and preferences.
[0315] A "device that generates information using an artificial intelligence model" is a system that uses artificial intelligence technology to generate guidance content optimized for a specific user based on the data it receives.
[0316] "Personalized information" refers to travel and sightseeing information that is customized according to the individual user's interests and current circumstances.
[0317] "Feedback and behavioral history" refers to a record of a user's reactions to the information they receive and their subsequent actions.
[0318] This invention is a system that provides users with personalized tourist information in real time. The system mainly consists of terminals, servers, and applications that connect them.
[0319] The device obtains the user's location information using GPS and Wi-Fi. In particular, using mobile devices such as smartphones and tablets allows for more precise location tracking. This enables location-based personalization.
[0320] The user's device sends data to the server, including pre-registered interest categories (e.g., history, art, food culture, etc.) and past browsing history. This data communication uses common internet protocols, and security is ensured by applying data encryption technology as needed.
[0321] The server searches a database based on received location information and user interest and history data. This database contains a large amount of tourist destination information, which the server filters to create the most suitable information set for the user. In this process, a generative AI model is used to generate user-specific guidance information based on the filtered data. The AI model generates text according to the user's interests, providing detailed background information and anecdotes.
[0322] As a concrete example, suppose a user is interested in history and is visiting an ancient Japanese capital. In this case, the server identifies nearby historical sites and cultural landmarks, and uses an AI model to generate guidance content that includes the historical significance and anecdotes behind them. The AI model might receive a prompt like this: "Please explain the history and important events of the historical building closest to your current location."
[0323] The generated guidance information is sent to the user's device and provided to the user via voice output. The device uses speech synthesis technology to convert the generated text into natural language speech, allowing the user to receive the information hands-free.
[0324] Furthermore, if the user provides additional questions via voice, the device can recognize this and resend the data to the server to provide even more detailed information. This allows for a more interactive and deeper understanding of the tourist experience.
[0325] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0326] Step 1:
[0327] The device uses GPS or Wi-Fi to obtain the user's current location.
[0328] Input: GPS sensor or Wi-Fi signal.
[0329] Processing: The terminal uses this data to perform triangulation and location algorithms to calculate the user's precise geographic coordinates.
[0330] Output: Geographic coordinate information such as latitude and longitude.
[0331] Specific operation: The device periodically updates its location information and tracks the user's movements in real time.
[0332] Step 2:
[0333] The device sends user interest and past history data to the server.
[0334] Input: Data on the user's interest categories and past visit history.
[0335] Processing: The terminal packages this data using a secure protocol and sends it to the server.
[0336] Output: Data about the user's interests and history received by the server.
[0337] Specific operation: Encrypt data using transport layer security to prevent unauthorized access by third parties.
[0338] Step 3:
[0339] The server searches the database based on location information and interest / history data received, and filters the information on tourist spots.
[0340] Input: Location information, interest data, history data.
[0341] Processing: The server uses SQL queries or similar database search methods to select relevant tourist destination information.
[0342] Output: A list of tourist spots that match the user's interests.
[0343] Specific operation: The server uses a scoring algorithm to prioritize selecting the most relevant spots.
[0344] Step 4:
[0345] The server uses an AI model to generate customized guidance content.
[0346] Input: Filtered tourist spot information.
[0347] Processing: Input prompt text into the generation AI model and generate user-specific detailed instructions.
[0348] Output: Customized guidance information.
[0349] Specific operation: The AI model utilizes natural language generation technology to create personalized explanatory text as information.
[0350] Step 5:
[0351] The server generates guidance information and sends it to the user's terminal.
[0352] Input: Guidance information.
[0353] Processing: The server generates a data packet and sends it to the terminal using a communication protocol.
[0354] Output: Guidance information received by the terminal.
[0355] Specific actions: To improve the reliability of data transmission, utilize error checking functions.
[0356] Step 6:
[0357] The terminal provides the user with the received guidance content using its voice output function.
[0358] Input: Guidance information.
[0359] Processing: The terminal's speech synthesis engine converts the text information into speech.
[0360] Output: User receives guidance via voice.
[0361] Specific actions: Adjust the tone and pronunciation of the synthesized speech to ensure smoothness and naturalness of the voice.
[0362] Step 7:
[0363] The device uses voice recognition to accept additional questions from the user.
[0364] Input: User's voice question.
[0365] Processing: The speech recognition engine converts the speech into text and sends the question to the server.
[0366] Output: Text data of the question.
[0367] Specific actions: Noise filtering and optimization of the speech model are performed to improve the accuracy of speech recognition.
[0368] Step 8:
[0369] The server will provide further information based on the user's additional questions.
[0370] Input: User question text.
[0371] Processing: The server retrieves relevant information from the database and generates information to answer the user's question.
[0372] Output: Additional information.
[0373] Specific operation: The server uses the AI model again to generate more detailed and relevant information.
[0374] (Application Example 1)
[0375] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0376] In modern urban tourism, travelers struggle to find places that match their interests amidst a vast amount of information. Furthermore, the information provided is often too generalized, resulting in a lack of personalized tourism experiences tailored to individual interests and preferences. Additionally, providing users with relevant information in real time while they are navigating urban areas presents a significant challenge.
[0377] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0378] In this invention, the server includes means for acquiring the user's location information, means for filtering information on tourist spots based on the location information and the user's interest and history data, and means for generating customized guidance content based on the filtered information using a generative artificial intelligence model. This enables the user to receive real-time, personalized tourist information tailored to their individual interests while they are in a city.
[0379] "Means for obtaining user location information" refers to devices or methods that have the function of identifying the user's current location and transmitting that information to a server.
[0380] "Means for filtering tourist spot information based on user interest and history data" refers to a function that selects relevant tourist destination information based on the user's past interests and visit history.
[0381] "Means for generating customized guidance content based on the filtered information using a generative artificial intelligence model" refers to a function that uses AI to process selected tourist information in a way that is suitable for individual users and to create original guidance content.
[0382] "Means of providing generated guidance content to users in audio format" refers to a function that converts text information into audio to convey guidance to users in an easy-to-understand manner.
[0383] "A means of recognizing user input through speech and providing additional information" refers to a function that understands voice-based questions and requests from users and provides information accordingly.
[0384] "Means for recording user responses and using them as learning data" refers to a function that accumulates data on users' actions and responses when using the system, and uses that data to improve future guidance.
[0385] "A means of acquiring and providing real-time tourist information within a city" refers to a function that instantly acquires and provides the most appropriate tourist information for a user as they move around the city.
[0386] "A means of providing additional details via voice based on user questions" refers to a function that can prepare detailed information in response to questions asked by the user and deliver it in voice format.
[0387] To realize this invention, the user's device first obtains its current location information using GPS or Wi-Fi. The device has the function to send this location information, along with the user's previously registered interests and history data, to a server. The server receives this data and searches a tourism database to filter for appropriate spot information.
[0388] Next, the server uses a generative AI model (specifically, OpenAI's ChatGPT API) based on the filtered information to generate personalized tourist information tailored to the user. The generated information is then transmitted from the terminal to the user using speech synthesis technology.
[0389] Users can ask questions using voice recognition technology, and their input is sent to the server. The server analyzes the questions, generates corresponding additional information, and provides it to the user again via voice. This allows users to obtain detailed information about tourist destinations.
[0390] Furthermore, the user's device records their responses, and this data is stored on the server and used as learning data to further personalize future guidance.
[0391] Specific example:
[0392] For example, if a user is a tourist interested in Japanese culture and is visiting a specific area of Tokyo, the server will prioritize providing information about traditional events and historical buildings rooted in that area. If the user asks, "What is the history of this building?" during the audio guidance, the server will search for the appropriate information and provide it to the user verbally.
[0393] Examples of prompts for a generative AI model:
[0394] "The user is in the city center. Please briefly describe the nearby historical landmarks."
[0395] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0396] Step 1:
[0397] The user's device uses GPS and Wi-Fi to obtain its current location. This location information is temporarily stored on the device while waiting for the next data transmission. The input is GPS data, and the output is coordinate data represented as location information.
[0398] Step 2:
[0399] The device collects location information and interest / history data previously registered by the user, and sends this information to the server. The input consists of location information and interest / history data, and the output is a request packet sent to the server.
[0400] Step 3:
[0401] The server receives the request packet and analyzes the location information and interest / history data. Based on the analysis results, it filters the relevant spot information from the tourism database. The input is the request packet, and the output is the filtered tourist spot information. Here, as a data processing step, spot information that matches the specified criteria is selected.
[0402] Step 4:
[0403] The server uses a generative artificial intelligence model to generate customized guidance content based on filtered tourist spot information. The input is filtered information, and the output is user-specific guidance content. Specifically, prompts are generated for the AI model.
[0404] Step 5:
[0405] The generated guidance content is converted into an audio file using speech synthesis technology and sent to the terminal. The input is the generated guidance content, and the output is an audio file.
[0406] Step 6:
[0407] The device plays an audio file and provides instructions to the user. The user can listen to this and make further questions or requests by voice. The input is an audio file, and the output is audio data recorded as the user's response.
[0408] Step 7:
[0409] Questions and requests made by the user are converted into text by the device's speech recognition function and sent to the server. The input is the user's voice, and the output is sent to the server as a text message.
[0410] Step 8:
[0411] The server receives user questions transcribed into text via speech recognition, searches for additional information based on the content, and generates further detailed guides using a generative AI model. The input is the user's question, and the output is the additional guidance content.
[0412] Step 9:
[0413] The generated additional informational content is converted back into an audio file and sent to the device. The device then plays this for the user, completing the provision of detailed information. The input is the additional informational content, and the output is the audio guide delivered to the user.
[0414] Step 10:
[0415] User behavior history and responses are recorded on the device and sent to the server's database. This data is used as training data to improve future guidance content. The input is user behavior data, and the output is stored as training data.
[0416] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0417] This invention combines an emotion engine with a system that provides personalized tourist information to users, and a specific embodiment thereof will be described. This system consists of a user's terminal, a server, and an application including the emotion engine.
[0418] First, the device uses GPS or Wi-Fi to obtain the user's location information. Based on this information, it prepares to collect information about the user's current location and nearby tourist attractions.
[0419] Next, the device collects the user's interest and history data and sends it to the server. This data is set by the user in advance and includes areas of interest such as "nature," "history," and "art."
[0420] Based on the received data, the server uses a generative artificial intelligence model to create customized guidance tailored to the user. At this stage, the emotion engine analyzes the user's emotions from their voice and input, and uses this information to appropriately adjust the guidance. For example, if it detects that the user is tired, it optimizes the guidance to recommend relaxing places to visit.
[0421] The generated guidance content is provided to the user as audio via the device. The tone of the guidance is adjusted based on the analysis results from the emotion engine. Furthermore, the device can integrate with the camera function to overlay digital information onto the real-world scenery. This allows users to enjoy an interactive experience that fully utilizes both sight and hearing.
[0422] When a user asks a question or responds to the guidance, the device uses its voice recognition function to analyze the content and send it to the server. The server stores the user's feedback and uses it as training data to improve future guidance. Data acquired by the emotion engine is also used for training, improving the accuracy of personalized guidance.
[0423] For example, if a user is interested in art and is visiting a museum in Tokyo, the emotion engine will sense their joy and excitement and provide highly tailored information and experiences, such as detailed background information related to the exhibits and other recommended spots. This allows users to receive advanced, personalized travel guidance, improving their overall travel satisfaction.
[0424] The following describes the processing flow.
[0425] Step 1:
[0426] The device uses GPS or Wi-Fi to determine the user's current location in order to obtain their location information. This location information serves as basic data for providing tourist information.
[0427] Step 2:
[0428] The device checks the user's pre-set interests and past travel history and sends them to the server. These include areas of interest such as history, nature, and art.
[0429] Step 3:
[0430] The server receives location information and interest / history data, searches the database, and filters out tourist spot information that is most relevant to the user.
[0431] Step 4:
[0432] Based on the information filtered by the server, a generative artificial intelligence model is used to generate customized guidance content. During this process, the emotion engine analyzes the user's voice and input data to determine the user's current emotional state.
[0433] Step 5:
[0434] Based on the analysis results of the emotion engine, the server adjusts the guidance content. For example, if the user is seeking relaxation, it will recommend calm places.
[0435] Step 6:
[0436] The server generates and sends the adjusted guidance content to the terminal, which then provides it to the user as audio. The audio guidance is played back in a tone that matches the user's emotional state.
[0437] Step 7:
[0438] As the user listens to the instructions and asks questions or makes responses, the device uses speech recognition to analyze the content and sends it to the server. This data is recorded as feedback on the responses.
[0439] Step 8:
[0440] The server collects user feedback as training data to improve the personalization accuracy of future guidance. Additionally, the results of the emotion engine's analysis are also stored as data and used for future guidance.
[0441] (Example 2)
[0442] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0443] When travelers choose places to visit and the order in which to visit them, they are only provided with general information, making it difficult to provide guidance tailored to their individual interests and circumstances. Furthermore, it is challenging to optimize the guidance content in real time by reflecting the user's emotions and feedback during their trip.
[0444] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0445] In this invention, the server includes means for acquiring the user's spatial location information, means for extracting information on travel spots, and means for generating personalized guidance content using a generative information processing model. This makes it possible to provide dynamically personalized guidance according to the user's interests and current emotional state.
[0446] "Spatial location information" refers to data used to identify a user's geographical location, and is acquired using technologies such as GPS and Wi-Fi.
[0447] "Interest and history data" refers to information based on a user's past interests and behavioral history, and reflects the user's preferences and interests.
[0448] A "generative information processing model" is an algorithm or method that generates personalized information based on user data, and primarily utilizes AI technology.
[0449] "Personalized guidance" refers to travel and sightseeing recommendations tailored to the user's specific interests and current circumstances.
[0450] "Audio output" refers to the means and devices used to convey information to users as sound, and is primarily done through speech generation technology.
[0451] "Speech recognition" is a technology in which a computer analyzes the voice spoken by a user and interprets it as text data or instructions.
[0452] "Analyzing emotional state" is a process that estimates what emotions the user is currently experiencing, based on their voice and text input.
[0453] An "image acquisition device" is a device such as a camera or video camera used to record real-world scenes and situations in digital format.
[0454] "Information overlay" is a technology that adds digital information to the real-world field of view, helping users intuitively understand their physical environment.
[0455] The embodiments for carrying out this invention are shown below.
[0456] First, the device uses GPS and Wi-Fi to acquire the user's spatial location information. This information is used to determine the user's current geographical location and is part of the criteria for selecting the next tourist spot to visit. The device has a built-in location sensor, and this data is acquired through dedicated application software.
[0457] Next, the user inputs their interests and historical data into the device. This includes pre-set areas of interest and past activity history. Through the application's settings screen, the user can select interests such as "nature," "history," and "art."
[0458] Subsequently, the server receives the user's location information and interest data transmitted from the terminal. The server utilizes a generative information processing model to generate personalized guidance based on this data. The generated guidance is optimized by AI to best suit the user's specific needs. As a concrete example of a prompt, the instruction "List tourist spots that the user is of high interest" is input to the AI model.
[0459] Furthermore, the server uses an emotion analysis engine to analyze voice and text input from the user, adjusting the guidance content according to the user's emotional state. This includes a function that, based on the process of analyzing the emotional state, recommends tourist spots that are appropriate for a user who wants to relax.
[0460] The generated guidance content is delivered to the user through the terminal via natural sound output. This allows the user to intuitively plan their trip while receiving voice guidance. Furthermore, the terminal uses a video acquisition device to overlay information onto the real-world scenery, enabling the user to receive information more interactively.
[0461] This format provides a dynamic sightseeing guide experience tailored to each user's individual interests and emotional state, while also utilizing user feedback as learning data, ensuring that the accuracy of the guide content improves over time.
[0462] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0463] Step 1:
[0464] The device acquires the user's spatial location information via GPS or Wi-Fi. Specifically, it collects current geographic location data using a built-in location sensor. This input data is processed through the device application and output as the user's coordinate information. This location information is useful as basic data when selecting tourist spots.
[0465] Step 2:
[0466] The user enters their interests and history data into the device. In the application's settings screen, they select categories of interest (e.g., "Nature," "History," "Art"). This information is stored in the system as user preference data and transferred to the server as input data.
[0467] Step 3:
[0468] The server uses a generative AI model to generate customized guidance content based on the user's location information and interest data received from the terminal. The AI model processes the input data as prompts and extracts and generates information that best suits the user's needs. As output, a user-optimized list of tourist information is generated.
[0469] Step 4:
[0470] The server analyzes the user's input (voice and text) using an emotion analysis engine. Voice and text data are taken in as input, and the user's emotional state is determined from their content. The output of this analysis provides information about the user's current emotional state, and the guidance content is adjusted accordingly.
[0471] Step 5:
[0472] The terminal provides the user with guidance content generated on the server via audio output. Specifically, the content is generated using speech synthesis technology and then spoken to the user. This output is transmitted in a tone that is easy for the user to hear.
[0473] Step 6:
[0474] The device utilizes a camera and AR technology to overlay digital information onto real-world scenery. Users capture their real-world space through the camera, and digital information related to sightseeing is displayed on top of that image. This enables the creation of a visual travel guide.
[0475] Step 7:
[0476] The user provides feedback on the guidance content, and the device analyzes it using speech recognition technology. The input voice and text data is converted into text format and sent to the server as feedback data. This data is incorporated as training data and will be reflected in the guidance content for future visits.
[0477] (Application Example 2)
[0478] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0479] Modern tourism demands the provision of information that caters to the diverse interests and emotions of visitors. However, conventional information systems only provide uniform information to visitors, making it difficult to personalize the experience based on their current emotions and specific interests. Therefore, there is a need to provide flexible information tailored to the individual needs of each visitor.
[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0481] In this invention, the server includes means for analyzing the user's emotions and adjusting the tone of the guidance, means for acquiring the user's location data, and means for filtering tourist destination information based on interest and history data. This makes it possible to provide highly personalized guidance that matches the visitor's emotions and interests.
[0482] A "user" is a visitor who uses the tourist information system.
[0483] "Location data" refers to information indicating the user's current location, and is obtained through GPS, Wi-Fi, etc.
[0484] "Interest and history data" refers to information related to the user's past interests and browsing history.
[0485] "Tourist destination" refers to a place that users visit.
[0486] "Information data" refers to detailed guide information about tourist destinations.
[0487] "Filtering" is the process of selecting only the information that matches the user's interests from the information that has been acquired.
[0488] A "generative artificial intelligence model" is an AI technology that generates user-specific guidance based on acquired data.
[0489] "Guidance content" refers to the collection of information provided to users as tourist information.
[0490] "Emotional analysis" refers to evaluating a user's emotional state based on voice and other inputs.
[0491] A "display device" is a visual output device used to display tourist information.
[0492] "Personalization accuracy" is an indicator that shows the degree to which guidance is adapted to each individual user.
[0493] "Adjusting the tone" refers to changing the tone or manner of speaking in the announcement.
[0494] This invention aims to build a system for personalizing tourist information. It primarily utilizes terminals, servers, an emotion engine, and a generative AI model.
[0495] The device utilizes GPS and Wi-Fi to obtain visitor location data. It also collects user interest and history data and sends it to a server. The server receives this data and uses a generative artificial intelligence model to generate personalized guidance for the visitor. The emotion engine analyzes the visitor's emotions from their voice and other inputs, and adjusts the tone of the guidance based on this analysis.
[0496] Furthermore, the terminal can display information data through its display device. For example, when visiting a traditional building in a tourist area, information about the building's history and background can be displayed visually, allowing visitors to gain a deeper understanding of the place.
[0497] User responses are recorded immediately and stored as training data. This data is used to improve the accuracy of personalized guidance for future visitors.
[0498] For example, if a visitor is interested in nature, the emotion engine analyzes their level of satisfaction and provides information about the natural environment and flora and fauna related to the visited location. Furthermore, if the system detects that a visitor is tired, it suggests nearby places to refresh, optimizing the guidance according to the visitor's state.
[0499] An example of a prompt to input into the generating AI model is: "The visitor is currently in a famous garden and is interested in nature. Please provide the visitor with detailed information about the plants and animals in the garden."
[0500] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0501] Step 1:
[0502] The device acquires the visitor's location data. Specifically, it obtains the current latitude and longitude data using a GPS module or Wi-Fi communication. This is output as location data and becomes input for the next step.
[0503] Step 2:
[0504] The server receives location data transmitted from the terminal, along with user interest and history data. The received data is input into a filtering algorithm, which then selects and extracts relevant tourist destination information. This results in the output of relevant tourist destination information, which is then used for subsequent processing.
[0505] Step 3:
[0506] The server inputs filtered tourist destination information into a generating artificial intelligence model. This model generates user-specific guidance content based on prompt messages. The generated guidance content is output and becomes the guidance information sent to the terminal.
[0507] Step 4:
[0508] The terminal outputs the generated guidance content as audio using speech synthesis technology. The guidance information is conveyed to the visitor via the audio output device. At this stage, the guidance content is provided to the user.
[0509] Step 5:
[0510] The terminal analyzes user input using speech recognition technology. The obtained input data is sent to a server, which then uses a regenerative AI model to generate additional information. Through this process, information is output that responds to user feedback and additional requests.
[0511] Step 6:
[0512] The device uses an emotion engine to analyze the user's emotions from their voice and facial expressions. Based on this analysis data, information is output to adjust the tone of the guidance, and this information is applied to the voice guidance.
[0513] Step 7:
[0514] The server records user responses and stores them in a learning database. This recorded data is analyzed to improve the accuracy of future guidance and is used as feedback information for guiding subsequent visitors.
[0515] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0516] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0517] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0518] [Third Embodiment]
[0519] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0520] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0521] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0522] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0523] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0524] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0525] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0526] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0527] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0528] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0529] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0530] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0531] This invention relates to a system that provides personalized tourist information to travelers, and the specific form of its implementation is described below. This system consists of a user's terminal, a server, and application software that connects them.
[0532] The user's device obtains location information using GPS or Wi-Fi. The device sends individual data to the server, including interest and history data that the user has registered in advance. This includes, for example, interest categories such as history and food culture, and past visit history.
[0533] Based on the received location information and interest / history data, the server searches for tourist spot information in the database and creates a filtered set of information. Based on this information, the server uses a generative artificial intelligence model to generate user-specific, customized guidance content.
[0534] As a concrete example, if the user is in an ancient Japanese capital, the server will prioritize extracting information on nearby historical temples and gardens, and use artificial intelligence to generate guidance content that includes their background and related anecdotes. This guidance content is then sent to the user's device and provided via voice output.
[0535] Through voice guidance, users can learn details about specific locations, and the device can even take additional questions from the user via voice recognition. This information is continuously transmitted to a server, providing further details tailored to the user's interests, such as the founding date of the temple or information about particularly famous abbots.
[0536] User responses and behavioral history are recorded on the server and used as learning data to create more refined suggestions in future guidance. This allows the system to continuously provide tourist information that is more tailored to each user's individual preferences.
[0537] The following describes the processing flow.
[0538] Step 1:
[0539] The device obtains location information using GPS or Wi-Fi to determine the user's current location. This location information is essential for subsequent data processing.
[0540] Step 2:
[0541] The device sends pre-set interest categories and past browsing history to the server. This provides the basic data needed to identify information that is highly relevant to the user.
[0542] Step 3:
[0543] The server searches for tourist spots in its database based on location information received and user interest / history data. Spots near the specified location are prioritized for extraction.
[0544] Step 4:
[0545] Based on filtered tourist spot information, the server uses a generative artificial intelligence model to generate customized guide content. This includes the history, culture, and interesting anecdotes of the place.
[0546] Step 5:
[0547] The server generates the guidance content and sends it to the terminal, which then provides it to the user via voice. Through voice guidance, the user can receive information not only visually but also aurally.
[0548] Step 6:
[0549] When a user asks a question or requests additional information in response to voice guidance, the device processes it using speech recognition. The question is converted into text data and sent to the server.
[0550] Step 7:
[0551] The server retrieves relevant information from the database based on the user's question and generates additional guidance. This new information provides detailed answers tailored to the user's interests.
[0552] Step 8:
[0553] The device provides newly generated information to the user via voice, continuing the guidance interaction. User feedback and behavioral data are recorded on the server and used for personalization in the future.
[0554] (Example 1)
[0555] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0556] Traditional tourist information systems provided users with uniform information, lacking sufficient customization to suit individual hobbies and interests. Furthermore, they struggled to offer real-time information based on the user's current location. Additionally, the lack of mechanisms to utilize user feedback and behavioral history for future guidance hindered service improvement based on experience.
[0557] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0558] In this invention, the server includes a device for acquiring the user's location, a device for selecting geographical spot data based on the location and the user's interests and history information, and a device for generating personalized guidance information based on the selected data using an artificial intelligence model. This makes it possible to provide timely and personalized tourist information based on each user's interests.
[0559] A "device for acquiring user location" is a device or technology for measuring and collecting location information to detect where a user is currently located.
[0560] "Interest and history information" refers to data based on individual areas of interest and past behavioral history that users have registered in advance.
[0561] A "device for selecting geographical spot data" is equipment or a process for selecting and filtering relevant tourist destinations and spot information based on the user's current location and preferences.
[0562] A "device that generates information using an artificial intelligence model" is a system that uses artificial intelligence technology to generate guidance content optimized for a specific user based on the data it receives.
[0563] "Personalized information" refers to travel and sightseeing information that is customized according to the individual user's interests and current circumstances.
[0564] "Feedback and behavioral history" refers to a record of a user's reactions to the information they receive and their subsequent actions.
[0565] This invention is a system that provides users with personalized tourist information in real time. The system mainly consists of terminals, servers, and applications that connect them.
[0566] The device obtains the user's location information using GPS and Wi-Fi. In particular, using mobile devices such as smartphones and tablets allows for more precise location tracking. This enables location-based personalization.
[0567] The user's device sends data to the server, including pre-registered interest categories (e.g., history, art, food culture, etc.) and past browsing history. This data communication uses common internet protocols, and security is ensured by applying data encryption technology as needed.
[0568] The server searches a database based on received location information and user interest and history data. This database contains a large amount of tourist destination information, which the server filters to create the most suitable information set for the user. In this process, a generative AI model is used to generate user-specific guidance information based on the filtered data. The AI model generates text according to the user's interests, providing detailed background information and anecdotes.
[0569] As a concrete example, suppose a user is interested in history and is visiting an ancient Japanese capital. In this case, the server identifies nearby historical sites and cultural landmarks, and uses an AI model to generate guidance content that includes the historical significance and anecdotes behind them. The AI model might receive a prompt like this: "Please explain the history and important events of the historical building closest to your current location."
[0570] The generated guidance information is sent to the user's device and provided to the user via voice output. The device uses speech synthesis technology to convert the generated text into natural language speech, allowing the user to receive the information hands-free.
[0571] Furthermore, if the user provides additional questions via voice, the device can recognize this and resend the data to the server to provide even more detailed information. This allows for a more interactive and deeper understanding of the tourist experience.
[0572] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0573] Step 1:
[0574] The device uses GPS or Wi-Fi to obtain the user's current location.
[0575] Input: GPS sensor or Wi-Fi signal.
[0576] Processing: The terminal uses this data to perform triangulation and location algorithms to calculate the user's precise geographic coordinates.
[0577] Output: Geographic coordinate information such as latitude and longitude.
[0578] Specific operation: The device periodically updates its location information and tracks the user's movements in real time.
[0579] Step 2:
[0580] The device sends user interest and past history data to the server.
[0581] Input: Data on the user's interest categories and past visit history.
[0582] Processing: The terminal packages this data using a secure protocol and sends it to the server.
[0583] Output: Data about the user's interests and history received by the server.
[0584] Specific operation: Encrypt data using transport layer security to prevent unauthorized access by third parties.
[0585] Step 3:
[0586] The server searches the database based on location information and interest / history data received, and filters the information on tourist spots.
[0587] Input: Location information, interest data, history data.
[0588] Processing: The server uses SQL queries or similar database search methods to select relevant tourist destination information.
[0589] Output: A list of tourist spots that match the user's interests.
[0590] Specific operation: The server uses a scoring algorithm to prioritize selecting the most relevant spots.
[0591] Step 4:
[0592] The server uses an AI model to generate customized guidance content.
[0593] Input: Filtered tourist spot information.
[0594] Processing: Input prompt text into the generation AI model and generate user-specific detailed instructions.
[0595] Output: Customized guidance information.
[0596] Specific operation: The AI model utilizes natural language generation technology to create personalized explanatory text as information.
[0597] Step 5:
[0598] The server generates guidance information and sends it to the user's terminal.
[0599] Input: Guidance information.
[0600] Processing: The server generates a data packet and sends it to the terminal using a communication protocol.
[0601] Output: Guidance information received by the terminal.
[0602] Specific actions: To improve the reliability of data transmission, utilize error checking functions.
[0603] Step 6:
[0604] The terminal provides the user with the received guidance content using its voice output function.
[0605] Input: Guidance information.
[0606] Processing: The terminal's speech synthesis engine converts the text information into speech.
[0607] Output: User receives guidance via voice.
[0608] Specific actions: Adjust the tone and pronunciation of the synthesized speech to ensure smoothness and naturalness of the voice.
[0609] Step 7:
[0610] The device uses voice recognition to accept additional questions from the user.
[0611] Input: User's voice question.
[0612] Processing: The speech recognition engine converts the speech into text and sends the question to the server.
[0613] Output: Text data of the question.
[0614] Specific actions: Noise filtering and optimization of the speech model are performed to improve the accuracy of speech recognition.
[0615] Step 8:
[0616] The server will provide further information based on the user's additional questions.
[0617] Input: User question text.
[0618] Processing: The server retrieves relevant information from the database and generates information to answer the user's question.
[0619] Output: Additional information.
[0620] Specific operation: The server uses the AI model again to generate more detailed and relevant information.
[0621] (Application Example 1)
[0622] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0623] In modern urban tourism, travelers struggle to find places that match their interests amidst a vast amount of information. Furthermore, the information provided is often too generalized, resulting in a lack of personalized tourism experiences tailored to individual interests and preferences. Additionally, providing users with relevant information in real time while they are navigating urban areas presents a significant challenge.
[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0625] In this invention, the server includes means for acquiring the user's location information, means for filtering information on tourist spots based on the location information and the user's interest and history data, and means for generating customized guidance content based on the filtered information using a generative artificial intelligence model. This enables the user to receive real-time, personalized tourist information tailored to their individual interests while they are in a city.
[0626] "Means for obtaining user location information" refers to devices or methods that have the function of identifying the user's current location and transmitting that information to a server.
[0627] "Means for filtering tourist spot information based on user interest and history data" refers to a function that selects relevant tourist destination information based on the user's past interests and visit history.
[0628] "Means for generating customized guidance content based on the filtered information using a generative artificial intelligence model" refers to a function that uses AI to process selected tourist information in a way that is suitable for individual users and to create original guidance content.
[0629] "Means of providing generated guidance content to users in audio format" refers to a function that converts text information into audio to convey guidance to users in an easy-to-understand manner.
[0630] "A means of recognizing user input through speech and providing additional information" refers to a function that understands voice-based questions and requests from the user and provides information accordingly.
[0631] "Means for recording user responses and using them as learning data" refers to a function that accumulates data on users' actions and responses when using the system, and uses that data to improve future guidance.
[0632] "A means of acquiring and providing real-time tourist information within a city" refers to a function that instantly acquires and provides the most appropriate tourist information for a user as they move around the city.
[0633] "A means of providing additional details via voice based on user questions" refers to a function that can prepare detailed information in response to questions asked by the user and deliver it in voice format.
[0634] To realize this invention, the user's device first obtains its current location information using GPS or Wi-Fi. The device has the function to send this location information, along with the user's previously registered interests and history data, to a server. The server receives this data and searches a tourism database to filter for appropriate spot information.
[0635] Next, the server uses a generative AI model (specifically, OpenAI's ChatGPT API) based on the filtered information to generate personalized tourist information tailored to the user. The generated information is then transmitted from the terminal to the user using speech synthesis technology.
[0636] Users can ask questions using voice recognition technology, and their input is sent to the server. The server analyzes the questions, generates corresponding additional information, and provides it to the user again via voice. This allows users to obtain detailed information about tourist destinations.
[0637] Furthermore, the user's device records their responses, and this data is stored on the server and used as learning data to further personalize future guidance.
[0638] Specific example:
[0639] For example, if a user is a tourist interested in Japanese culture and is visiting a specific area of Tokyo, the server will prioritize providing information about traditional events and historical buildings rooted in that area. If the user asks, "What is the history of this building?" during the audio guidance, the server will search for the appropriate information and provide it to the user verbally.
[0640] Examples of prompts for a generative AI model:
[0641] "The user is in the city center. Please briefly describe the nearby historical landmarks."
[0642] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0643] Step 1:
[0644] The user's device uses GPS and Wi-Fi to obtain its current location. This location information is temporarily stored on the device while waiting for the next data transmission. The input is GPS data, and the output is coordinate data represented as location information.
[0645] Step 2:
[0646] The device collects location information and interest / history data previously registered by the user, and sends this information to the server. The input consists of location information and interest / history data, and the output is a request packet sent to the server.
[0647] Step 3:
[0648] The server receives the request packet and analyzes the location information and interest / history data. Based on the analysis results, it filters the relevant spot information from the tourism database. The input is the request packet, and the output is the filtered tourist spot information. Here, as a data processing step, spot information that matches the specified criteria is selected.
[0649] Step 4:
[0650] The server uses a generative artificial intelligence model to generate customized guidance content based on filtered tourist spot information. The input is filtered information, and the output is user-specific guidance content. Specifically, prompts are created for the AI model.
[0651] Step 5:
[0652] The generated guidance content is converted into an audio file using speech synthesis technology and sent to the terminal. The input is the generated guidance content, and the output is an audio file.
[0653] Step 6:
[0654] The device plays an audio file and provides instructions to the user. The user can listen to this and make further questions or requests by voice. The input is an audio file, and the output is audio data recorded as the user's response.
[0655] Step 7:
[0656] Questions and requests made by the user are converted into text by the device's speech recognition function and sent to the server. The input is the user's voice, and the output is sent to the server as a text message.
[0657] Step 8:
[0658] The server receives user questions transcribed into text via speech recognition, searches for additional information based on the content, and generates further detailed guides using a generative AI model. The input is the user's question, and the output is the additional guidance content.
[0659] Step 9:
[0660] The generated additional informational content is converted back into an audio file and sent to the device. The device then plays this for the user, completing the provision of detailed information. The input is the additional informational content, and the output is the audio guide delivered to the user.
[0661] Step 10:
[0662] User behavior history and responses are recorded on the device and sent to the server's database. This data is used as training data to improve future guidance content. The input is user behavior data, and the output is stored as training data.
[0663] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0664] This invention combines an emotion engine with a system that provides personalized tourist information to users, and a specific embodiment thereof will be described. This system consists of a user's terminal, a server, and an application including the emotion engine.
[0665] First, the device uses GPS or Wi-Fi to obtain the user's location information. Based on this information, it prepares to collect information about the user's current location and nearby tourist attractions.
[0666] Next, the device collects the user's interest and history data and sends it to the server. This data is set by the user in advance and includes areas of interest such as "nature," "history," and "art."
[0667] Based on the received data, the server uses a generative artificial intelligence model to create customized guidance tailored to the user. At this stage, the emotion engine analyzes the user's emotions from their voice and input, and uses this information to appropriately adjust the guidance. For example, if it detects that the user is tired, it optimizes the guidance to recommend relaxing places to visit.
[0668] The generated guidance content is provided to the user as audio via the device. The tone of the guidance is adjusted based on the analysis results from the emotion engine. Furthermore, the device can integrate with the camera function to overlay digital information onto the real-world scenery. This allows users to enjoy an interactive experience that fully utilizes both sight and hearing.
[0669] When a user asks a question or responds to the guidance, the device uses its voice recognition function to analyze the content and send it to the server. The server stores the user's feedback and uses it as training data to improve future guidance. Data acquired by the emotion engine is also used for training, improving the accuracy of personalized guidance.
[0670] For example, if a user is interested in art and is visiting a museum in Tokyo, the emotion engine will sense their joy and excitement and provide highly tailored information and experiences, such as detailed background information related to the exhibits and other recommended spots. This allows users to receive advanced, personalized travel guidance, improving their overall travel satisfaction.
[0671] The following describes the processing flow.
[0672] Step 1:
[0673] The device uses GPS or Wi-Fi to determine the user's current location in order to obtain their location information. This location information serves as basic data for providing tourist information.
[0674] Step 2:
[0675] The device checks the user's pre-set interests and past travel history and sends them to the server. These include areas of interest such as history, nature, and art.
[0676] Step 3:
[0677] The server receives location information and interest / history data, searches the database, and filters out tourist spot information that is most relevant to the user.
[0678] Step 4:
[0679] Based on the information filtered by the server, a generative artificial intelligence model is used to generate customized guidance content. During this process, the emotion engine analyzes the user's voice and input data to determine the user's current emotional state.
[0680] Step 5:
[0681] Based on the analysis results of the emotion engine, the server adjusts the guidance content. For example, if the user is seeking relaxation, it will recommend calm places.
[0682] Step 6:
[0683] The server generates and sends the adjusted guidance content to the terminal, which then provides it to the user as audio. The audio guidance is played back in a tone that matches the user's emotional state.
[0684] Step 7:
[0685] As the user listens to the instructions and asks questions or makes responses, the device uses speech recognition to analyze the content and sends it to the server. This data is recorded as feedback on the responses.
[0686] Step 8:
[0687] The server collects user feedback as training data to improve the personalization accuracy of future guidance. Additionally, the results of the emotion engine's analysis are also stored as data and used for future guidance.
[0688] (Example 2)
[0689] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0690] When travelers choose places to visit and the order in which to visit them, they are only provided with general information, making it difficult to provide guidance tailored to their individual interests and circumstances. Furthermore, it is challenging to optimize the guidance content in real time by reflecting the user's emotions and feedback during their trip.
[0691] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0692] In this invention, the server includes means for acquiring the user's spatial location information, means for extracting information on travel spots, and means for generating personalized guidance content using a generative information processing model. This makes it possible to provide dynamically personalized guidance according to the user's interests and current emotional state.
[0693] "Spatial location information" refers to data used to identify a user's geographical location, and is acquired using technologies such as GPS and Wi-Fi.
[0694] "Interest and history data" refers to information based on a user's past interests and behavioral history, and reflects the user's preferences and interests.
[0695] A "generative information processing model" is an algorithm or method that generates personalized information based on user data, and primarily utilizes AI technology.
[0696] "Personalized guidance" refers to travel and sightseeing recommendations tailored to the user's specific interests and current circumstances.
[0697] "Audio output" refers to the means and devices used to convey information to users as sound, and is primarily done through speech generation technology.
[0698] "Speech recognition" is a technology in which a computer analyzes the voice spoken by a user and interprets it as text data or instructions.
[0699] "Analyzing emotional state" is a process that estimates what emotions the user is currently experiencing, based on their voice and text input.
[0700] An "image acquisition device" is a device such as a camera or video camera used to record real-world scenes and situations in digital format.
[0701] "Information overlay" is a technology that adds digital information to the real-world field of view, helping users intuitively understand their physical environment.
[0702] The embodiments for carrying out this invention are shown below.
[0703] First, the device uses GPS and Wi-Fi to acquire the user's spatial location information. This information is used to determine the user's current geographical location and is part of the criteria for selecting the next tourist spot to visit. The device has a built-in location sensor, and this data is acquired through dedicated application software.
[0704] Next, the user inputs their interests and historical data into the device. This includes pre-set areas of interest and past activity history. Through the application's settings screen, the user can select interests such as "nature," "history," and "art."
[0705] Subsequently, the server receives the user's location information and interest data transmitted from the terminal. The server utilizes a generative information processing model to generate personalized guidance based on this data. The generated guidance is optimized by AI to best suit the user's specific needs. As a concrete example of a prompt, the instruction "List tourist spots that the user is of high interest" is input to the AI model.
[0706] Furthermore, the server uses an emotion analysis engine to analyze voice and text input from the user, adjusting the guidance content according to the user's emotional state. This includes a function that, based on the process of analyzing the emotional state, recommends tourist spots that are appropriate for a user who wants to relax.
[0707] The generated guidance content is delivered to the user through the terminal via natural sound output. This allows the user to intuitively plan their trip while receiving voice guidance. Furthermore, the terminal uses a video acquisition device to overlay information onto the real-world scenery, enabling the user to receive information more interactively.
[0708] This format provides a dynamic sightseeing guide experience tailored to each user's individual interests and emotional state, while also utilizing user feedback as learning data, ensuring that the accuracy of the guide content improves over time.
[0709] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0710] Step 1:
[0711] The device acquires the user's spatial location information via GPS or Wi-Fi. Specifically, it collects current geographic location data using a built-in location sensor. This input data is processed through the device application and output as the user's coordinate information. This location information is useful as basic data when selecting tourist spots.
[0712] Step 2:
[0713] The user enters their interests and history data into the device. In the application's settings screen, they select categories of interest (e.g., "Nature," "History," "Art"). This information is stored in the system as user preference data and transferred to the server as input data.
[0714] Step 3:
[0715] The server uses a generative AI model to generate customized guidance content based on the user's location information and interest data received from the terminal. The AI model processes the input data as prompts and extracts and generates information that best suits the user's needs. As output, a user-optimized list of tourist information is generated.
[0716] Step 4:
[0717] The server analyzes the user's input (voice and text) using an emotion analysis engine. Voice and text data are taken in as input, and the user's emotional state is determined from their content. The output of this analysis provides information about the user's current emotional state, and the guidance content is adjusted accordingly.
[0718] Step 5:
[0719] The terminal provides the user with guidance content generated on the server via audio output. Specifically, the content is generated using speech synthesis technology and then spoken to the user. This output is transmitted in a tone that is easy for the user to hear.
[0720] Step 6:
[0721] The device utilizes a camera and AR technology to overlay digital information onto real-world scenery. Users capture their real-world space through the camera, and digital information related to sightseeing is displayed on top of that image. This enables the creation of a visual travel guide.
[0722] Step 7:
[0723] The user provides feedback on the guidance content, and the device analyzes it using speech recognition technology. The input voice and text data is converted into text format and sent to the server as feedback data. This data is incorporated as training data and will be reflected in the guidance content for future visits.
[0724] (Application Example 2)
[0725] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0726] Modern tourism demands the provision of information that caters to the diverse interests and emotions of visitors. However, conventional information systems only provide uniform information to visitors, making it difficult to personalize the experience based on their current emotions and specific interests. Therefore, there is a need to provide flexible information tailored to the individual needs of each visitor.
[0727] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0728] In this invention, the server includes means for analyzing the user's emotions and adjusting the tone of the guidance, means for acquiring the user's location data, and means for filtering tourist destination information based on interest and history data. This makes it possible to provide highly personalized guidance that matches the visitor's emotions and interests.
[0729] A "user" is a visitor who uses the tourist information system.
[0730] "Location data" refers to information indicating the user's current location, and is obtained through GPS, Wi-Fi, etc.
[0731] "Interest and history data" refers to information related to the user's past interests and browsing history.
[0732] "Tourist destination" refers to a place that users visit.
[0733] "Information data" refers to detailed guide information about tourist destinations.
[0734] "Filtering" is the process of selecting only the information that matches the user's interests from the information that has been acquired.
[0735] A "generative artificial intelligence model" is an AI technology that generates user-specific guidance based on acquired data.
[0736] "Guidance content" refers to the collection of information provided to users as tourist information.
[0737] "Emotional analysis" refers to evaluating a user's emotional state based on voice and other inputs.
[0738] A "display device" is a visual output device used to display tourist information.
[0739] "Personalization accuracy" is an indicator that shows the degree to which guidance is adapted to each individual user.
[0740] "Adjusting the tone" refers to changing the tone or manner of speaking in the announcement.
[0741] This invention aims to build a system for personalizing tourist information. It primarily utilizes terminals, servers, an emotion engine, and a generative AI model.
[0742] The device utilizes GPS and Wi-Fi to obtain visitor location data. It also collects user interest and history data and sends it to a server. The server receives this data and uses a generative artificial intelligence model to generate personalized guidance for the visitor. The emotion engine analyzes the visitor's emotions from their voice and other inputs, and adjusts the tone of the guidance based on this analysis.
[0743] Furthermore, the terminal can display information data through its display device. For example, when visiting a traditional building in a tourist area, information about the building's history and background can be displayed visually, allowing visitors to gain a deeper understanding of the place.
[0744] User responses are recorded immediately and stored as training data. This data is used to improve the accuracy of personalized guidance for future visitors.
[0745] For example, if a visitor is interested in nature, the emotion engine analyzes their level of satisfaction and provides information about the natural environment and flora and fauna related to the visited location. Furthermore, if the system detects that a visitor is tired, it suggests nearby places to refresh, optimizing the guidance according to the visitor's state.
[0746] An example of a prompt to input into the generating AI model is: "The visitor is currently in a famous garden and is interested in nature. Please provide the visitor with detailed information about the plants and animals in the garden."
[0747] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0748] Step 1:
[0749] The device acquires the visitor's location data. Specifically, it obtains the current latitude and longitude data using a GPS module or Wi-Fi communication. This is output as location data and becomes input for the next step.
[0750] Step 2:
[0751] The server receives location data transmitted from the terminal, along with user interest and history data. The received data is input into a filtering algorithm, which then selects and extracts relevant tourist destination information. This results in the output of relevant tourist destination information, which is then used for subsequent processing.
[0752] Step 3:
[0753] The server inputs filtered tourist destination information into a generating artificial intelligence model. This model generates user-specific guidance content based on prompt messages. The generated guidance content is output and becomes the guidance information sent to the terminal.
[0754] Step 4:
[0755] The terminal outputs the generated guidance content as audio using speech synthesis technology. The guidance information is conveyed to the visitor via the audio output device. At this stage, the guidance content is provided to the user.
[0756] Step 5:
[0757] The terminal analyzes user input using speech recognition technology. The obtained input data is sent to a server, which then uses a regenerative AI model to generate additional information. Through this process, information is output that responds to user feedback and additional requests.
[0758] Step 6:
[0759] The device uses an emotion engine to analyze the user's emotions from their voice and facial expressions. Based on this analysis data, information is output to adjust the tone of the guidance, and this information is applied to the voice guidance.
[0760] Step 7:
[0761] The server records user responses and stores them in a learning database. This recorded data is analyzed to improve the accuracy of future guidance and is used as feedback information for guiding subsequent visitors.
[0762] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0763] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0764] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0765] [Fourth Embodiment]
[0766] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0767] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0768] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0769] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0770] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0771] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0772] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0773] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0774] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0775] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0776] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0777] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0778] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0779] This invention relates to a system that provides personalized tourist information to travelers, and the specific form of its implementation is described below. This system consists of a user's terminal, a server, and application software that connects them.
[0780] The user's device obtains location information using GPS or Wi-Fi. The device sends individual data to the server, including interest and history data that the user has registered in advance. This includes, for example, interest categories such as history and food culture, and past visit history.
[0781] Based on the received location information and interest / history data, the server searches for tourist spot information in the database and creates a filtered set of information. Based on this information, the server uses a generative artificial intelligence model to generate user-specific, customized guidance content.
[0782] As a concrete example, if the user is in an ancient Japanese capital, the server will prioritize extracting information on nearby historical temples and gardens, and use artificial intelligence to generate guidance content that includes their background and related anecdotes. This guidance content is then sent to the user's device and provided via voice output.
[0783] Through voice guidance, users can learn details about specific locations, and the device can even take additional questions from the user via voice recognition. This information is continuously transmitted to a server, providing further details tailored to the user's interests, such as the founding date of the temple or information about particularly famous abbots.
[0784] User responses and behavioral history are recorded on the server and used as learning data to create more refined suggestions in future guidance. This allows the system to continuously provide tourist information that is more tailored to each user's individual preferences.
[0785] The following describes the processing flow.
[0786] Step 1:
[0787] The device obtains location information using GPS or Wi-Fi to determine the user's current location. This location information is essential for subsequent data processing.
[0788] Step 2:
[0789] The device sends pre-set interest categories and past browsing history to the server. This provides the basic data needed to identify information that is highly relevant to the user.
[0790] Step 3:
[0791] The server searches for tourist spots in its database based on location information received and user interest / history data. Spots near the specified location are prioritized for extraction.
[0792] Step 4:
[0793] Based on filtered tourist spot information, the server uses a generative artificial intelligence model to generate customized guide content. This includes the history, culture, and interesting anecdotes of the place.
[0794] Step 5:
[0795] The server generates the guidance content and sends it to the terminal, which then provides it to the user via voice. Through voice guidance, the user can receive information not only visually but also aurally.
[0796] Step 6:
[0797] When a user asks a question or requests additional information in response to voice guidance, the device processes it using speech recognition. The question is converted into text data and sent to the server.
[0798] Step 7:
[0799] The server retrieves relevant information from the database based on the user's question and generates additional guidance. This new information provides detailed answers tailored to the user's interests.
[0800] Step 8:
[0801] The device provides newly generated information to the user via voice, continuing the guidance interaction. User feedback and behavioral data are recorded on the server and used for personalization in the future.
[0802] (Example 1)
[0803] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0804] Traditional tourist information systems provided users with uniform information, lacking sufficient customization to suit individual hobbies and interests. Furthermore, they struggled to offer real-time information based on the user's current location. Additionally, the lack of mechanisms to utilize user feedback and behavioral history for future guidance hindered service improvement based on experience.
[0805] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0806] In this invention, the server includes a device for acquiring the user's location, a device for selecting geographical spot data based on the location and the user's interests and history information, and a device for generating personalized guidance information based on the selected data using an artificial intelligence model. This makes it possible to provide timely and personalized tourist information based on each user's interests.
[0807] A "device for acquiring user location" is a device or technology for measuring and collecting location information to detect where a user is currently located.
[0808] "Interest and history information" refers to data based on individual areas of interest and past behavioral history that users have registered in advance.
[0809] A "device for selecting geographical spot data" is equipment or a process for selecting and filtering relevant tourist destinations and spot information based on the user's current location and preferences.
[0810] A "device that generates information using an artificial intelligence model" is a system that uses artificial intelligence technology to generate guidance content optimized for a specific user based on the data it receives.
[0811] "Personalized information" refers to travel and sightseeing information that is customized according to the individual user's interests and current circumstances.
[0812] "Feedback and behavioral history" refers to a record of a user's reactions to the information they receive and their subsequent actions.
[0813] This invention is a system that provides users with personalized tourist information in real time. The system mainly consists of terminals, servers, and applications that connect them.
[0814] The device obtains the user's location information using GPS and Wi-Fi. In particular, using mobile devices such as smartphones and tablets allows for more precise location tracking. This enables location-based personalization.
[0815] The user's device sends data to the server, including pre-registered interest categories (e.g., history, art, food culture, etc.) and past browsing history. This data communication uses common internet protocols, and security is ensured by applying data encryption technology as needed.
[0816] The server searches a database based on received location information and user interest and history data. This database contains a large amount of tourist destination information, which the server filters to create the most suitable information set for the user. In this process, a generative AI model is used to generate user-specific guidance information based on the filtered data. The AI model generates text according to the user's interests, providing detailed background information and anecdotes.
[0817] As a concrete example, suppose a user is interested in history and is visiting an ancient Japanese capital. In this case, the server identifies nearby historical sites and cultural landmarks, and uses an AI model to generate guidance content that includes the historical significance and anecdotes behind them. The AI model might receive a prompt like this: "Please explain the history and important events of the historical building closest to your current location."
[0818] The generated guidance information is sent to the user's device and provided to the user via voice output. The device uses speech synthesis technology to convert the generated text into natural language speech, allowing the user to receive the information hands-free.
[0819] Furthermore, if the user provides additional questions via voice, the device can recognize this and resend the data to the server to provide even more detailed information. This allows for a more interactive and deeper understanding of the tourist experience.
[0820] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0821] Step 1:
[0822] The device uses GPS or Wi-Fi to obtain the user's current location.
[0823] Input: GPS sensor or Wi-Fi signal.
[0824] Processing: The terminal uses this data to perform triangulation and location algorithms to calculate the user's precise geographic coordinates.
[0825] Output: Geographic coordinate information such as latitude and longitude.
[0826] Specific operation: The device periodically updates its location information and tracks the user's movements in real time.
[0827] Step 2:
[0828] The device sends user interest and past history data to the server.
[0829] Input: Data on the user's interest categories and past visit history.
[0830] Processing: The terminal packages this data using a secure protocol and sends it to the server.
[0831] Output: Data about the user's interests and history received by the server.
[0832] Specific operation: Encrypt data using transport layer security to prevent unauthorized access by third parties.
[0833] Step 3:
[0834] The server searches the database based on location information and interest / history data received, and filters the information on tourist spots.
[0835] Input: Location information, interest data, history data.
[0836] Processing: The server uses SQL queries or similar database search methods to select relevant tourist destination information.
[0837] Output: A list of tourist spots that match the user's interests.
[0838] Specific operation: The server uses a scoring algorithm to prioritize selecting the most relevant spots.
[0839] Step 4:
[0840] The server uses an AI model to generate customized guidance content.
[0841] Input: Filtered tourist spot information.
[0842] Processing: Input prompt text into the generation AI model and generate user-specific detailed instructions.
[0843] Output: Customized guidance information.
[0844] Specific operation: The AI model utilizes natural language generation technology to create personalized explanatory text as information.
[0845] Step 5:
[0846] The server generates guidance information and sends it to the user's terminal.
[0847] Input: Guidance information.
[0848] Processing: The server generates a data packet and sends it to the terminal using a communication protocol.
[0849] Output: Guidance information received by the terminal.
[0850] Specific actions: To improve the reliability of data transmission, utilize error checking functions.
[0851] Step 6:
[0852] The terminal provides the user with the received guidance content using its voice output function.
[0853] Input: Guidance information.
[0854] Processing: The terminal's speech synthesis engine converts the text information into speech.
[0855] Output: User receives guidance via voice.
[0856] Specific actions: Adjust the tone and pronunciation of the synthesized speech to ensure smoothness and naturalness of the voice.
[0857] Step 7:
[0858] The device uses voice recognition to accept additional questions from the user.
[0859] Input: User's voice question.
[0860] Processing: The speech recognition engine converts the speech into text and sends the question to the server.
[0861] Output: Text data of the question.
[0862] Specific actions: Noise filtering and optimization of the speech model are performed to improve the accuracy of speech recognition.
[0863] Step 8:
[0864] The server will provide further information based on the user's additional questions.
[0865] Input: User question text.
[0866] Processing: The server retrieves relevant information from the database and generates information to answer the user's question.
[0867] Output: Additional information.
[0868] Specific operation: The server uses the AI model again to generate more detailed and relevant information.
[0869] (Application Example 1)
[0870] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0871] In modern urban tourism, travelers struggle to find places that match their interests amidst a vast amount of information. Furthermore, the information provided is often too generalized, resulting in a lack of personalized tourism experiences tailored to individual interests and preferences. Additionally, providing users with relevant information in real time while they are navigating urban areas presents a significant challenge.
[0872] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0873] In this invention, the server includes means for acquiring the user's location information, means for filtering information on tourist spots based on the location information and the user's interest and history data, and means for generating customized guidance content based on the filtered information using a generative artificial intelligence model. This enables the user to receive real-time, personalized tourist information tailored to their individual interests while they are in a city.
[0874] "Means for obtaining user location information" refers to devices or methods that have the function of identifying the user's current location and transmitting that information to a server.
[0875] "Means for filtering tourist spot information based on user interest and history data" refers to a function that selects relevant tourist destination information based on the user's past interests and visit history.
[0876] "Means for generating customized guidance content based on the filtered information using a generative artificial intelligence model" refers to a function that uses AI to process selected tourist information in a way that is suitable for individual users and to create original guidance content.
[0877] "Means of providing generated guidance content to users in audio format" refers to a function that converts text information into audio to convey guidance to users in an easy-to-understand manner.
[0878] "A means of recognizing user input through speech and providing additional information" refers to a function that understands voice-based questions and requests from the user and provides information accordingly.
[0879] "Means for recording user responses and using them as learning data" refers to a function that accumulates data on users' actions and responses when using the system, and uses that data to improve future guidance.
[0880] "A means of acquiring and providing real-time tourist information within a city" refers to a function that instantly acquires and provides the most appropriate tourist information for a user as they move around the city.
[0881] "A means of providing additional details via voice based on user questions" refers to a function that can prepare detailed information in response to questions asked by the user and deliver it in voice format.
[0882] To realize this invention, the user's device first obtains its current location information using GPS or Wi-Fi. The device has the function to send this location information, along with the user's previously registered interests and history data, to a server. The server receives this data and searches a tourism database to filter for appropriate spot information.
[0883] Next, the server uses a generative AI model (specifically, OpenAI's ChatGPT API) based on the filtered information to generate personalized tourist information tailored to the user. The generated information is then transmitted from the terminal to the user using speech synthesis technology.
[0884] Users can ask questions using voice recognition technology, and their input is sent to the server. The server analyzes the questions, generates corresponding additional information, and provides it to the user again via voice. This allows users to obtain detailed information about tourist destinations.
[0885] Furthermore, the user's device records their responses, and this data is stored on the server and used as learning data to further personalize future guidance.
[0886] Specific example:
[0887] For example, if a user is a tourist interested in Japanese culture and is visiting a specific area of Tokyo, the server will prioritize providing information about traditional events and historical buildings rooted in that area. If the user asks, "What is the history of this building?" during the audio guidance, the server will search for the appropriate information and provide it to the user verbally.
[0888] Examples of prompts for a generative AI model:
[0889] "The user is in the city center. Please briefly describe the nearby historical landmarks."
[0890] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0891] Step 1:
[0892] The user's device uses GPS and Wi-Fi to obtain its current location. This location information is temporarily stored on the device while waiting for the next data transmission. The input is GPS data, and the output is coordinate data represented as location information.
[0893] Step 2:
[0894] The device collects location information and interest / history data previously registered by the user, and sends this information to the server. The input consists of location information and interest / history data, and the output is a request packet sent to the server.
[0895] Step 3:
[0896] The server receives the request packet and analyzes the location information and interest / history data. Based on the analysis results, it filters the relevant spot information from the tourism database. The input is the request packet, and the output is the filtered tourist spot information. Here, as a data processing step, spot information that matches the specified criteria is selected.
[0897] Step 4:
[0898] The server uses a generative artificial intelligence model to generate customized guidance content based on filtered tourist spot information. The input is filtered information, and the output is user-specific guidance content. Specifically, prompts are created for the AI model.
[0899] Step 5:
[0900] The generated guidance content is converted into an audio file using speech synthesis technology and sent to the terminal. The input is the generated guidance content, and the output is an audio file.
[0901] Step 6:
[0902] The device plays an audio file and provides instructions to the user. The user can listen to this and make further questions or requests by voice. The input is an audio file, and the output is audio data recorded as the user's response.
[0903] Step 7:
[0904] Questions and requests made by the user are converted into text by the device's speech recognition function and sent to the server. The input is the user's voice, and the output is sent to the server as a text message.
[0905] Step 8:
[0906] The server receives user questions transcribed into text via speech recognition, searches for additional information based on the content, and generates further detailed guides using a generative AI model. The input is the user's question, and the output is the additional guidance content.
[0907] Step 9:
[0908] The generated additional informational content is converted back into an audio file and sent to the device. The device then plays this for the user, completing the provision of detailed information. The input is the additional informational content, and the output is the audio guide delivered to the user.
[0909] Step 10:
[0910] User behavior history and responses are recorded on the device and sent to the server's database. This data is used as training data to improve future guidance content. The input is user behavior data, and the output is stored as training data.
[0911] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0912] This invention combines an emotion engine with a system that provides personalized tourist information to users, and a specific embodiment thereof will be described. This system consists of a user's terminal, a server, and an application including the emotion engine.
[0913] First, the device uses GPS or Wi-Fi to obtain the user's location information. Based on this information, it prepares to collect information about the user's current location and nearby tourist attractions.
[0914] Next, the device collects the user's interest and history data and sends it to the server. This data is set by the user in advance and includes areas of interest such as "nature," "history," and "art."
[0915] Based on the received data, the server uses a generative artificial intelligence model to create customized guidance tailored to the user. At this stage, the emotion engine analyzes the user's emotions from their voice and input, and uses this information to appropriately adjust the guidance. For example, if it detects that the user is tired, it optimizes the guidance to recommend relaxing places to visit.
[0916] The generated guidance content is provided to the user as audio via the device. The tone of the guidance is adjusted based on the analysis results from the emotion engine. Furthermore, the device can integrate with the camera function to overlay digital information onto the real-world scenery. This allows users to enjoy an interactive experience that fully utilizes both sight and hearing.
[0917] When a user asks a question or responds to the guidance, the device uses its voice recognition function to analyze the content and send it to the server. The server stores the user's feedback and uses it as training data to improve future guidance. Data acquired by the emotion engine is also used for training, improving the accuracy of personalized guidance.
[0918] For example, if a user is interested in art and is visiting a museum in Tokyo, the emotion engine will sense their joy and excitement and provide highly tailored information and experiences, such as detailed background information related to the exhibits and other recommended spots. This allows users to receive advanced, personalized travel guidance, improving their overall travel satisfaction.
[0919] The following describes the processing flow.
[0920] Step 1:
[0921] The device uses GPS or Wi-Fi to determine the user's current location in order to obtain their location information. This location information serves as basic data for providing tourist information.
[0922] Step 2:
[0923] The device checks the user's pre-set interests and past travel history and sends them to the server. These include areas of interest such as history, nature, and art.
[0924] Step 3:
[0925] The server receives location information and interest / history data, searches the database, and filters out tourist spot information that is most relevant to the user.
[0926] Step 4:
[0927] Based on the information filtered by the server, a generative artificial intelligence model is used to generate customized guidance content. During this process, the emotion engine analyzes the user's voice and input data to determine the user's current emotional state.
[0928] Step 5:
[0929] Based on the analysis results of the emotion engine, the server adjusts the guidance content. For example, if the user is seeking relaxation, it will recommend calm places.
[0930] Step 6:
[0931] The server generates and sends the adjusted guidance content to the terminal, which then provides it to the user as audio. The audio guidance is played back in a tone that matches the user's emotional state.
[0932] Step 7:
[0933] As the user listens to the instructions and asks questions or makes responses, the device uses speech recognition to analyze the content and sends it to the server. This data is recorded as feedback on the responses.
[0934] Step 8:
[0935] The server collects user feedback as training data to improve the personalization accuracy of future guidance. Additionally, the results of the emotion engine's analysis are also stored as data and used for future guidance.
[0936] (Example 2)
[0937] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0938] When travelers choose places to visit and the order in which to visit them, they are only provided with general information, making it difficult to provide guidance tailored to their individual interests and circumstances. Furthermore, it is challenging to optimize the guidance content in real time by reflecting the user's emotions and feedback during their trip.
[0939] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0940] In this invention, the server includes means for acquiring the user's spatial location information, means for extracting information on travel spots, and means for generating personalized guidance content using a generative information processing model. This makes it possible to provide dynamically personalized guidance according to the user's interests and current emotional state.
[0941] "Spatial location information" refers to data used to identify a user's geographical location, and is acquired using technologies such as GPS and Wi-Fi.
[0942] "Interest and history data" refers to information based on a user's past interests and behavioral history, and reflects the user's preferences and interests.
[0943] A "generative information processing model" is an algorithm or method that generates personalized information based on user data, and primarily utilizes AI technology.
[0944] "Personalized guidance" refers to travel and sightseeing recommendations tailored to the user's specific interests and current circumstances.
[0945] "Audio output" refers to the means and devices used to convey information to users as sound, and is primarily done through speech generation technology.
[0946] "Speech recognition" is a technology in which a computer analyzes the voice spoken by a user and interprets it as text data or instructions.
[0947] "Analyzing emotional state" is a process that estimates what emotions the user is currently experiencing, based on their voice and text input.
[0948] An "image acquisition device" is a device such as a camera or video camera used to record real-world scenes and situations in digital format.
[0949] "Information overlay" is a technology that adds digital information to the real-world field of view, helping users intuitively understand their physical environment.
[0950] The embodiments for carrying out this invention are shown below.
[0951] First, the device uses GPS and Wi-Fi to acquire the user's spatial location information. This information is used to determine the user's current geographical location and is part of the criteria for selecting the next tourist spot to visit. The device has a built-in location sensor, and this data is acquired through dedicated application software.
[0952] Next, the user inputs their interests and historical data into the device. This includes pre-set areas of interest and past activity history. Through the application's settings screen, the user can select interests such as "nature," "history," and "art."
[0953] Subsequently, the server receives the user's location information and interest data transmitted from the terminal. The server utilizes a generative information processing model to generate personalized guidance based on this data. The generated guidance is optimized by AI to best suit the user's specific needs. As a concrete example of a prompt, the instruction "List tourist spots that the user is of high interest" is input to the AI model.
[0954] Furthermore, the server uses an emotion analysis engine to analyze voice and text input from the user, adjusting the guidance content according to the user's emotional state. This includes a function that, based on the process of analyzing the emotional state, recommends tourist spots that are appropriate for a user who wants to relax.
[0955] The generated guidance content is delivered to the user through the terminal via natural sound output. This allows the user to intuitively plan their trip while receiving voice guidance. Furthermore, the terminal uses a video acquisition device to overlay information onto the real-world scenery, enabling the user to receive information more interactively.
[0956] This format provides a dynamic sightseeing guide experience tailored to each user's individual interests and emotional state, while also utilizing user feedback as learning data, ensuring that the accuracy of the guide content improves over time.
[0957] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0958] Step 1:
[0959] The device acquires the user's spatial location information via GPS or Wi-Fi. Specifically, it collects current geographic location data using a built-in location sensor. This input data is processed through the device application and output as the user's coordinate information. This location information is useful as basic data when selecting tourist spots.
[0960] Step 2:
[0961] The user enters their interests and history data into the device. In the application's settings screen, they select categories of interest (e.g., "Nature," "History," "Art"). This information is stored in the system as user preference data and transferred to the server as input data.
[0962] Step 3:
[0963] The server uses a generative AI model to generate customized guidance content based on the user's location information and interest data received from the terminal. The AI model processes the input data as prompts and extracts and generates information that best suits the user's needs. As output, a user-optimized list of tourist information is generated.
[0964] Step 4:
[0965] The server analyzes the user's input (voice and text) using an emotion analysis engine. Voice and text data are taken in as input, and the user's emotional state is determined from their content. The output of this analysis provides information about the user's current emotional state, and the guidance content is adjusted accordingly.
[0966] Step 5:
[0967] The terminal provides the user with guidance content generated on the server via audio output. Specifically, the content is generated using speech synthesis technology and then spoken to the user. This output is transmitted in a tone that is easy for the user to hear.
[0968] Step 6:
[0969] The device utilizes a camera and AR technology to overlay digital information onto real-world scenery. Users capture their real-world space through the camera, and digital information related to sightseeing is displayed on top of that image. This enables the creation of a visual travel guide.
[0970] Step 7:
[0971] The user provides feedback on the guidance content, and the device analyzes it using speech recognition technology. The input voice and text data is converted into text format and sent to the server as feedback data. This data is incorporated as training data and will be reflected in the guidance content for future visits.
[0972] (Application Example 2)
[0973] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0974] Modern tourism demands the provision of information that caters to the diverse interests and emotions of visitors. However, conventional information systems only provide uniform information to visitors, making it difficult to personalize the experience based on their current emotions and specific interests. Therefore, there is a need to provide flexible information tailored to the individual needs of each visitor.
[0975] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0976] In this invention, the server includes means for analyzing the user's emotions and adjusting the tone of the guidance, means for acquiring the user's location data, and means for filtering tourist destination information based on interest and history data. This makes it possible to provide highly personalized guidance that matches the visitor's emotions and interests.
[0977] A "user" is a visitor who uses the tourist information system.
[0978] "Location data" refers to information indicating the user's current location, and is obtained through GPS, Wi-Fi, etc.
[0979] "Interest and history data" refers to information related to the user's past interests and browsing history.
[0980] "Tourist destination" refers to a place that users visit.
[0981] "Information data" refers to detailed guide information about tourist destinations.
[0982] "Filtering" is the process of selecting only the information that matches the user's interests from the information that has been acquired.
[0983] A "generative artificial intelligence model" is an AI technology that generates user-specific guidance based on acquired data.
[0984] "Guidance content" refers to the collection of information provided to users as tourist information.
[0985] "Emotional analysis" refers to evaluating a user's emotional state based on voice and other inputs.
[0986] A "display device" is a visual output device used to display tourist information.
[0987] "Personalization accuracy" is an indicator that shows the degree to which guidance is adapted to each individual user.
[0988] "Adjusting the tone" refers to changing the tone or manner of speaking in the announcement.
[0989] This invention aims to build a system for personalizing tourist information. It primarily utilizes terminals, servers, an emotion engine, and a generative AI model.
[0990] The device utilizes GPS and Wi-Fi to obtain visitor location data. It also collects user interest and history data and sends it to a server. The server receives this data and uses a generative artificial intelligence model to generate personalized guidance for the visitor. The emotion engine analyzes the visitor's emotions from their voice and other inputs, and adjusts the tone of the guidance based on this analysis.
[0991] Furthermore, the terminal can display information data through its display device. For example, when visiting a traditional building in a tourist area, information about the building's history and background can be displayed visually, allowing visitors to gain a deeper understanding of the place.
[0992] User responses are recorded immediately and stored as training data. This data is used to improve the accuracy of personalized guidance for future visitors.
[0993] For example, if a visitor is interested in nature, the emotion engine analyzes their level of satisfaction and provides information about the natural environment and flora and fauna related to the visited location. Furthermore, if the system detects that a visitor is tired, it suggests nearby places to refresh, optimizing the guidance according to the visitor's state.
[0994] An example of a prompt to input into the generating AI model is: "The visitor is currently in a famous garden and is interested in nature. Please provide the visitor with detailed information about the plants and animals in the garden."
[0995] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0996] Step 1:
[0997] The device acquires the visitor's location data. Specifically, it obtains the current latitude and longitude data using a GPS module or Wi-Fi communication. This is output as location data and becomes input for the next step.
[0998] Step 2:
[0999] The server receives location data transmitted from the terminal, along with user interest and history data. The received data is input into a filtering algorithm, which then selects and extracts relevant tourist destination information. This results in the output of relevant tourist destination information, which is then used for subsequent processing.
[1000] Step 3:
[1001] The server inputs filtered tourist destination information into a generating artificial intelligence model. This model generates user-specific guidance content based on prompt messages. The generated guidance content is output and becomes the guidance information sent to the terminal.
[1002] Step 4:
[1003] The terminal outputs the generated guidance content as audio using speech synthesis technology. The guidance information is conveyed to the visitor via the audio output device. At this stage, the guidance content is provided to the user.
[1004] Step 5:
[1005] The terminal analyzes user input using speech recognition technology. The obtained input data is sent to a server, which then uses a regenerative AI model to generate additional information. Through this process, information is output that responds to user feedback and additional requests.
[1006] Step 6:
[1007] The device uses an emotion engine to analyze the user's emotions from their voice and facial expressions. Based on this analysis data, information is output to adjust the tone of the guidance, and this information is applied to the voice guidance.
[1008] Step 7:
[1009] The server records user responses and stores them in a learning database. This recorded data is analyzed to improve the accuracy of future guidance and is used as feedback information for guiding subsequent visitors.
[1010] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1011] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1012] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1013] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1014] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1015] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1016] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1017] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1018] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1019] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1020] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1021] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1022] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1023] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1024] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1025] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1026] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1027] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1028] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1029] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1030] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1031] The following is further disclosed regarding the embodiments described above.
[1032] (Claim 1)
[1033] Means for obtaining the user's location information,
[1034] A means for filtering information on tourist spots based on the aforementioned location information and user interest / history data,
[1035] A means for generating customized guidance content based on the filtered information using a generative artificial intelligence model,
[1036] A means of providing the generated guidance content to the user via voice,
[1037] A means of recognizing user input via speech and providing additional information,
[1038] A system that includes means for recording user responses and using them as training data.
[1039] (Claim 2)
[1040] The system according to claim 1, further comprising means for using a camera via a terminal to superimpose and display digital information onto a real-world landscape.
[1041] (Claim 3)
[1042] The system according to claim 1, further comprising means for improving the personalization accuracy of the generated guidance content based on the aforementioned learning data.
[1043] "Example 1"
[1044] (Claim 1)
[1045] A device for acquiring the user's location,
[1046] A device for selecting geographical spot data based on the aforementioned location and user interests and history information,
[1047] A device that generates personalized guidance information based on the selected data using an artificial intelligence model,
[1048] A device that provides generated guidance information to the user via voice,
[1049] A device that recognizes user voice input and provides additional information,
[1050] A system that includes a device for recording user responses and using them as training data.
[1051] (Claim 2)
[1052] The system according to claim 1, further comprising a device that overlays and displays electronic information onto a real-world landscape using video functions via a visual device.
[1053] (Claim 3)
[1054] The system according to claim 1, further comprising a device for improving the personalization accuracy of the guidance information generated based on the aforementioned training data.
[1055] "Application Example 1"
[1056] (Claim 1)
[1057] Means for obtaining the user's location information,
[1058] A means for filtering information on tourist spots based on the aforementioned location information and user interest / history data,
[1059] A means for generating customized guidance content based on the filtered information using a generative artificial intelligence model,
[1060] A means of providing the generated guidance content to the user via voice,
[1061] A means of recognizing user input via speech and providing additional information,
[1062] A means of recording user responses and using them as training data,
[1063] A means of obtaining and providing real-time tourist information within cities,
[1064] A system that includes a means of providing additional details via voice based on user questions.
[1065] (Claim 2)
[1066] The system according to claim 1, further comprising means for using a camera via a terminal to superimpose and display digital information onto a real-world landscape, and means for providing information specifically for urban tourism.
[1067] (Claim 3)
[1068] The system according to claim 1, further comprising means for improving the personalization accuracy of the generated guidance content based on the aforementioned learning data, and means for analyzing the user's interests based on their movement history in the city.
[1069] "Example 2 of combining an emotion engine"
[1070] (Claim 1)
[1071] A means of obtaining the user's spatial location information,
[1072] A means for extracting travel spot information based on the aforementioned spatial location information and user interest / history data,
[1073] A means for generating personalized guidance content based on the extracted information using a generated information processing model,
[1074] A means of providing the generated guidance content to the user via audio output,
[1075] A means of recognizing user input via speech and providing additional information,
[1076] A means of recording user responses and using them as training data,
[1077] A system that includes means for analyzing the user's emotional state and adjusting the guidance content accordingly.
[1078] (Claim 2)
[1079] The system according to claim 1, further comprising means for using a video acquisition device via a terminal to superimpose and display information onto a real-world scene.
[1080] (Claim 3)
[1081] The system according to claim 1, further comprising means for improving the personalization accuracy of the generated guidance content based on the aforementioned learning data and emotional information.
[1082] "Application example 2 when combining with an emotional engine"
[1083] (Claim 1)
[1084] A means of obtaining user location data,
[1085] A means for filtering information on tourist destinations based on the aforementioned location data and user interest / history data,
[1086] A means for generating customized guidance content based on the filtered information using a generative artificial intelligence model,
[1087] A means of providing the generated guidance content to the user via voice,
[1088] A means of analyzing user emotions and adjusting the tone of guidance,
[1089] A means of recognizing user input via speech and providing additional information,
[1090] A system that includes means for recording user responses and using them as training data.
[1091] (Claim 2)
[1092] The system according to claim 1, further comprising means for using a display device via a terminal to superimpose and display information data onto the real field of view.
[1093] (Claim 3)
[1094] The system according to claim 1, further comprising means for improving the personalization accuracy of the generated guidance content based on the aforementioned learning data. [Explanation of Symbols]
[1095] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for obtaining the user's location information, A means for filtering information on tourist spots based on the aforementioned location information and user interest / history data, A means for generating customized guidance content based on the filtered information using a generative artificial intelligence model, A means of providing the generated guidance content to the user via voice, A means of recognizing user input via speech and providing additional information, A means of recording user responses and using them as training data, A means of obtaining and providing real-time tourist information within cities, A system that includes a means of providing additional details via voice based on user questions.
2. The system according to claim 1, further comprising means for using a camera via a terminal to superimpose and display digital information onto a real-world landscape, and means for providing information specifically for urban tourism.
3. The system according to claim 1, further comprising means for improving the personalization accuracy of the generated guidance content based on the aforementioned learning data, and means for analyzing the user's interests based on their movement history in the city.
Citation Information
Patent Citations
JP2022180282A