System

The AI and AR-based system addresses the guide shortage and cultural engagement issues in tourist destinations by offering personalized, locally dialect-specific guidance and information updates.

JP2026023465APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024125400
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Regional tourist destinations face a labor shortage of tourist guides, lack of information on lesser-known sites, and insufficient engagement for solo travelers, leading to inadequate guidance and cultural experience.

Method used

A system utilizing AI characters speaking local dialects and augmented reality to provide tourist information, answer questions, and update databases with user input, enhancing guidance and cultural engagement.

Benefits of technology

The system addresses the guide shortage by providing personalized, culturally immersive experiences and up-to-date information through AI and AR, improving visitor satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026023465000001_ABST
    Figure 2026023465000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring location information; means for searching for information on a tourist attraction based on the acquired location information; means for providing the searched information on the tourist attraction to a user; means for displaying a AI character that speaks in a local dialect and providing a sightseeing guide; means for receiving a question from the user, generating an answer according to the content of the question, and providing the answer to the user; means for AR-displaying the AI character based on the location information of the user; and means for storing tourist spot information provided by the user in a database and updating a AI model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In regional tourist destinations, the aging of tourist guides is causing a serious labor shortage, making it difficult to provide sufficient information and guidance on tourist attractions. Furthermore, in minor tourist destinations, famous sites are not well known, making it difficult to attract tourists' interest. Furthermore, with the number of solo travelers on the rise, there is a lack of ways to alleviate feelings of loneliness and provide tourist experiences that allow for a deeper connection to local culture. To solve these issues, a new guide system is needed that can respond to individual tourist demands without relying on the aging population or labor shortages. [Means for solving the problem]

[0005] The present invention is a system that includes a means for acquiring location information, a means for searching for tourist attraction information based on the acquired location information, a means for providing the user with the searched tourist attraction information, a means for displaying an AI character that speaks in the local dialect and providing a tour guide, a means for accepting questions from the user and generating answers based on the question content and providing them to the user, a means for displaying the AI ​​character in AR based on the user's location information, and a means for storing tourist attraction information provided by the user in a database and updating the AI ​​model. This system can alleviate the shortage of tour guides and information on famous places in regional tourist destinations. Furthermore, by using an AI character that speaks in the local dialect and AR functions, it is possible to provide tourists with a deeper local experience and reduce their sense of loneliness.

[0006] "Location Information" means data that indicates the physical location of a particular device or user using GPS or other location-determining technology.

[0007] "Information about tourist attractions" refers to all the information that visitors want to know, such as the highlights of the tourist destination and its surrounding areas, its history, culture, and how to get there.

[0008] A "local dialect" refers to the unique words, accents, and expressions used in a particular region, and is a linguistic form based on the culture and history of that region.

[0009] An "AI character" is a character based on artificial intelligence, a computer-generated virtual personality that interacts with users and provides information.

[0010] "AR display" is a technology that uses augmented reality to overlay computer graphics and information onto real-world images.

[0011] A "database" is a collection of data organized for efficient storage, retrieval, and management.

[0012] An "AI model" is software trained using machine learning algorithms, a mathematical model used to make predictions or classifications for specific tasks.

[0013] "Natural language processing (NLP)" is a general term for technologies and methods that enable computers to understand, generate, and respond to human language (natural language).

[0014] "Training data" is a dataset used to build and train an AI model, and is used as teacher data to improve the model's accuracy. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention is a system that complements the lack of tourist guides in local tourist destinations and provides tourists with a new experience. This system utilizes location information and features an AI character that speaks in the local dialect to provide tourist guidance. Below, we will explain each component of the system and its operation in detail.

[0037] Location information acquisition means

[0038] Device:

[0039] When a user visits a tourist spot, the device acquires the current location information using GPS, which is then sent to the server.

[0040] Tourist attraction information search methods

[0041] server:

[0042] Based on the received location information, the server searches a database for information on nearby tourist attractions, including detailed descriptions, photos, and reviews of the attractions.

[0043] Tourist attraction information provision method

[0044] server:

[0045] The searched tourist attraction information is transmitted to the user terminal.

[0046] Device:

[0047] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is displayed.

[0048] Tourist guide using AI characters in local dialects

[0049] Device:

[0050] The tourist attraction information displayed is provided to the user by an AI character speaking in the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[0051] Chat-style question and answering tool

[0052] User:

[0053] Users can type questions into a chat window on their device (e.g., "What are some recommended restaurants nearby?").

[0054] Device:

[0055] The entered question is sent to the server.

[0056] server:

[0057] The received question is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[0058] Device:

[0059] The generated answer is displayed in the chat window, and the AI ​​character also responds verbally in the local dialect.

[0060] AR display guide

[0061] User:

[0062] The user selects the AR mode on the device and activates the camera.

[0063] Device:

[0064] The camera captures the video and sends it to the server in real time.

[0065] server:

[0066] Based on the location information and captured video, the display position of the AI ​​character is calculated and sent to the device.

[0067] Device:

[0068] Based on the received data, an AI character is superimposed on the captured image. For example, if you stand in front of a tourist spot, an AI guide character will appear in AR and begin explaining.

[0069] A learning tool for user-provided information on famous places

[0070] User:

[0071] Users post photos and descriptions of new tourist spots to the app (e.g., when they discover a new photo spot).

[0072] Device:

[0073] The post content is sent to the server.

[0074] server:

[0075] The new information is stored in a database and used as training data for the AI ​​model.

[0076] server:

[0077] The learning results will be used to provide the latest tourist spot information to other users the next time they visit the same location.

[0078] Specific examples

[0079] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server obtains information about the peak of a nearby tourist attraction called "Mount XX," and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[0080] As a result, the system of the present invention can solve problems such as a lack of tourist guides and information specific to a particular region, and provide tourists with new experiences.

[0081] The processing flow will be explained below.

[0082] Specific processing steps will be described below.

[0083] Step 1:

[0084] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[0085] Step 2:

[0086] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[0087] Step 3:

[0088] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and configured.

[0089] Step 4:

[0090] Server: Sends the data of the configured tourist guide AI character to the terminal.

[0091] Step 5:

[0092] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[0093] Step 6:

[0094] User: Visits a tourist spot and starts moving around with the device in hand.

[0095] Step 7:

[0096] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[0097] Step 8:

[0098] Server: Based on the received location information, search the database for nearby tourist attractions.

[0099] Step 9:

[0100] Server: Generates relevant tourist attraction information and sends it to the terminal.

[0101] Step 10:

[0102] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[0103] Step 11:

[0104] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[0105] Step 12:

[0106] Terminal: Sends the entered question to the server.

[0107] Step 13:

[0108] Server: Analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[0109] Step 14:

[0110] Server: Sends the generated answer to the device.

[0111] Step 15:

[0112] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[0113] Step 16:

[0114] User: Select the AR mode on the device and activate the camera.

[0115] Step 17:

[0116] Terminal: Captures camera images and sends them to the server in real time.

[0117] Step 18:

[0118] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the data to the device.

[0119] Step 19:

[0120] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI provides tourist information using AR display.

[0121] Step 20:

[0122] Users: When they discover a new tourist spot, they post a photo and description to the app.

[0123] Step 21:

[0124] Device: Sends the posted content (photos, text, location information, etc.) to the server.

[0125] Step 22:

[0126] Server: Stores new tourist spot information in a database and uses it as training data for the AI ​​model.

[0127] Step 23:

[0128] Server: Updates the AI ​​model based on the training data to improve the accuracy of the tourist guide AI.

[0129] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[0130] Example 1

[0131] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0132] In regional tourist destinations, a lack of tourist guides means that tourists cannot obtain sufficient information. Furthermore, tourist information is often provided in standard Japanese, leaving few opportunities to experience the unique culture and dialects of the region. Furthermore, there is a lack of systems that can quickly and accurately respond to the diverse questions tourists ask. Another issue is how to efficiently incorporate new tourist information provided by users and utilize it in future tourism services.

[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0134] In this invention, the server includes a means for acquiring location information, a means for searching for information on tourist spots based on the acquired location information, a means for providing the user with the searched information on tourist spots, a means for displaying an AI character that speaks in the local dialect and providing tourist information, a means for accepting questions from the user, generating answers based on the questions, and providing the answer to the user, a means for displaying the AI ​​character in augmented reality based on the user's location information, and a means for storing the tourist spot information provided by the user in a database and updating the AI ​​model. This solves the problem of a lack of tourist guides and allows users to have new experiences through local information and dialects. Furthermore, the server can respond quickly and accurately to user questions and provide the latest tourist information that is continuously updated.

[0135] "Location information" is information that indicates the user's current geographical location.

[0136] "Means of acquisition" refers to the method or apparatus by which a device or system collects location information or other data.

[0137] "Tourist destination" refers to a place or attraction that tourists are expected to visit.

[0138] A "search means" refers to a method or device for finding required information from a database or information source based on specific criteria.

[0139] "Means for providing" refers to the method or device for making information or services available to users.

[0140] "Local dialect" refers to linguistic expressions and ways of speaking that are unique to a region.

[0141] "Artificial intelligence characters" refer to virtual people or characters with specific roles that are generated using AI technology.

[0142] "Tourist information" refers to activities and services that provide information and guide tourists about tourist destinations and attractions.

[0143] "Means for accepting queries" refers to a method or device for accepting and processing inquiries or questions from users.

[0144] "Answer generation means" refers to a method or device that generates an appropriate response to a user's question.

[0145] "Augmented reality display" refers to a technology that overlays digital information on the real world.

[0146] A "database" refers to a collection of information that is systematically organized and made available for efficient search and use.

[0147] An "artificial intelligence model" refers to a conceptual model of an AI system that is trained based on large amounts of data to perform specific tasks.

[0148] The present invention is a system for supplementing the lack of tourist guides in tourist destinations and providing tourists with information and experiences unique to the region. This system utilizes location information and provides tourist guidance using an AI character that speaks in the local dialect. Specific embodiments of this system are described in detail below.

[0149] Hardware and Software Configuration

[0150] Device:

[0151] The device is a user's smartphone or tablet. This device is equipped with a GPS function, a camera, a microphone, and a speaker. The device also has network connectivity (Wi-Fi or mobile data).

[0152] server:

[0153] A cloud server is used to perform location analysis, database management, natural language processing (NLP), and augmented reality (AR) calculations. The following software and services run on the server.

[0154] DBMS (Database Management System): Manages tourist spot information, user-provided information, etc.

[0155] NLP engine: Analyzes user questions and generates appropriate answers

[0156] AR engine: Calculates the display position of the AI ​​character

[0157] Program processing overview

[0158] Device:

[0159] 1. When a user visits a tourist spot, the device uses GPS to obtain current location information, which is then sent to the server in real time.

[0160] 2. The received tourist spot information is displayed on the user interface. For example, it might say, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[0161] 3. The displayed tourist attraction information is then provided to the user by an AI character speaking in the local dialect. For example, the guide will say, "If you go straight ahead, you will come to a famous local shrine."

[0162] server:

[0163] 1. The server searches the database based on the received location information and obtains information about tourist attractions in the vicinity of the location (detailed descriptions, photos, reviews, etc.).

[0164] 2. Analyze the received question using natural language processing technology and generate an appropriate answer.

[0165] 3. Based on the received camera footage and location information, the display position of the AI ​​character is calculated and sent to the device.

[0166] Specific examples

[0167] As a concrete example, consider the case where a user visits a mountain called "Mount XX" in a local tourist spot. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server retrieves information about the peak of nearby tourist spot "Mount XX" from a database and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[0168] Prompt Sentence Examples

[0169] The following prompt sentences could be input to the generative AI model:

[0170] Example: "Please explain in detail the system that displays an AI character that guides users in the local dialect when they visit a local tourist spot."

[0171] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[0172] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0173] Step 1:

[0174] Terminal: When a user visits a tourist spot, the terminal uses a GPS module to obtain current location information. The input is the user's physical location information, and the output is GPS coordinate data (latitude and longitude). Specifically, the GPS module determines the location of the spot, and the data is processed internally on the terminal to generate coordinate information. This coordinate information is then sent to the server.

[0175] Step 2:

[0176] Terminal: Sends the acquired location information to the server. The input is the GPS coordinate data acquired in step 1, and the output is the location data sent to the server. Specifically, the terminal uses a network communication module (e.g., Wi-Fi or mobile data communication) to send the location data to the specified endpoint on the server.

[0177] Step 3:

[0178] Server: Based on the received location information, the server searches a database. The input is location data, and the output is information about tourist attractions near the location. Specifically, the server performs a database query to obtain detailed information (e.g., name, description, photos, reviews, etc.) about the nearest tourist attraction based on the location information.

[0179] Step 4:

[0180] Server: Sends the searched tourist attraction information to the user's device. The input is the tourist attraction information retrieved from the database, and the output is the tourist attraction information data sent to the device. Specifically, the server packages the tourist attraction information in an appropriate data format (e.g., JSON format) and sends it to the user's device via the network.

[0181] Step 5:

[0182] Terminal: The terminal displays the information it receives on its user interface. The input is tourist attraction information data sent from the server, and the output is tourist attraction information that is visually displayed to the user. Specifically, the terminal parses the received JSON data and displays the information in an appropriate format on the application's user interface. For example, it might display, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[0183] Step 6:

[0184] Terminal: An AI character speaking in the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is the visually displayed tourist attraction information, and the output is audio guidance information. Specifically, the terminal inputs text data into a speech synthesis engine (e.g., a TTS engine), which generates and plays audio data in the local dialect.

[0185] Step 7:

[0186] User: The user types a question into the chat window of the terminal. The input is the user's text-based question, and the output is the question data stored in the terminal. In concrete terms, the user types a question into the chat window of the application and presses the send button.

[0187] Step 8:

[0188] Terminal: Sends the entered question to the server. The input is the question data from the user, and the output is the question data sent to the server. In concrete terms, the terminal sends the question data to the server via the network communication module.

[0189] Step 9:

[0190] Server: Analyzes the received question using natural language processing technology and generates an appropriate answer. The input is the user's question data, and the output is the generated answer data. Specifically, the server uses a natural language processing engine to analyze the question and generate the optimal answer from a database or knowledge base.

[0191] Step 10:

[0192] Server: Sends the generated answer to the terminal. The input is the answer data, and the output is the answer data sent to the terminal. Specifically, the server packages the answer data in an appropriate data format and sends it to the user's terminal via the network.

[0193] Step 11:

[0194] Terminal: The terminal displays the generated answer in the chat window, and the AI ​​character also answers aloud in the local dialect. The input is the answer data sent from the server, and the output is the answer information provided visually and aloud. Specifically, the terminal analyzes the received answer data and displays it as text in the chat window. It also uses a speech synthesis engine to generate and play back audio data in the local dialect.

[0195] Step 12:

[0196] User: The user selects the AR mode on the device and activates the camera. The input is the user's operation, and the output is the video data captured by the camera. Specifically, the user presses the AR mode button in the application to activate the device's camera.

[0197] Step 13:

[0198] Terminal: Captures camera images and sends them to the server in real time. The input is the camera's image data, and the output is the image data sent to the server. Specifically, the terminal acquires image data from the camera module, compresses it, and sends it to the server.

[0199] Step 14:

[0200] Server: Based on location information and camera footage, calculates the display position of the AI ​​character and sends this to the device. The input is location information and video data, and the output is character display position data. Specifically, the server uses the AR engine to analyze the local geographic information and video data and calculate the optimal character display position.

[0201] Step 15:

[0202] Terminal: Based on the received data, an AI character is superimposed on the camera image. The input is character display position data and camera image data, and the output is the character displayed in AR. In concrete terms, the terminal superimposes a virtual character on the camera image based on the received position data. For example, when you stand in front of a tourist attraction, an AI guide character will appear in AR and begin to explain.

[0203] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[0204] (Application example 1)

[0205] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0206] Currently, there is a significant shortage of tourist guides and store guides in regional tourist destinations and shopping malls. Furthermore, tourists and visitors have limited access to information specific to their region, raising concerns about declining satisfaction. Furthermore, information about tourist destinations and stores is often provided in standard Japanese, which means that the characteristics and atmosphere of the region cannot be fully conveyed. This makes it difficult to maximize the appeal of tourist destinations and shopping malls and provide visitors with personalized, attractive guides. Therefore, there is a need to develop a system that can efficiently provide tourist guides and store guides and offer new experiences to visitors.

[0207] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0208] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving tourist attraction information provided by the user in a database and updating the AI ​​model, means for searching for store information within a shopping mall from the database, means for using an AI character that provides guidance in the local dialect and displaying it on a smartphone camera image, means for saving new store information and photos posted by the user in the database, and means for users to input questions in a chat format and for analyzing and answering them using natural language processing technology. This enables tourist spots and shopping malls to efficiently provide visitors with personalized guides and information using the local dialect.

[0209] "Location information acquisition means" refers to a system in which a user's terminal acquires the current location using GPS or other technology.

[0210] The "tourist attraction information search means" is a function that allows the server to search the database for information on the relevant tourist attraction based on the acquired location information.

[0211] The "tourist attraction information providing means" is a function that transmits information about the searched tourist attractions to the user terminal and displays it on the user interface.

[0212] "AI characters that speak in local dialects" refers to artificial intelligence characters that provide audio guidance using a dialect specific to the region.

[0213] The "question answering means" is a function that accepts questions from users, analyzes them, and generates and provides appropriate answers.

[0214] "Augmented reality display means" is a technology that displays virtual information superimposed on real-world images, and in this case refers to the function of displaying an AI character superimposed on a real-world scene.

[0215] The "user-provided information storage means" is a function that stores information and photos of tourist spots provided by users in a database.

[0216] "AI model updating means" refers to a method for retraining an artificial intelligence model based on stored user-provided information to improve its accuracy.

[0217] The "shopping mall information search means" is a function for searching for information on nearby stores based on the current location within the shopping mall.

[0218] "Natural language processing technology" is a technology that allows a computer to analyze a user's question, understand human language, and generate an appropriate answer.

[0219] This system complements the lack of tourist guides and store guides in shopping malls and local tourist destinations, providing visitors with a personalized new experience. This system uses an AI character that provides audio guidance in the local dialect and displays the information on the smartphone camera using augmented reality (AR). It also includes a function that uses natural language processing (NLP) to generate appropriate answers to user questions.

[0220] First, when a user visits a tourist spot or shopping mall, the location information acquisition means uses the smartphone's GPS to acquire current location information. This location information is sent to the server. Based on the received location information, the server searches a database for information on tourist attractions and stores in the vicinity of the location. For example, it provides information such as, "There is a recommended restaurant 200 meters from your current location."

[0221] Next, the tourist attraction information provider sends the searched tourist attraction and store information to the user's smartphone. The displayed information is guided by an AI character speaking the local dialect. For example, guidance such as "If you turn right at the next street, you will find a delicious Japanese restaurant" is provided.

[0222] The system also includes a chat-style question-and-answer mechanism. Users enter questions into a chat window on their smartphone, which is then sent to the server. The server uses natural language processing technology to analyze the question and generate an appropriate answer. For example, in response to the question, "What are some recommended cafes nearby?", the server will respond with, "The recommended nearby cafe is on the second floor of the mall."

[0223] Furthermore, as an augmented reality display method, when a user activates the smartphone camera and selects AR mode, an AI character is superimposed on the camera image, allowing the user to receive guidance while viewing the AI ​​character over the real-world scene.

[0224] The user-provided information storage means provides a mechanism that allows users to post new tourist spots and store information to the app. The posted information is sent to the server and stored in a database. This information is used as learning data for the AI ​​model, and the next time other users visit the same location, more up-to-date information will be provided.

[0225] As a concrete example, consider a smartphone app for tourists visiting a local shopping mall. As a user walks through the mall, their current location is acquired via GPS and sent to a server. The server uses this location information to search for nearby store information and provides information such as, "There's a new cafe 200 meters from your current location." When a user asks in a chat window, "What are some recommended restaurants nearby?", an answer analyzed using NLP technology is displayed. In addition, when using AR mode, an AI character appears on the camera image, providing specific directions.

[0226] The tour guide using the generative AI model generates a detailed scenario using the following prompt sentence:

[0227] ---

[0228] When a user visits a local shopping mall, provide an application that uses the smartphone's location information to provide voice guidance on recommended stores and tourist spots in the area. The application should include a feature that uses an AI character that speaks the local dialect, overlaying the AI ​​character on the screen in AR, and provide specific guidance. It should also have the ability to search a database based on GPS coordinates, analyze user questions using NLP technology, and allow users to post new store information. Please explain with specific examples.

[0229] ---

[0230] In this way, the present invention aims to make up for the lack of guides at tourist spots and shopping malls and provide visitors with a new experience.

[0231] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0232] Step 1:

[0233] The server receives GPS location information sent from the user's device. The input is the coordinate data of the current location obtained from the smartphone's GPS sensor. Based on this data, the server searches a database for information on relevant tourist attractions and stores in the shopping mall. As output, it generates information on the tourist attractions and stores found.

[0234] Step 2:

[0235] The server sends the searched tourist attraction and store information to the user terminal. The tourist attraction and store information generated in step 1 is the input. The server formats this and sends it to the user terminal in an appropriate format. The output is information data to be displayed on the user terminal.

[0236] Step 3:

[0237] The device uses an AI character to provide audio guidance in the local dialect based on the received information. The inputs include tourist attraction and store information sent from the server and the AI ​​character's audio data. The device combines this information to output audio and display on the screen. The output is audio guidance and visual guidance information presented to the user.

[0238] Step 4:

[0239] The user enters a question into the chat window on the smartphone. The input is a text question. The device sends this to the server. The output is the user question data.

[0240] Step 5:

[0241] The server analyzes the received question using natural language processing (NLP) technology and generates an appropriate answer. The input is the question data entered by the user. The server inputs this data into an NLP model and generates an answer based on the analysis results. The output is an appropriate answer text.

[0242] Step 6:

[0243] The server sends the generated answer to the user terminal. The input is the answer text generated in step 5. The server formats it and sends it to the user terminal. The output is the answer data to be displayed on the user terminal.

[0244] Step 7:

[0245] The device displays the received response in the user's chat window and responds audibly if necessary. The inputs are the response data sent from the server and the AI ​​character's voice data. The device combines these and outputs text and voice. The output is a visual and voice response to the user.

[0246] Step 8:

[0247] The user activates the smartphone camera and selects AR mode. The input is the camera image. The device captures it and sends it to the server. The output is real-time video data sent to the server.

[0248] Step 9:

[0249] The server calculates the display position of the AI ​​character based on the received camera image and location information, and sends this data to the device. The inputs are camera image data and GPS location information. The server processes this data and generates display position data for the AI ​​character. The output is sent to the device.

[0250] Step 10:

[0251] The device overlays the AI ​​character on the camera image based on the received AI character position data. The inputs are camera image data and AI character position data. The device combines these and displays them. The output is an augmented reality image from the user's perspective.

[0252] Step 11:

[0253] Users post information about new tourist spots and shops to the app. The input is text and image data. The device sends this to the server. The output is the posted data sent to the server.

[0254] Step 12:

[0255] The server stores the received posting data in a database and uses it as training data for the AI ​​model. The input includes tourist spot and store information data posted by users. The server stores this in the database and retrains the AI ​​model to improve its accuracy. As output, the new information data is stored in the database and an updated AI model is generated.

[0256] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0257] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR display, and an emotion engine to realize a tourist guide that responds to the user's emotions. Each component and its operation are explained in detail below.

[0258] Location information acquisition means

[0259] Device:

[0260] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[0261] Tourist attraction information search methods

[0262] server:

[0263] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[0264] Tourist attraction information provision method

[0265] server:

[0266] Tourist attraction information from search results is sent to the device.

[0267] Device:

[0268] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[0269] Tourist guide using AI characters in local dialects

[0270] Device:

[0271] The displayed tourist spot information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[0272] Chat-style question and answering tool

[0273] User:

[0274] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[0275] Device:

[0276] The entered question is sent to the server.

[0277] server:

[0278] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[0279] Device:

[0280] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[0281] AR display guide

[0282] User:

[0283] The user selects the AR mode on the device and activates the camera.

[0284] Device:

[0285] The camera captures the video and sends it to the server in real time.

[0286] server:

[0287] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[0288] Device:

[0289] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides sightseeing information.

[0290] A learning tool for user-provided information on famous places

[0291] User:

[0292] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[0293] Device:

[0294] The post content is sent to the server.

[0295] server:

[0296] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[0297] server:

[0298] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[0299] Response methods using emotion engines

[0300] Device:

[0301] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[0302] server:

[0303] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[0304] Device:

[0305] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[0306] Specific examples

[0307] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[0308] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In addition, by using AR mode, an AI character will appear on the camera image, providing visual guidance to the destination.

[0309] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas to help the user continue sightseeing in comfort.

[0310] As a result, the system of the present invention can provide a personalized tourist experience by addressing the lack of tourist guides and information specific to the region, as well as responding appropriately to the user's emotions.

[0311] The processing flow will be explained below.

[0312] Specific processing steps will be described below.

[0313] Step 1:

[0314] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[0315] Step 2:

[0316] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[0317] Step 3:

[0318] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and set up.

[0319] Step 4:

[0320] Server: Sends the data of the configured tourist guide AI character to the terminal.

[0321] Step 5:

[0322] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[0323] Step 6:

[0324] User: Visits a tourist spot and starts moving around with the device in hand.

[0325] Step 7:

[0326] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[0327] Step 8:

[0328] Server: Based on the received location information, search the database for nearby tourist attractions.

[0329] Step 9:

[0330] Server: Generates relevant tourist attraction information and sends it to the terminal.

[0331] Step 10:

[0332] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[0333] Step 11:

[0334] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[0335] Step 12:

[0336] Terminal: Sends the entered question to the server.

[0337] Step 13:

[0338] Server: Analyzes the question using natural language processing technology and generates an appropriate answer.

[0339] Step 14:

[0340] Server: Sends the generated answer to the device.

[0341] Step 15:

[0342] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[0343] Step 16:

[0344] User: Select the AR mode on the device and activate the camera.

[0345] Step 17:

[0346] Terminal: Captures camera images and sends them to the server in real time.

[0347] Step 18:

[0348] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the 3D model data and location information to the device.

[0349] Step 19:

[0350] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI character provides sightseeing information.

[0351] Step 20:

[0352] Users: When they discover a new tourist spot, they post a photo and description to the app.

[0353] Step 21:

[0354] Terminal: Sends the posted content to the server.

[0355] Step 22:

[0356] Server: Saves new tourist spot information in a database and uses it as training data for the AI ​​model.

[0357] Step 23:

[0358] Server: Updates the AI ​​model based on the learning results and provides the latest tourist spot information to other users from the next time onwards.

[0359] Step 24:

[0360] Terminal: Analyzes the user's facial expressions and voice, and recognizes their emotional state using an emotion engine.

[0361] Step 25:

[0362] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine.

[0363] Step 26:

[0364] Terminal: Providing information according to emotions.

[0365] Specific examples

[0366] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of a nearby tourist spot called "Mount XX," and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters away from here," and sends it to the device.

[0367] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In addition, by using AR mode, an AI character will appear on the camera image, providing visual guidance to the destination.

[0368] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas to help the user continue sightseeing in comfort.

[0369] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[0370] Example 2

[0371] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0372] In traditional tourist destinations, the number of guides is limited, making it difficult to provide individualized and satisfying tours to all visitors. There is also a lack of information that responds to visitors' emotions and personal preferences. Furthermore, there is a lack of technological means to provide advanced guidance and guidance at tourist destinations. This prevents visitors from fully enjoying the destinations, hindering the revitalization of the tourism industry as a whole.

[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0374] In this invention, the server includes a means for acquiring location information, a means for searching for tourist attraction information, and a means for providing tourist attraction information to users, thereby making it possible to provide personalized tourist information to visitors and increase their satisfaction.

[0375] "Location information" is data that indicates a user's current geographic location.

[0376] "Tourist attraction" refers to a place or facility that is of interest to visitors for tourism purposes.

[0377] "Searching for information" is the act of locating relevant information from databases or sources based on specific criteria.

[0378] "Providing" is the act of giving or displaying information or services to a user.

[0379] An "artificial intelligence character" is a virtual guide character programmed using AI technology.

[0380] "Augmented reality display" is a technology that displays digital information and objects overlaid on real-world scenery.

[0381] A "database" is a collection of information for efficiently storing, searching, and managing large amounts of data.

[0382] An "artificial intelligence model" is an implementation of an AI algorithm that learns and makes inferences based on data.

[0383] "Emotional state" is information that indicates the psychological and physiological state of the user.

[0384] This invention is a system that compensates for the lack of guides in local tourist destinations and provides personalized tourist experiences. This system incorporates GPS functionality, natural language processing (NLP) technology, AR technology, and an emotion recognition engine. Below, we will explain in detail each component that makes up the system and its operation.

[0385] Location information acquisition means

[0386] Device: When a user visits a tourist spot, the device uses its GPS function to obtain current location information. This location information is periodically sent to the server. For example, the GPS module of a smartphone can be used to obtain the user's location and send it to the server as an HTTP request.

[0387] Tourist attraction information search methods

[0388] Server: Based on the location information received from the device, the server searches a database for information on nearby tourist attractions. This database contains detailed descriptions of tourist attractions, photos, user reviews, etc. The server executes SQL queries to search for this information.

[0389] Tourist attraction information provision method

[0390] Server: Sends tourist attraction information obtained as search results to the terminal.

[0391] Terminal: The terminal displays the information received from the server on a user interface. For example, it may display information such as, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[0392] Tourist guide using AI characters in local dialects

[0393] Device: The displayed tourist attraction information is provided to the user by an AI character that speaks the local dialect. For example, the device uses a speech synthesis engine to generate voice guidance in the local dialect, such as "If you go straight ahead, you will find a famous local shrine."

[0394] Chat-style question and answering tool

[0395] User: If the user has a question, they type it into the chat window on their device. For example, "What are some recommended restaurants nearby?"

[0396] Terminal: The terminal sends the entered question to the server.

[0397] Server: The server analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[0398] Device: The device will display the answers received from the server in a chat window, and the guide AI character will read the answers aloud in the local dialect.

[0399] AR display guide

[0400] User: The user selects the AR mode on the device and activates the camera.

[0401] Terminal: The terminal captures the camera image and transmits it to the server in real time.

[0402] Server: The server calculates the display position of the AI ​​character based on the location information and captured video, and sends the AI ​​character's 3D model data and location information to the device.

[0403] Terminal: Based on the received data, the terminal displays an AI character superimposed on the camera image. The guide AI character provides sightseeing information.

[0404] A learning tool for user-provided information on famous places

[0405] User: When a user finds a new tourist spot, they post a photo and description of it to the app. For example, they post something like, "I found a new photo spot."

[0406] Terminal: The terminal sends the posted content to the server.

[0407] Server: The server saves new tourist spot information in a database and uses it as learning data for the AI ​​model. It then updates the AI ​​model based on the learning results, enabling it to provide the latest tourist spot information to other users from the next time onwards.

[0408] Response methods using emotion engines

[0409] Device: The device analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[0410] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine.

[0411] Device: The device provides information based on the user's emotions. For example, if the device detects that the user is tired, it will suggest nearby rest spots or cafes.

[0412] Specific examples

[0413] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[0414] Based on this information, the device will have an AI guide character provide voice guidance in the local dialect. For example, it might say, "500 meters ahead there is a peak with a spectacular view." In addition, when using AR mode, the AI ​​character will appear on the camera image, providing visual guidance to the destination.

[0415] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas. This allows the system of the present invention to address the lack of tourist guides and local information, and respond appropriately to the user's emotions, providing a personalized sightseeing experience.

[0416] Prompt Sentence Examples

[0417] The following are some examples of input prompts for a generative AI model:

[0418] "Please introduce us to a new tourist guide system. This system will address the lack of guides in local tourist destinations and provide a personalized tourist experience. Specifically, it will use location information acquisition, AI characters that speak the local dialect, AR displays, and an emotion engine to provide a tourist guide that responds to the user's emotions."

[0419] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0420] Step 1: Obtaining location information

[0421] Terminal: When a user visits a tourist spot, the terminal's GPS function is activated and acquires current location information. The input is geographic coordinate data from the GPS sensor, and the output is current location information. This location information is sent to the server at regular intervals.

[0422] Specific operation: The device uses the built-in GPS module to obtain geographic coordinates (latitude and longitude) and transmits them to the server as location information.

[0423] Step 2: Search for tourist attraction information

[0424] Server: Receives the acquired location information and searches for information on nearby tourist attractions. The input is the location information and the output is a list of nearby tourist attractions. The server searches the database using an SQL query and obtains the search results.

[0425] Specific operation: Based on the received location information, the server performs a range search on the tourist attraction database to obtain detailed information about related tourist attractions.

[0426] Step 3: Providing tourist attraction information

[0427] Server: Sends tourist attraction information obtained as search results to the terminal. The input is tourist attraction information, and the output is the transmission of information to the terminal.

[0428] Specific operation: The server composes the search results in JSON format and sends them to the terminal as an HTTP response.

[0429] Step 4: View tourist attraction information

[0430] Terminal: The terminal displays tourist attraction information received from the server on a user interface. The input is tourist attraction information from the server, and the output is the information displayed on the display.

[0431] Specific operation: The terminal parses the received JSON data, binds the information to the user interface, and displays it.

[0432] Step 5: Tourist guide in local dialect

[0433] Terminal: An AI character speaking the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is text information about the tourist attraction, and the output is synthesized speech.

[0434] Specific operation: The device passes the text information to a TTS (Text-to-Speech) engine, synthesizes it in the local dialect, and plays it back through the speaker.

[0435] Step 6: Chat-style Q&A

[0436] User: The user types a question into the chat window on the device. The input is the text of the user's question.

[0437] Terminal: The terminal sends the input question to the server. The output is the question data sent to the server.

[0438] Specific behavior: The user types a question into the chat interface, the device captures it, and sends it to the server via an HTTP POST request.

[0439] Server: The server analyzes the question content using NLP technology and generates an appropriate answer. The input is the user's question text, and the output is the generated answer text.

[0440] What it does: The server uses an NLP model to analyze the question and generate an appropriate answer from a database or other source.

[0441] Terminal: The terminal displays the answer received from the server in a chat window and has an AI character read it aloud in the local dialect. The input is the answer text from the server, and the output is the displayed answer and audio.

[0442] Specific operation: The device displays the received reply text in the chat window, synthesizes it using the TTS engine, and plays it back.

[0443] Step 7: Guided by AR display

[0444] User: The user selects the device's AR mode and activates the camera. The input is the AR mode selection action.

[0445] Terminal: The terminal captures camera images and transmits them to the server in real time. The input is the camera image and the output is the video data sent to the server.

[0446] Specific operation: The device captures camera footage and streams it to the server along with location information.

[0447] Server: The server calculates the display position of the AI ​​character based on the location information and video data, and sends the 3D model data to the terminal. The input is the camera image and location information, and the output is the display position of the AI ​​character and 3D model data.

[0448] Specific operation: The server uses the AR calculation engine to calculate the appropriate display position and send the 3D model data to the terminal.

[0449] Terminal: Based on the received data, the terminal displays an AI character superimposed on the camera image. The input is 3D model data and display position information, and the output is an AI character superimposed on the camera image.

[0450] What it does: The device uses an AR library to render a 3D model overlaid on the camera image.

[0451] Step 8: Learning user-provided points of interest information

[0452] User: When a user finds a new tourist spot, he / she posts a photo and description of the spot to the app. The input is a photo of the new tourist spot and text information.

[0453] Terminal: The terminal sends the posted content to the server. The output is the user posted data sent to the server.

[0454] Specific operation: The device collects posting data including information entered by the user and sends it to the server as an HTTP POST request.

[0455] Server: The server stores the received information in a database and uses it as training data for the AI ​​model. The input is new tourist attraction information, and the output is an updated database and AI model.

[0456] Specific operation: The server stores the received post data in a database and reflects it in the AI ​​model in the next learning phase.

[0457] Step 9: Respond with the Emotion Engine

[0458] Terminal: The terminal analyzes the user's facial expressions and voice and recognizes their emotional state using an emotion engine. The input is the user's facial expression data and voice data, and the output is the recognized emotional information.

[0459] How it works: The device uses a camera and microphone to capture the user's facial expressions and voice, and applies emotion analysis algorithms to identify their emotional state.

[0460] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine. The input is the recognized emotion information, and the output is adjusted tourist attraction information.

[0461] Specific operation: The server dynamically adjusts the method and content of providing tourist attraction information based on the emotion data.

[0462] Terminal: The terminal provides tailored information to the user. For example, if the terminal recognizes that the user is tired, it suggests nearby rest spots or cafes. The input is tailored tourist attraction information, and the output is the provided information and guidance.

[0463] Specific operation: Based on the adjusted information, the device updates the user interface and voice guidance to provide the user with the most appropriate information.

[0464] (Application example 2)

[0465] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0466] There is a shortage of tourist guides in local tourist destinations, and it is difficult to provide tourists with a personalized guided experience. Furthermore, there is a demand for dynamic information provision based on the user's emotions and location information. Furthermore, tourist guides using local dialects and visual guidance using augmented reality displays are insufficient. A means is needed to address these issues and provide a richer tourist experience.

[0467] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0468] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving the tourist attraction information provided by the user in a database and updating the AI ​​model, means for analyzing emotions from the user's facial expressions and voice and providing information based on the emotions, means for capturing camera footage and transmitting it to the server in real time, and means for calculating the display position of the AI ​​character based on the location information and the captured footage and superimposing the AI ​​character on the camera footage. This makes it possible to address the lack of tourist destination guides and local information, and to provide personalized tour guidance based on the user's emotions, as well as a visual guide experience.

[0469] Key Word Definitions

[0470] The "means for acquiring location information" is a means for acquiring the user's current location information using the GPS function and periodically transmitting it to the server.

[0471] The "means for searching tourist attraction information" is a means for searching a database for information on tourist attractions that exist in the vicinity of a point based on the acquired location information.

[0472] The "means for providing tourist attraction information to the user" refers to a means for displaying information about the searched tourist attraction to the user through a user interface.

[0473] "A means for displaying an AI character that speaks in the local dialect and provides tourist guidance" is a means for an AI character displayed to the user to provide tourist guidance using the local dialect.

[0474] The "means for generating an answer according to the question content and providing it to the user" is a means for analyzing the question entered by the user, generating an appropriate answer, and providing it to the user in a chat window or by voice.

[0475] "Means for displaying AI characters in augmented reality" refers to a means for displaying a 3D model of an AI character overlaid on camera images based on location information and captured images.

[0476] "Means for saving tourist spot information in a database and updating the AI ​​model" refers to a means for saving new tourist spot information provided by users in a database and using that information to learn and update the AI ​​model.

[0477] The "means for analyzing emotions and providing information based on emotions" refers to a means for analyzing a user's facial expressions and voice to recognize emotions and providing information according to those emotions.

[0478] "Means for capturing camera images and transmitting them to a server in real time" refers to means for capturing images using a camera on a user terminal and transmitting the images to a server in real time.

[0479] "Means for calculating the display position and superimposing and displaying an AI character on the camera image" refers to means for calculating the display position of an AI character based on the acquired position information and camera image, and superimposing and displaying the AI ​​character on the camera image at that position.

[0480] MODE FOR CARRYING OUT THE INVENTION

[0481] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR displays, an emotion engine, and more to provide a tourist guide that responds to the user's emotions. Each component and its operation are described in detail below.

[0482] Location information acquisition means

[0483] Device:

[0484] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[0485] Tourist attraction information search methods

[0486] server:

[0487] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[0488] Tourist attraction information provision method

[0489] server:

[0490] Tourist attraction information from search results is sent to the device.

[0491] Device:

[0492] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[0493] Tourist guide using AI characters in local dialects

[0494] Device:

[0495] The displayed tourist attraction information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[0496] Chat-style question and answering tool

[0497] User:

[0498] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[0499] Device:

[0500] The entered question is sent to the server.

[0501] server:

[0502] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[0503] Device:

[0504] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[0505] AR display guide

[0506] User:

[0507] The user selects the AR mode on the device and activates the camera.

[0508] Device:

[0509] The camera captures the video and sends it to the server in real time.

[0510] server:

[0511] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[0512] Device:

[0513] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides tourist information.

[0514] A learning tool for user-provided information on famous places

[0515] User:

[0516] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[0517] Device:

[0518] The post content is sent to the server.

[0519] server:

[0520] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[0521] server:

[0522] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[0523] Response methods using emotion engines

[0524] Device:

[0525] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[0526] server:

[0527] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[0528] Device:

[0529] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[0530] Specific examples

[0531] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant location, and generates a guide message, such as "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device. Based on this guide message, an AI guide character on the device provides audio guidance in the local dialect. In addition, by using AR mode, an AI character appears on the camera image, providing visual guidance to the destination. Furthermore, if the emotion engine detects that the user appears tired, it will suggest nearby rest spots and relaxation areas, helping the user continue sightseeing comfortably.

[0532] As a result, the system of the present invention can provide a personalized tourist experience by addressing the lack of tourist guides, the lack of local information, and the user's emotions.

[0533] Example prompt sentence:

[0534] A user visits a local tourist destination and wears smart glasses to travel to the nearest tourist attraction from their current location. The glasses provide appropriate information based on the user's emotions, and audio guidance is provided in the local dialect. Information about tourist attractions is displayed in AR.

[0535] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0536] Program processing flow

[0537] Step 1:

[0538] Obtaining location information

[0539] (input)

[0540] The terminal uses the GPS function to obtain the user's current location information.

[0541] (Data processing)

[0542] Location information is obtained in the form of latitude and longitude and is updated periodically.

[0543] (output)

[0544] The acquired location information is temporarily stored on the device and then sent to the server.

[0545] Step 2:

[0546] Search for tourist attraction information

[0547] (input)

[0548] Based on the location information received from the terminal, the server searches a database for information on tourist attractions in the vicinity of the relevant location.

[0549] (Data calculation)

[0550] Execute a database query to retrieve a list of attractions within a specified radius.

[0551] (output)

[0552] Search results include detailed information about tourist attractions (descriptions, photos, and user reviews).

[0553] Step 3:

[0554] Providing tourist attraction information

[0555] (input)

[0556] The server transmits the retrieved tourist attraction information to the terminal.

[0557] (Data calculation)

[0558] The acquired tourist attraction information is converted into an appropriate format and sent via the API.

[0559] (output)

[0560] The terminal displays the received tourist attraction information on a user interface.

[0561] Step 4:

[0562] Guided by an AI character who speaks the local dialect

[0563] (input)

[0564] Based on tourist attraction information received from the server, the device displays an AI character that speaks the local dialect and provides audio guidance.

[0565] (Data calculation)

[0566] Using speech synthesis technology, tourist attraction information is read out in the local dialect.

[0567] (output)

[0568] The user is provided with audio tourist information.

[0569] Step 5:

[0570] Chat-style question and answering

[0571] (input)

[0572] The user types a question into the chat window on the terminal.

[0573] (Data calculation)

[0574] The entered question is sent to the server, where it is analyzed using natural language processing technology and an appropriate answer is generated.

[0575] (output)

[0576] The generated answer is sent to the device, displayed in the chat window, and read aloud by an AI character.

[0577] Step 6:

[0578] AR display guide

[0579] (input)

[0580] The user selects the AR mode on the device and activates the camera.

[0581] (Data processing)

[0582] The device captures camera footage and transmits it to the server in real time.

[0583] (output)

[0584] The server calculates the display position of the AI ​​character based on the location information and captured video, and sends the AI ​​character's 3D model data and location information to the device. The device then displays the AI ​​character overlaid on the camera image based on the received data.

[0585] Step 7:

[0586] Learning user-provided points of interest information

[0587] (input)

[0588] When users discover a new tourist spot, they post a photo and description to the app.

[0589] (Data calculation)

[0590] The posted content is sent to the server and saved in the database as new tourist spot information.

[0591] (output)

[0592] Based on this, the AI ​​model is updated and the latest tourist spot information is provided to other users from the next time onwards.

[0593] Step 8:

[0594] Emotional engine response

[0595] (input)

[0596] The device analyzes the user's facial expressions and voice and uses an emotion engine to recognize the user's emotional state.

[0597] (Data calculation)

[0598] The emotion engine analyzes emotions based on the data it acquires and adjusts the information it provides accordingly.

[0599] (output)

[0600] The server selects appropriate tourist attraction information and guidance content based on the user's emotions and transmits them to the terminal, which then provides the adjusted information to the user.

[0601] Prompt Sentence Examples

[0602] A user visits a local tourist destination and wears smart glasses to travel to the nearest tourist attraction from their current location. The glasses provide appropriate information based on the user's emotions, and audio guidance is provided in the local dialect. Information about tourist attractions is displayed in AR.

[0603] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0604] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0605] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0606] [Second embodiment]

[0607] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0608] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0609] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0610] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0611] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0612] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0613] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0614] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0615] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0616] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0617] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0618] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0619] This invention is a system that complements the lack of tourist guides in local tourist destinations and provides tourists with a new experience. This system utilizes location information and features an AI character that speaks in the local dialect to provide tourist guidance. Below, we will explain each component of the system and its operation in detail.

[0620] Location information acquisition means

[0621] Device:

[0622] When a user visits a tourist spot, the device acquires the current location information using GPS, which is then sent to the server.

[0623] Tourist attraction information search methods

[0624] server:

[0625] Based on the received location information, the server searches a database for information on nearby tourist attractions, including detailed descriptions, photos, and reviews of the attractions.

[0626] Tourist attraction information provision method

[0627] server:

[0628] The searched tourist attraction information is transmitted to the user terminal.

[0629] Device:

[0630] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is displayed.

[0631] Tourist guide using AI characters in local dialects

[0632] Device:

[0633] The tourist attraction information displayed is provided to the user by an AI character speaking in the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[0634] Chat-style question and answering tool

[0635] User:

[0636] Users can type questions into a chat window on their device (e.g., "What are some recommended restaurants nearby?").

[0637] Device:

[0638] The entered question is sent to the server.

[0639] server:

[0640] The received question is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[0641] Device:

[0642] The generated answer is displayed in the chat window, and the AI ​​character also responds verbally in the local dialect.

[0643] AR display guide

[0644] User:

[0645] The user selects the AR mode on the device and activates the camera.

[0646] Device:

[0647] The camera captures the video and sends it to the server in real time.

[0648] server:

[0649] Based on the location information and captured video, the display position of the AI ​​character is calculated and sent to the device.

[0650] Device:

[0651] Based on the received data, an AI character is superimposed on the captured image. For example, if you stand in front of a tourist spot, an AI guide character will appear in AR and begin explaining.

[0652] A learning tool for user-provided information on famous places

[0653] User:

[0654] Users post photos and descriptions of new tourist spots to the app (e.g., when they discover a new photo spot).

[0655] Device:

[0656] The post content is sent to the server.

[0657] server:

[0658] The new information is stored in a database and used as training data for the AI ​​model.

[0659] server:

[0660] The learning results will be used to provide the latest tourist spot information to other users the next time they visit the same location.

[0661] Specific examples

[0662] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server obtains information about the peak of a nearby tourist attraction called "Mount XX," and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[0663] As a result, the system of the present invention can solve problems such as a lack of tourist guides and information specific to a particular region, and provide tourists with new experiences.

[0664] The processing flow will be explained below.

[0665] Specific processing steps will be described below.

[0666] Step 1:

[0667] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[0668] Step 2:

[0669] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[0670] Step 3:

[0671] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and configured.

[0672] Step 4:

[0673] Server: Sends the data of the configured tourist guide AI character to the terminal.

[0674] Step 5:

[0675] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[0676] Step 6:

[0677] User: Visits a tourist spot and starts moving around with the device in hand.

[0678] Step 7:

[0679] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[0680] Step 8:

[0681] Server: Based on the received location information, search the database for nearby tourist attractions.

[0682] Step 9:

[0683] Server: Generates relevant tourist attraction information and sends it to the terminal.

[0684] Step 10:

[0685] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[0686] Step 11:

[0687] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[0688] Step 12:

[0689] Terminal: Sends the entered question to the server.

[0690] Step 13:

[0691] Server: Analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[0692] Step 14:

[0693] Server: Sends the generated answer to the device.

[0694] Step 15:

[0695] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[0696] Step 16:

[0697] User: Select the AR mode on the device and activate the camera.

[0698] Step 17:

[0699] Terminal: Captures camera images and sends them to the server in real time.

[0700] Step 18:

[0701] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the data to the device.

[0702] Step 19:

[0703] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI provides tourist information using AR display.

[0704] Step 20:

[0705] Users: When they discover a new tourist spot, they post a photo and description to the app.

[0706] Step 21:

[0707] Device: Sends the posted content (photos, text, location information, etc.) to the server.

[0708] Step 22:

[0709] Server: Stores new tourist spot information in a database and uses it as training data for the AI ​​model.

[0710] Step 23:

[0711] Server: Updates the AI ​​model based on the training data to improve the accuracy of the tourist guide AI.

[0712] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[0713] Example 1

[0714] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0715] In regional tourist destinations, a lack of tourist guides means that tourists cannot obtain sufficient information. Furthermore, tourist information is often provided in standard Japanese, leaving few opportunities to experience the unique culture and dialects of the region. Furthermore, there is a lack of systems that can quickly and accurately respond to the diverse questions tourists ask. Another issue is how to efficiently incorporate new tourist information provided by users and utilize it in future tourism services.

[0716] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0717] In this invention, the server includes a means for acquiring location information, a means for searching for information on tourist spots based on the acquired location information, a means for providing the user with the searched information on tourist spots, a means for displaying an AI character that speaks in the local dialect and providing tourist information, a means for accepting questions from the user, generating answers based on the questions, and providing the answer to the user, a means for displaying the AI ​​character in augmented reality based on the user's location information, and a means for storing the tourist spot information provided by the user in a database and updating the AI ​​model. This solves the problem of a lack of tourist guides and allows users to have new experiences through local information and dialects. Furthermore, the server can respond quickly and accurately to user questions and provide the latest tourist information that is continuously updated.

[0718] "Location information" is information that indicates the user's current geographical location.

[0719] "Means of acquisition" refers to the method or apparatus by which a device or system collects location information or other data.

[0720] "Tourist destination" refers to a place or attraction that tourists are expected to visit.

[0721] A "search means" refers to a method or device for finding required information from a database or information source based on specific criteria.

[0722] "Means for providing" refers to the method or device for making information or services available to users.

[0723] "Local dialect" refers to linguistic expressions and ways of speaking that are unique to a region.

[0724] "Artificial intelligence characters" refer to virtual people or characters with specific roles that are generated using AI technology.

[0725] "Tourist information" refers to activities and services that provide information and guide tourists about tourist destinations and attractions.

[0726] "Means for accepting queries" refers to a method or device for accepting and processing inquiries or questions from users.

[0727] "Answer generation means" refers to a method or device that generates an appropriate response to a user's question.

[0728] "Augmented reality display" refers to a technology that overlays digital information on the real world.

[0729] A "database" refers to a collection of information that is systematically organized and made available for efficient search and use.

[0730] An "artificial intelligence model" refers to a conceptual model of an AI system that is trained based on large amounts of data to perform specific tasks.

[0731] The present invention is a system for supplementing the lack of tourist guides in tourist destinations and providing tourists with information and experiences unique to the region. This system utilizes location information and provides tourist guidance using an AI character that speaks in the local dialect. Specific embodiments of this system are described in detail below.

[0732] Hardware and Software Configuration

[0733] Device:

[0734] The device is a user's smartphone or tablet. This device is equipped with a GPS function, a camera, a microphone, and a speaker. The device also has network connectivity (Wi-Fi or mobile data).

[0735] server:

[0736] A cloud server is used to perform location analysis, database management, natural language processing (NLP), and augmented reality (AR) calculations. The following software and services run on the server.

[0737] DBMS (Database Management System): Manages tourist spot information, user-provided information, etc.

[0738] NLP engine: Analyzes user questions and generates appropriate answers

[0739] AR engine: Calculates the display position of the AI ​​character

[0740] Program processing overview

[0741] Device:

[0742] 1. When a user visits a tourist spot, the device uses GPS to obtain current location information, which is then sent to the server in real time.

[0743] 2. The received tourist spot information is displayed on the user interface. For example, it might say, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[0744] 3. The displayed tourist attraction information is then provided to the user by an AI character speaking in the local dialect. For example, the guide will say, "If you go straight ahead, you will come to a famous local shrine."

[0745] server:

[0746] 1. The server searches the database based on the received location information and obtains information about tourist attractions in the vicinity of the location (detailed descriptions, photos, reviews, etc.).

[0747] 2. Analyze the received question using natural language processing technology and generate an appropriate answer.

[0748] 3. Based on the received camera footage and location information, the display position of the AI ​​character is calculated and sent to the device.

[0749] Specific examples

[0750] As a concrete example, consider the case where a user visits a mountain called "Mount XX" in a local tourist spot. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server retrieves information about the peak of nearby tourist spot "Mount XX" from a database and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[0751] Prompt Sentence Examples

[0752] The following prompt sentences could be input to the generative AI model:

[0753] Example: "Please explain in detail the system that displays an AI character that guides users in the local dialect when they visit a local tourist spot."

[0754] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[0755] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0756] Step 1:

[0757] Terminal: When a user visits a tourist spot, the terminal uses a GPS module to obtain current location information. The input is the user's physical location information, and the output is GPS coordinate data (latitude and longitude). Specifically, the GPS module determines the location of the spot, and the data is processed internally on the terminal to generate coordinate information. This coordinate information is then sent to the server.

[0758] Step 2:

[0759] Terminal: Sends the acquired location information to the server. The input is the GPS coordinate data acquired in step 1, and the output is the location data sent to the server. Specifically, the terminal uses a network communication module (e.g., Wi-Fi or mobile data communication) to send the location data to the specified endpoint on the server.

[0760] Step 3:

[0761] Server: Based on the received location information, the server searches a database. The input is location data, and the output is information about tourist attractions near the location. Specifically, the server performs a database query to obtain detailed information (e.g., name, description, photos, reviews, etc.) about the nearest tourist attraction based on the location information.

[0762] Step 4:

[0763] Server: Sends the searched tourist attraction information to the user's device. The input is the tourist attraction information retrieved from the database, and the output is the tourist attraction information data sent to the device. Specifically, the server packages the tourist attraction information in an appropriate data format (e.g., JSON format) and sends it to the user's device via the network.

[0764] Step 5:

[0765] Terminal: The terminal displays the information it receives on its user interface. The input is tourist attraction information data sent from the server, and the output is tourist attraction information that is visually displayed to the user. Specifically, the terminal parses the received JSON data and displays the information in an appropriate format on the application's user interface. For example, it might display, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[0766] Step 6:

[0767] Terminal: An AI character speaking in the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is the visually displayed tourist attraction information, and the output is audio guidance information. Specifically, the terminal inputs text data into a speech synthesis engine (e.g., a TTS engine), which generates and plays audio data in the local dialect.

[0768] Step 7:

[0769] User: The user types a question into the chat window of the terminal. The input is the user's text-based question, and the output is the question data stored in the terminal. In concrete terms, the user types a question into the chat window of the application and presses the send button.

[0770] Step 8:

[0771] Terminal: Sends the entered question to the server. The input is the question data from the user, and the output is the question data sent to the server. In concrete terms, the terminal sends the question data to the server via the network communication module.

[0772] Step 9:

[0773] Server: Analyzes the received question using natural language processing technology and generates an appropriate answer. The input is the user's question data, and the output is the generated answer data. Specifically, the server uses a natural language processing engine to analyze the question and generate the optimal answer from a database or knowledge base.

[0774] Step 10:

[0775] Server: Sends the generated answer to the terminal. The input is the answer data, and the output is the answer data sent to the terminal. Specifically, the server packages the answer data in an appropriate data format and sends it to the user's terminal via the network.

[0776] Step 11:

[0777] Terminal: The terminal displays the generated answer in the chat window, and the AI ​​character also answers aloud in the local dialect. The input is the answer data sent from the server, and the output is the answer information provided visually and aloud. Specifically, the terminal analyzes the received answer data and displays it as text in the chat window. It also uses a speech synthesis engine to generate and play back audio data in the local dialect.

[0778] Step 12:

[0779] User: The user selects the AR mode on the device and activates the camera. The input is the user's operation, and the output is the video data captured by the camera. Specifically, the user presses the AR mode button in the application to activate the device's camera.

[0780] Step 13:

[0781] Terminal: Captures camera images and sends them to the server in real time. The input is the camera's image data, and the output is the image data sent to the server. Specifically, the terminal acquires image data from the camera module, compresses it, and sends it to the server.

[0782] Step 14:

[0783] Server: Based on location information and camera footage, calculates the display position of the AI ​​character and sends this to the device. The input is location information and video data, and the output is character display position data. Specifically, the server uses the AR engine to analyze the local geographic information and video data and calculate the optimal character display position.

[0784] Step 15:

[0785] Terminal: Based on the received data, an AI character is superimposed on the camera image. The input is character display position data and camera image data, and the output is the character displayed in AR. In concrete terms, the terminal superimposes a virtual character on the camera image based on the received position data. For example, when you stand in front of a tourist attraction, an AI guide character will appear in AR and begin to explain.

[0786] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[0787] (Application example 1)

[0788] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0789] Currently, there is a significant shortage of tourist guides and store guides in regional tourist destinations and shopping malls. Furthermore, tourists and visitors have limited access to information specific to their region, raising concerns about declining satisfaction. Furthermore, information about tourist destinations and stores is often provided in standard Japanese, which means that the characteristics and atmosphere of the region cannot be fully conveyed. This makes it difficult to maximize the appeal of tourist destinations and shopping malls and provide visitors with personalized, attractive guides. Therefore, there is a need to develop a system that can efficiently provide tourist guides and store guides and offer new experiences to visitors.

[0790] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0791] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving tourist attraction information provided by the user in a database and updating the AI ​​model, means for searching for store information within a shopping mall from the database, means for using an AI character that provides guidance in the local dialect and displaying it on a smartphone camera image, means for saving new store information and photos posted by the user in the database, and means for users to input questions in a chat format and for analyzing and answering them using natural language processing technology. This enables tourist spots and shopping malls to efficiently provide visitors with personalized guides and information using the local dialect.

[0792] "Location information acquisition means" refers to a system in which a user's terminal acquires the current location using GPS or other technology.

[0793] The "tourist attraction information search means" is a function that allows the server to search the database for information on the relevant tourist attraction based on the acquired location information.

[0794] The "tourist attraction information providing means" is a function that transmits information about the searched tourist attractions to the user terminal and displays it on the user interface.

[0795] "AI characters that speak in local dialects" refers to artificial intelligence characters that provide audio guidance using a dialect specific to the region.

[0796] The "question answering means" is a function that accepts questions from users, analyzes them, and generates and provides appropriate answers.

[0797] "Augmented reality display means" is a technology that displays virtual information superimposed on real-world images, and in this case refers to the function of displaying an AI character superimposed on a real-world scene.

[0798] The "user-provided information storage means" is a function that stores information and photos of tourist spots provided by users in a database.

[0799] "AI model updating means" refers to a method for retraining an artificial intelligence model based on stored user-provided information to improve its accuracy.

[0800] The "shopping mall information search means" is a function for searching for information on nearby stores based on the current location within the shopping mall.

[0801] "Natural language processing technology" is a technology that allows a computer to analyze a user's question, understand human language, and generate an appropriate answer.

[0802] This system complements the lack of tourist guides and store guides in shopping malls and local tourist destinations, providing visitors with a personalized new experience. This system uses an AI character that provides audio guidance in the local dialect and displays the information on the smartphone camera using augmented reality (AR). It also includes a function that uses natural language processing (NLP) to generate appropriate answers to user questions.

[0803] First, when a user visits a tourist spot or shopping mall, the location information acquisition means uses the smartphone's GPS to acquire current location information. This location information is sent to the server. Based on the received location information, the server searches a database for information on tourist attractions and stores in the vicinity of the location. For example, it provides information such as, "There is a recommended restaurant 200 meters from your current location."

[0804] Next, the tourist attraction information provider sends the searched tourist attraction and store information to the user's smartphone. The displayed information is guided by an AI character speaking the local dialect. For example, guidance such as "If you turn right at the next street, you will find a delicious Japanese restaurant" is provided.

[0805] The system also includes a chat-style question-and-answer mechanism. Users enter questions into a chat window on their smartphone, which is then sent to the server. The server uses natural language processing technology to analyze the question and generate an appropriate answer. For example, in response to the question, "What are some recommended cafes nearby?", the server will respond with, "The recommended nearby cafe is on the second floor of the mall."

[0806] Furthermore, as an augmented reality display method, when a user activates the smartphone camera and selects AR mode, an AI character is superimposed on the camera image, allowing the user to receive guidance while viewing the AI ​​character over the real-world scene.

[0807] The user-provided information storage means provides a mechanism that allows users to post new tourist spots and store information to the app. The posted information is sent to the server and stored in a database. This information is used as learning data for the AI ​​model, and the next time other users visit the same location, more up-to-date information will be provided.

[0808] As a concrete example, consider a smartphone app for tourists visiting a local shopping mall. As a user walks through the mall, their current location is acquired via GPS and sent to a server. The server uses this location information to search for nearby store information and provides information such as, "There's a new cafe 200 meters from your current location." When a user asks in a chat window, "What are some recommended restaurants nearby?", an answer analyzed using NLP technology is displayed. In addition, when using AR mode, an AI character appears on the camera image, providing specific directions.

[0809] The tour guide using the generative AI model generates a detailed scenario using the following prompt sentence:

[0810] ---

[0811] When a user visits a local shopping mall, provide an application that uses the smartphone's location information to provide voice guidance on recommended stores and tourist spots in the area. The application should include a feature that uses an AI character that speaks the local dialect, overlaying the AI ​​character on the screen in AR, and provide specific guidance. It should also have the ability to search a database based on GPS coordinates, analyze user questions using NLP technology, and allow users to post new store information. Please explain with specific examples.

[0812] ---

[0813] In this way, the present invention aims to make up for the lack of guides at tourist spots and shopping malls and provide visitors with a new experience.

[0814] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0815] Step 1:

[0816] The server receives GPS location information sent from the user's device. The input is the coordinate data of the current location obtained from the smartphone's GPS sensor. Based on this data, the server searches a database for information on relevant tourist attractions and stores in the shopping mall. As output, it generates information on the tourist attractions and stores found.

[0817] Step 2:

[0818] The server sends the searched tourist attraction and store information to the user terminal. The tourist attraction and store information generated in step 1 is the input. The server formats this and sends it to the user terminal in an appropriate format. The output is information data to be displayed on the user terminal.

[0819] Step 3:

[0820] The device uses an AI character to provide audio guidance in the local dialect based on the received information. The inputs include tourist attraction and store information sent from the server and the AI ​​character's audio data. The device combines this information to output audio and display on the screen. The output is audio guidance and visual guidance information presented to the user.

[0821] Step 4:

[0822] The user enters a question into the chat window on the smartphone. The input is a text question. The device sends this to the server. The output is the user question data.

[0823] Step 5:

[0824] The server analyzes the received question using natural language processing (NLP) technology and generates an appropriate answer. The input is the question data entered by the user. The server inputs this data into an NLP model and generates an answer based on the analysis results. The output is an appropriate answer text.

[0825] Step 6:

[0826] The server sends the generated answer to the user terminal. The input is the answer text generated in step 5. The server formats it and sends it to the user terminal. The output is the answer data to be displayed on the user terminal.

[0827] Step 7:

[0828] The device displays the received response in the user's chat window and responds audibly if necessary. The inputs are the response data sent from the server and the AI ​​character's voice data. The device combines these and outputs text and voice. The output is a visual and voice response to the user.

[0829] Step 8:

[0830] The user activates the smartphone camera and selects AR mode. The input is the camera image. The device captures it and sends it to the server. The output is real-time video data sent to the server.

[0831] Step 9:

[0832] The server calculates the display position of the AI ​​character based on the received camera image and location information, and sends this data to the device. The inputs are camera image data and GPS location information. The server processes this data and generates display position data for the AI ​​character. The output is sent to the device.

[0833] Step 10:

[0834] The device overlays the AI ​​character on the camera image based on the received AI character position data. The inputs are camera image data and AI character position data. The device combines these and displays them. The output is an augmented reality image from the user's perspective.

[0835] Step 11:

[0836] Users post information about new tourist spots and shops to the app. The input is text and image data. The device sends this to the server. The output is the posted data sent to the server.

[0837] Step 12:

[0838] The server stores the received posting data in a database and uses it as training data for the AI ​​model. The input includes tourist spot and store information data posted by users. The server stores this in the database and retrains the AI ​​model to improve its accuracy. As output, the new information data is stored in the database and an updated AI model is generated.

[0839] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0840] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR display, and an emotion engine to realize a tourist guide that responds to the user's emotions. Each component and its operation are explained in detail below.

[0841] Location information acquisition means

[0842] Device:

[0843] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[0844] Tourist attraction information search methods

[0845] server:

[0846] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[0847] Tourist attraction information provision method

[0848] server:

[0849] Tourist attraction information from search results is sent to the device.

[0850] Device:

[0851] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[0852] Tourist guide using AI characters in local dialects

[0853] Device:

[0854] The displayed tourist spot information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[0855] Chat-style question and answering tool

[0856] User:

[0857] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[0858] Device:

[0859] The entered question is sent to the server.

[0860] server:

[0861] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[0862] Device:

[0863] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[0864] AR display guide

[0865] User:

[0866] The user selects the AR mode on the device and activates the camera.

[0867] Device:

[0868] The camera captures the video and sends it to the server in real time.

[0869] server:

[0870] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[0871] Device:

[0872] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides sightseeing information.

[0873] A learning tool for user-provided information on famous places

[0874] User:

[0875] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[0876] Device:

[0877] The post content is sent to the server.

[0878] server:

[0879] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[0880] server:

[0881] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[0882] Response methods using emotion engines

[0883] Device:

[0884] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[0885] server:

[0886] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[0887] Device:

[0888] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[0889] Specific examples

[0890] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[0891] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In addition, by using AR mode, an AI character will appear on the camera image, providing visual guidance to the destination.

[0892] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas to help the user continue sightseeing in comfort.

[0893] As a result, the system of the present invention can provide a personalized tourist experience by addressing the lack of tourist guides and information specific to the region, as well as responding appropriately to the user's emotions.

[0894] The processing flow will be explained below.

[0895] Specific processing steps will be described below.

[0896] Step 1:

[0897] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[0898] Step 2:

[0899] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[0900] Step 3:

[0901] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and set up.

[0902] Step 4:

[0903] Server: Sends the data of the configured tourist guide AI character to the terminal.

[0904] Step 5:

[0905] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[0906] Step 6:

[0907] User: Visits a tourist spot and starts moving around with the device in hand.

[0908] Step 7:

[0909] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[0910] Step 8:

[0911] Server: Based on the received location information, search the database for nearby tourist attractions.

[0912] Step 9:

[0913] Server: Generates relevant tourist attraction information and sends it to the terminal.

[0914] Step 10:

[0915] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[0916] Step 11:

[0917] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[0918] Step 12:

[0919] Terminal: Sends the entered question to the server.

[0920] Step 13:

[0921] Server: Analyzes the question using natural language processing technology and generates an appropriate answer.

[0922] Step 14:

[0923] Server: Sends the generated answer to the device.

[0924] Step 15:

[0925] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[0926] Step 16:

[0927] User: Select the AR mode on the device and activate the camera.

[0928] Step 17:

[0929] Terminal: Captures camera images and sends them to the server in real time.

[0930] Step 18:

[0931] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the 3D model data and location information to the device.

[0932] Step 19:

[0933] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI character provides sightseeing information.

[0934] Step 20:

[0935] Users: When they discover a new tourist spot, they post a photo and description to the app.

[0936] Step 21:

[0937] Terminal: Sends the posted content to the server.

[0938] Step 22:

[0939] Server: Saves new tourist spot information in a database and uses it as training data for the AI ​​model.

[0940] Step 23:

[0941] Server: Updates the AI ​​model based on the learning results and provides the latest tourist spot information to other users from the next time onwards.

[0942] Step 24:

[0943] Terminal: Analyzes the user's facial expressions and voice, and recognizes their emotional state using an emotion engine.

[0944] Step 25:

[0945] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine.

[0946] Step 26:

[0947] Terminal: Providing information according to emotions.

[0948] Specific examples

[0949] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of a nearby tourist spot called "Mount XX," and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters away from here," and sends it to the device.

[0950] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In addition, by using AR mode, an AI character will appear on the camera image, providing visual guidance to the destination.

[0951] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas to help the user continue sightseeing in comfort.

[0952] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[0953] Example 2

[0954] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0955] In traditional tourist destinations, the number of guides is limited, making it difficult to provide individualized and satisfying tours to all visitors. There is also a lack of information that responds to visitors' emotions and personal preferences. Furthermore, there is a lack of technological means to provide advanced guidance and guidance at tourist destinations. This prevents visitors from fully enjoying the destinations, hindering the revitalization of the tourism industry as a whole.

[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0957] In this invention, the server includes a means for acquiring location information, a means for searching for tourist attraction information, and a means for providing tourist attraction information to users, thereby making it possible to provide personalized tourist information to visitors and increase their satisfaction.

[0958] "Location information" is data that indicates a user's current geographic location.

[0959] "Tourist attraction" refers to a place or facility that is of interest to visitors for tourism purposes.

[0960] "Searching for information" is the act of locating relevant information from databases or sources based on specific criteria.

[0961] "Providing" is the act of giving or displaying information or services to a user.

[0962] An "artificial intelligence character" is a virtual guide character programmed using AI technology.

[0963] "Augmented reality display" is a technology that displays digital information and objects overlaid on real-world scenery.

[0964] A "database" is a collection of information for efficiently storing, searching, and managing large amounts of data.

[0965] An "artificial intelligence model" is an implementation of an AI algorithm that learns and makes inferences based on data.

[0966] "Emotional state" is information that indicates the psychological and physiological state of the user.

[0967] This invention is a system that compensates for the lack of guides in local tourist destinations and provides personalized tourist experiences. This system incorporates GPS functionality, natural language processing (NLP) technology, AR technology, and an emotion recognition engine. Below, we will explain in detail each component that makes up the system and its operation.

[0968] Location information acquisition means

[0969] Device: When a user visits a tourist spot, the device uses its GPS function to obtain current location information. This location information is periodically sent to the server. For example, the GPS module of a smartphone can be used to obtain the user's location and send it to the server as an HTTP request.

[0970] Tourist attraction information search methods

[0971] Server: Based on the location information received from the device, the server searches a database for information on nearby tourist attractions. This database contains detailed descriptions of tourist attractions, photos, user reviews, etc. The server executes SQL queries to search for this information.

[0972] Tourist attraction information provision method

[0973] Server: Sends tourist attraction information obtained as search results to the terminal.

[0974] Terminal: The terminal displays the information received from the server on a user interface. For example, it may display information such as, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[0975] Tourist guide using AI characters in local dialects

[0976] Device: The displayed tourist attraction information is provided to the user by an AI character that speaks the local dialect. For example, the device uses a speech synthesis engine to generate voice guidance in the local dialect, such as "If you go straight ahead, you will find a famous local shrine."

[0977] Chat-style question and answering tool

[0978] User: If the user has a question, they type it into the chat window on their device. For example, "What are some recommended restaurants nearby?"

[0979] Terminal: The terminal sends the entered question to the server.

[0980] Server: The server analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[0981] Device: The device will display the answers received from the server in a chat window, and the guide AI character will read the answers aloud in the local dialect.

[0982] AR display guide

[0983] User: The user selects the AR mode on the device and activates the camera.

[0984] Terminal: The terminal captures the camera image and transmits it to the server in real time.

[0985] Server: The server calculates the display position of the AI ​​character based on the location information and captured video, and sends the AI ​​character's 3D model data and location information to the device.

[0986] Terminal: Based on the received data, the terminal displays an AI character superimposed on the camera image. The guide AI character provides sightseeing information.

[0987] A learning tool for user-provided information on famous places

[0988] User: When a user finds a new tourist spot, they post a photo and description of it to the app. For example, they post something like, "I found a new photo spot."

[0989] Terminal: The terminal sends the posted content to the server.

[0990] Server: The server saves new tourist spot information in a database and uses it as learning data for the AI ​​model. It then updates the AI ​​model based on the learning results, enabling it to provide the latest tourist spot information to other users from the next time onwards.

[0991] Response methods using emotion engines

[0992] Device: The device analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[0993] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine.

[0994] Device: The device provides information based on the user's emotions. For example, if the device detects that the user is tired, it will suggest nearby rest spots or cafes.

[0995] Specific examples

[0996] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[0997] Based on this information, the device will have an AI guide character provide voice guidance in the local dialect. For example, it might say, "500 meters ahead there is a peak with a spectacular view." In addition, when using AR mode, the AI ​​character will appear on the camera image, providing visual guidance to the destination.

[0998] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas. This allows the system of the present invention to address the lack of tourist guides and local information, and respond appropriately to the user's emotions, providing a personalized sightseeing experience.

[0999] Prompt Sentence Examples

[1000] The following are some examples of input prompts for a generative AI model:

[1001] "Please introduce us to a new tourist guide system. This system will address the lack of guides in local tourist destinations and provide a personalized tourist experience. Specifically, it will use location information acquisition, AI characters that speak the local dialect, AR displays, and an emotion engine to provide a tourist guide that responds to the user's emotions."

[1002] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1003] Step 1: Obtaining location information

[1004] Terminal: When a user visits a tourist spot, the terminal's GPS function is activated and acquires current location information. The input is geographic coordinate data from the GPS sensor, and the output is current location information. This location information is sent to the server at regular intervals.

[1005] Specific operation: The device uses the built-in GPS module to obtain geographic coordinates (latitude and longitude) and transmits them to the server as location information.

[1006] Step 2: Search for tourist attraction information

[1007] Server: Receives the acquired location information and searches for information on nearby tourist attractions. The input is the location information and the output is a list of nearby tourist attractions. The server searches the database using an SQL query and obtains the search results.

[1008] Specific operation: Based on the received location information, the server performs a range search on the tourist attraction database to obtain detailed information about related tourist attractions.

[1009] Step 3: Providing tourist attraction information

[1010] Server: Sends tourist attraction information obtained as search results to the terminal. The input is tourist attraction information, and the output is the transmission of information to the terminal.

[1011] Specific operation: The server composes the search results in JSON format and sends them to the terminal as an HTTP response.

[1012] Step 4: View tourist attraction information

[1013] Terminal: The terminal displays tourist attraction information received from the server on a user interface. The input is tourist attraction information from the server, and the output is the information displayed on the display.

[1014] Specific operation: The terminal parses the received JSON data, binds the information to the user interface, and displays it.

[1015] Step 5: Tourist guide in local dialect

[1016] Terminal: An AI character speaking the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is text information about the tourist attraction, and the output is synthesized speech.

[1017] Specific operation: The device passes the text information to a TTS (Text-to-Speech) engine, synthesizes it in the local dialect, and plays it back through the speaker.

[1018] Step 6: Chat-style Q&A

[1019] User: The user types a question into the chat window on the device. The input is the text of the user's question.

[1020] Terminal: The terminal sends the input question to the server. The output is the question data sent to the server.

[1021] Specific behavior: The user types a question into the chat interface, the device captures it, and sends it to the server via an HTTP POST request.

[1022] Server: The server analyzes the question content using NLP technology and generates an appropriate answer. The input is the user's question text, and the output is the generated answer text.

[1023] What it does: The server uses an NLP model to analyze the question and generate an appropriate answer from a database or other source.

[1024] Terminal: The terminal displays the answer received from the server in a chat window and has an AI character read it aloud in the local dialect. The input is the answer text from the server, and the output is the displayed answer and audio.

[1025] Specific operation: The device displays the received reply text in the chat window, synthesizes it using the TTS engine, and plays it back.

[1026] Step 7: Guided by AR display

[1027] User: The user selects the device's AR mode and activates the camera. The input is the AR mode selection action.

[1028] Terminal: The terminal captures camera images and transmits them to the server in real time. The input is the camera image and the output is the video data sent to the server.

[1029] Specific operation: The device captures camera footage and streams it to the server along with location information.

[1030] Server: The server calculates the display position of the AI ​​character based on the location information and video data, and sends the 3D model data to the terminal. The input is the camera image and location information, and the output is the display position of the AI ​​character and 3D model data.

[1031] Specific operation: The server uses the AR calculation engine to calculate the appropriate display position and send the 3D model data to the terminal.

[1032] Terminal: Based on the received data, the terminal displays an AI character superimposed on the camera image. The input is 3D model data and display position information, and the output is an AI character superimposed on the camera image.

[1033] What it does: The device uses an AR library to render a 3D model overlaid on the camera image.

[1034] Step 8: Learning user-provided points of interest information

[1035] User: When a user finds a new tourist spot, he / she posts a photo and description of the spot to the app. The input is a photo of the new tourist spot and text information.

[1036] Terminal: The terminal sends the posted content to the server. The output is the user posted data sent to the server.

[1037] Specific operation: The device collects posting data including information entered by the user and sends it to the server as an HTTP POST request.

[1038] Server: The server stores the received information in a database and uses it as training data for the AI ​​model. The input is new tourist attraction information, and the output is an updated database and AI model.

[1039] Specific operation: The server stores the received post data in a database and reflects it in the AI ​​model in the next learning phase.

[1040] Step 9: Respond with the Emotion Engine

[1041] Terminal: The terminal analyzes the user's facial expressions and voice and recognizes their emotional state using an emotion engine. The input is the user's facial expression data and voice data, and the output is the recognized emotional information.

[1042] How it works: The device uses a camera and microphone to capture the user's facial expressions and voice, and applies emotion analysis algorithms to identify their emotional state.

[1043] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine. The input is the recognized emotion information, and the output is adjusted tourist attraction information.

[1044] Specific operation: The server dynamically adjusts the method and content of providing tourist attraction information based on the emotion data.

[1045] Terminal: The terminal provides tailored information to the user. For example, if the terminal recognizes that the user is tired, it suggests nearby rest spots or cafes. The input is tailored tourist attraction information, and the output is the provided information and guidance.

[1046] Specific operation: Based on the adjusted information, the device updates the user interface and voice guidance to provide the user with the most appropriate information.

[1047] (Application example 2)

[1048] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1049] There is a shortage of tourist guides in local tourist destinations, and it is difficult to provide tourists with a personalized guided experience. Furthermore, there is a demand for dynamic information provision based on the user's emotions and location information. Furthermore, tourist guides using local dialects and visual guidance using augmented reality displays are insufficient. A means is needed to address these issues and provide a richer tourist experience.

[1050] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1051] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving the tourist attraction information provided by the user in a database and updating the AI ​​model, means for analyzing emotions from the user's facial expressions and voice and providing information based on the emotions, means for capturing camera footage and transmitting it to the server in real time, and means for calculating the display position of the AI ​​character based on the location information and the captured footage and superimposing the AI ​​character on the camera footage. This makes it possible to address the lack of tourist destination guides and local information, and to provide personalized tour guidance based on the user's emotions, as well as a visual guide experience.

[1052] Key Word Definitions

[1053] The "means for acquiring location information" is a means for acquiring the user's current location information using the GPS function and periodically transmitting it to the server.

[1054] The "means for searching tourist attraction information" is a means for searching a database for information on tourist attractions that exist in the vicinity of a point based on the acquired location information.

[1055] The "means for providing tourist attraction information to the user" refers to a means for displaying information about the searched tourist attraction to the user through a user interface.

[1056] "A means for displaying an AI character that speaks in the local dialect and provides tourist guidance" is a means for an AI character displayed to the user to provide tourist guidance using the local dialect.

[1057] The "means for generating an answer according to the question content and providing it to the user" is a means for analyzing the question entered by the user, generating an appropriate answer, and providing it to the user in a chat window or by voice.

[1058] "Means for displaying AI characters in augmented reality" refers to a means for displaying a 3D model of an AI character overlaid on camera images based on location information and captured images.

[1059] "Means for saving tourist spot information in a database and updating the AI ​​model" refers to a means for saving new tourist spot information provided by users in a database and using that information to learn and update the AI ​​model.

[1060] The "means for analyzing emotions and providing information based on emotions" refers to a means for analyzing a user's facial expressions and voice to recognize emotions and providing information according to those emotions.

[1061] "Means for capturing camera images and transmitting them to a server in real time" refers to means for capturing images using a camera on a user terminal and transmitting the images to a server in real time.

[1062] "Means for calculating the display position and superimposing and displaying an AI character on the camera image" refers to means for calculating the display position of an AI character based on the acquired position information and camera image, and superimposing and displaying the AI ​​character on the camera image at that position.

[1063] MODE FOR CARRYING OUT THE INVENTION

[1064] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR displays, an emotion engine, and more to provide a tourist guide that responds to the user's emotions. Each component and its operation are described in detail below.

[1065] Location information acquisition means

[1066] Device:

[1067] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[1068] Tourist attraction information search methods

[1069] server:

[1070] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[1071] Tourist attraction information provision method

[1072] server:

[1073] Tourist attraction information from search results is sent to the device.

[1074] Device:

[1075] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[1076] Tourist guide using AI characters in local dialects

[1077] Device:

[1078] The displayed tourist attraction information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[1079] Chat-style question and answering tool

[1080] User:

[1081] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[1082] Device:

[1083] The entered question is sent to the server.

[1084] server:

[1085] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[1086] Device:

[1087] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[1088] AR display guide

[1089] User:

[1090] The user selects the AR mode on the device and activates the camera.

[1091] Device:

[1092] The camera captures the video and sends it to the server in real time.

[1093] server:

[1094] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[1095] Device:

[1096] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides tourist information.

[1097] A learning tool for user-provided information on famous places

[1098] User:

[1099] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[1100] Device:

[1101] The post content is sent to the server.

[1102] server:

[1103] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[1104] server:

[1105] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[1106] Response methods using emotion engines

[1107] Device:

[1108] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[1109] server:

[1110] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[1111] Device:

[1112] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[1113] Specific examples

[1114] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant location, and generates a guide message, such as "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device. Based on this guide message, an AI guide character on the device provides audio guidance in the local dialect. In addition, by using AR mode, an AI character appears on the camera image, providing visual guidance to the destination. Furthermore, if the emotion engine detects that the user appears tired, it will suggest nearby rest spots and relaxation areas, helping the user continue sightseeing comfortably.

[1115] As a result, the system of the present invention can provide a personalized tourist experience by addressing the lack of tourist guides, the lack of local information, and the user's emotions.

[1116] Example prompt sentence:

[1117] A user visits a local tourist destination and wears smart glasses to travel to the nearest tourist attraction from their current location. The glasses provide appropriate information based on the user's emotions, and audio guidance is provided in the local dialect. Information about tourist attractions is displayed in AR.

[1118] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1119] Program processing flow

[1120] Step 1:

[1121] Obtaining location information

[1122] (input)

[1123] The terminal uses the GPS function to obtain the user's current location information.

[1124] (Data processing)

[1125] Location information is obtained in the form of latitude and longitude and is updated periodically.

[1126] (output)

[1127] The acquired location information is temporarily stored on the device and then sent to the server.

[1128] Step 2:

[1129] Search for tourist attraction information

[1130] (input)

[1131] Based on the location information received from the terminal, the server searches a database for information on tourist attractions in the vicinity of the relevant location.

[1132] (Data calculation)

[1133] Execute a database query to retrieve a list of attractions within a specified radius.

[1134] (output)

[1135] Search results include detailed information about tourist attractions (descriptions, photos, and user reviews).

[1136] Step 3:

[1137] Providing tourist attraction information

[1138] (input)

[1139] The server transmits the retrieved tourist attraction information to the terminal.

[1140] (Data calculation)

[1141] The acquired tourist attraction information is converted into an appropriate format and sent via the API.

[1142] (output)

[1143] The terminal displays the received tourist attraction information on a user interface.

[1144] Step 4:

[1145] Guided by an AI character who speaks the local dialect

[1146] (input)

[1147] Based on tourist attraction information received from the server, the device displays an AI character that speaks the local dialect and provides audio guidance.

[1148] (Data calculation)

[1149] Using speech synthesis technology, tourist attraction information is read out in the local dialect.

[1150] (output)

[1151] The user is provided with audio tourist information.

[1152] Step 5:

[1153] Chat-style question and answering

[1154] (input)

[1155] The user types a question into the chat window on the terminal.

[1156] (Data calculation)

[1157] The entered question is sent to the server, where it is analyzed using natural language processing technology and an appropriate answer is generated.

[1158] (output)

[1159] The generated answer is sent to the device, displayed in the chat window, and read aloud by an AI character.

[1160] Step 6:

[1161] AR display guide

[1162] (input)

[1163] The user selects the AR mode on the device and activates the camera.

[1164] (Data processing)

[1165] The device captures camera footage and transmits it to the server in real time.

[1166] (output)

[1167] The server calculates the display position of the AI ​​character based on the location information and captured video, and sends the AI ​​character's 3D model data and location information to the device. The device then displays the AI ​​character overlaid on the camera image based on the received data.

[1168] Step 7:

[1169] Learning user-provided points of interest information

[1170] (input)

[1171] When users discover a new tourist spot, they post a photo and description to the app.

[1172] (Data calculation)

[1173] The posted content is sent to the server and saved in the database as new tourist spot information.

[1174] (output)

[1175] Based on this, the AI ​​model is updated and the latest tourist spot information is provided to other users from the next time onwards.

[1176] Step 8:

[1177] Emotional engine response

[1178] (input)

[1179] The device analyzes the user's facial expressions and voice and uses an emotion engine to recognize the user's emotional state.

[1180] (Data calculation)

[1181] The emotion engine analyzes emotions based on the data it acquires and adjusts the information it provides accordingly.

[1182] (output)

[1183] The server selects appropriate tourist attraction information and guidance content based on the user's emotions and transmits them to the terminal, which then provides the adjusted information to the user.

[1184] Prompt Sentence Examples

[1185] A user visits a local tourist destination and wears smart glasses to travel to the nearest tourist attraction from their current location. The glasses provide appropriate information based on the user's emotions, and audio guidance is provided in the local dialect. Information about tourist attractions is displayed in AR.

[1186] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1187] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1188] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1189] [Third embodiment]

[1190] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1191] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1192] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1193] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1194] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1195] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1196] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1197] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1198] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1199] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1200] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1201] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1202] This invention is a system that complements the lack of tourist guides in local tourist destinations and provides tourists with a new experience. This system utilizes location information and features an AI character that speaks in the local dialect to provide tourist guidance. Below, we will explain each component of the system and its operation in detail.

[1203] Location information acquisition means

[1204] Device:

[1205] When a user visits a tourist spot, the device acquires the current location information using GPS, which is then sent to the server.

[1206] Tourist attraction information search methods

[1207] server:

[1208] Based on the received location information, the server searches a database for information on nearby tourist attractions, including detailed descriptions, photos, and reviews of the attractions.

[1209] Tourist attraction information provision method

[1210] server:

[1211] The searched tourist attraction information is transmitted to the user terminal.

[1212] Device:

[1213] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is displayed.

[1214] Tourist guide using AI characters in local dialects

[1215] Device:

[1216] The tourist attraction information displayed is provided to the user by an AI character speaking in the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[1217] Chat-style question and answering tool

[1218] User:

[1219] Users can type questions into a chat window on their device (e.g., "What are some recommended restaurants nearby?").

[1220] Device:

[1221] The entered question is sent to the server.

[1222] server:

[1223] The received question is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[1224] Device:

[1225] The generated answer is displayed in the chat window, and the AI ​​character also responds verbally in the local dialect.

[1226] AR display guide

[1227] User:

[1228] The user selects the AR mode on the device and activates the camera.

[1229] Device:

[1230] The camera captures the video and sends it to the server in real time.

[1231] server:

[1232] Based on the location information and captured video, the display position of the AI ​​character is calculated and sent to the device.

[1233] Device:

[1234] Based on the received data, an AI character is superimposed on the captured image. For example, if you stand in front of a tourist spot, an AI guide character will appear in AR and begin explaining.

[1235] A learning tool for user-provided information on famous places

[1236] User:

[1237] Users post photos and descriptions of new tourist spots to the app (e.g., when they discover a new photo spot).

[1238] Device:

[1239] The post content is sent to the server.

[1240] server:

[1241] The new information is stored in a database and used as training data for the AI ​​model.

[1242] server:

[1243] The learning results will be used to provide the latest tourist spot information to other users the next time they visit the same location.

[1244] Specific examples

[1245] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server obtains information about the peak of a nearby tourist attraction called "Mount XX," and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[1246] As a result, the system of the present invention can solve problems such as a lack of tourist guides and information specific to a particular region, and provide tourists with new experiences.

[1247] The processing flow will be explained below.

[1248] Specific processing steps will be described below.

[1249] Step 1:

[1250] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[1251] Step 2:

[1252] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[1253] Step 3:

[1254] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and configured.

[1255] Step 4:

[1256] Server: Sends the data of the configured tourist guide AI character to the terminal.

[1257] Step 5:

[1258] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[1259] Step 6:

[1260] User: Visits a tourist spot and starts moving around with the device in hand.

[1261] Step 7:

[1262] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[1263] Step 8:

[1264] Server: Based on the received location information, search the database for nearby tourist attractions.

[1265] Step 9:

[1266] Server: Generates relevant tourist attraction information and sends it to the terminal.

[1267] Step 10:

[1268] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[1269] Step 11:

[1270] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[1271] Step 12:

[1272] Terminal: Sends the entered question to the server.

[1273] Step 13:

[1274] Server: Analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[1275] Step 14:

[1276] Server: Sends the generated answer to the device.

[1277] Step 15:

[1278] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[1279] Step 16:

[1280] User: Select the AR mode on the device and activate the camera.

[1281] Step 17:

[1282] Terminal: Captures camera images and sends them to the server in real time.

[1283] Step 18:

[1284] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the data to the device.

[1285] Step 19:

[1286] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI provides tourist information using AR display.

[1287] Step 20:

[1288] Users: When they discover a new tourist spot, they post a photo and description to the app.

[1289] Step 21:

[1290] Device: Sends the posted content (photos, text, location information, etc.) to the server.

[1291] Step 22:

[1292] Server: Stores new tourist spot information in a database and uses it as training data for the AI ​​model.

[1293] Step 23:

[1294] Server: Updates the AI ​​model based on the training data to improve the accuracy of the tourist guide AI.

[1295] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[1296] Example 1

[1297] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1298] In regional tourist destinations, a lack of tourist guides means that tourists cannot obtain sufficient information. Furthermore, tourist information is often provided in standard Japanese, leaving few opportunities to experience the unique culture and dialects of the region. Furthermore, there is a lack of systems that can quickly and accurately respond to the diverse questions tourists ask. Another issue is how to efficiently incorporate new tourist information provided by users and utilize it in future tourism services.

[1299] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1300] In this invention, the server includes a means for acquiring location information, a means for searching for information on tourist spots based on the acquired location information, a means for providing the user with the searched information on tourist spots, a means for displaying an AI character that speaks in the local dialect and providing tourist information, a means for accepting questions from the user, generating answers based on the questions, and providing the answer to the user, a means for displaying the AI ​​character in augmented reality based on the user's location information, and a means for storing the tourist spot information provided by the user in a database and updating the AI ​​model. This solves the problem of a lack of tourist guides and allows users to have new experiences through local information and dialects. Furthermore, the server can respond quickly and accurately to user questions and provide the latest tourist information that is continuously updated.

[1301] "Location information" is information that indicates the user's current geographical location.

[1302] "Means of acquisition" refers to the method or apparatus by which a device or system collects location information or other data.

[1303] "Tourist destination" refers to a place or attraction that tourists are expected to visit.

[1304] A "search means" refers to a method or device for finding required information from a database or information source based on specific criteria.

[1305] "Means for providing" refers to the method or device for making information or services available to users.

[1306] "Local dialect" refers to linguistic expressions and ways of speaking that are unique to a region.

[1307] "Artificial intelligence characters" refer to virtual people or characters with specific roles that are generated using AI technology.

[1308] "Tourist information" refers to activities and services that provide information and guide tourists about tourist destinations and attractions.

[1309] "Means for accepting queries" refers to a method or device for accepting and processing inquiries or questions from users.

[1310] "Answer generation means" refers to a method or device that generates an appropriate response to a user's question.

[1311] "Augmented reality display" refers to a technology that overlays digital information on the real world.

[1312] A "database" refers to a collection of information that is systematically organized and made available for efficient search and use.

[1313] An "artificial intelligence model" refers to a conceptual model of an AI system that is trained based on large amounts of data to perform specific tasks.

[1314] The present invention is a system for supplementing the lack of tourist guides in tourist destinations and providing tourists with information and experiences unique to the region. This system utilizes location information and provides tourist guidance using an AI character that speaks in the local dialect. Specific embodiments of this system are described in detail below.

[1315] Hardware and Software Configuration

[1316] Device:

[1317] The device is a user's smartphone or tablet. This device is equipped with a GPS function, a camera, a microphone, and a speaker. The device also has network connectivity (Wi-Fi or mobile data).

[1318] server:

[1319] A cloud server is used to perform location analysis, database management, natural language processing (NLP), and augmented reality (AR) calculations. The following software and services run on the server.

[1320] DBMS (Database Management System): Manages tourist spot information, user-provided information, etc.

[1321] NLP engine: Analyzes user questions and generates appropriate answers

[1322] AR engine: Calculates the display position of the AI ​​character

[1323] Program processing overview

[1324] Device:

[1325] 1. When a user visits a tourist spot, the device uses GPS to obtain current location information, which is then sent to the server in real time.

[1326] 2. The received tourist spot information is displayed on the user interface. For example, it might say, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[1327] 3. The displayed tourist attraction information is then provided to the user by an AI character speaking in the local dialect. For example, the guide will say, "If you go straight ahead, you will come to a famous local shrine."

[1328] server:

[1329] 1. The server searches the database based on the received location information and obtains information about tourist attractions in the vicinity of the location (detailed descriptions, photos, reviews, etc.).

[1330] 2. Analyze the received question using natural language processing technology and generate an appropriate answer.

[1331] 3. Based on the received camera footage and location information, the display position of the AI ​​character is calculated and sent to the device.

[1332] Specific examples

[1333] As a concrete example, consider the case where a user visits a mountain called "Mount XX" in a local tourist spot. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server retrieves information about the peak of nearby tourist spot "Mount XX" from a database and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[1334] Prompt Sentence Examples

[1335] The following prompt sentences could be input to the generative AI model:

[1336] Example: "Please explain in detail the system that displays an AI character that guides users in the local dialect when they visit a local tourist spot."

[1337] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[1338] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1339] Step 1:

[1340] Terminal: When a user visits a tourist spot, the terminal uses a GPS module to obtain current location information. The input is the user's physical location information, and the output is GPS coordinate data (latitude and longitude). Specifically, the GPS module determines the location of the spot, and the data is processed internally on the terminal to generate coordinate information. This coordinate information is then sent to the server.

[1341] Step 2:

[1342] Terminal: Sends the acquired location information to the server. The input is the GPS coordinate data acquired in step 1, and the output is the location data sent to the server. Specifically, the terminal uses a network communication module (e.g., Wi-Fi or mobile data communication) to send the location data to the specified endpoint on the server.

[1343] Step 3:

[1344] Server: Based on the received location information, the server searches a database. The input is location data, and the output is information about tourist attractions near the location. Specifically, the server performs a database query to obtain detailed information (e.g., name, description, photos, reviews, etc.) about the nearest tourist attraction based on the location information.

[1345] Step 4:

[1346] Server: Sends the searched tourist attraction information to the user's device. The input is the tourist attraction information retrieved from the database, and the output is the tourist attraction information data sent to the device. Specifically, the server packages the tourist attraction information in an appropriate data format (e.g., JSON format) and sends it to the user's device via the network.

[1347] Step 5:

[1348] Terminal: The terminal displays the information it receives on its user interface. The input is tourist attraction information data sent from the server, and the output is tourist attraction information that is visually displayed to the user. Specifically, the terminal parses the received JSON data and displays the information in an appropriate format on the application's user interface. For example, it might display, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[1349] Step 6:

[1350] Terminal: An AI character speaking in the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is the visually displayed tourist attraction information, and the output is audio guidance information. Specifically, the terminal inputs text data into a speech synthesis engine (e.g., a TTS engine), which generates and plays audio data in the local dialect.

[1351] Step 7:

[1352] User: The user types a question into the chat window of the terminal. The input is the user's text-based question, and the output is the question data stored in the terminal. In concrete terms, the user types a question into the chat window of the application and presses the send button.

[1353] Step 8:

[1354] Terminal: Sends the entered question to the server. The input is the question data from the user, and the output is the question data sent to the server. In concrete terms, the terminal sends the question data to the server via the network communication module.

[1355] Step 9:

[1356] Server: Analyzes the received question using natural language processing technology and generates an appropriate answer. The input is the user's question data, and the output is the generated answer data. Specifically, the server uses a natural language processing engine to analyze the question and generate the optimal answer from a database or knowledge base.

[1357] Step 10:

[1358] Server: Sends the generated answer to the terminal. The input is the answer data, and the output is the answer data sent to the terminal. Specifically, the server packages the answer data in an appropriate data format and sends it to the user's terminal via the network.

[1359] Step 11:

[1360] Terminal: The terminal displays the generated answer in the chat window, and the AI ​​character also answers aloud in the local dialect. The input is the answer data sent from the server, and the output is the answer information provided visually and aloud. Specifically, the terminal analyzes the received answer data and displays it as text in the chat window. It also uses a speech synthesis engine to generate and play back audio data in the local dialect.

[1361] Step 12:

[1362] User: The user selects the AR mode on the device and activates the camera. The input is the user's operation, and the output is the video data captured by the camera. Specifically, the user presses the AR mode button in the application to activate the device's camera.

[1363] Step 13:

[1364] Terminal: Captures camera images and sends them to the server in real time. The input is the camera's image data, and the output is the image data sent to the server. Specifically, the terminal acquires image data from the camera module, compresses it, and sends it to the server.

[1365] Step 14:

[1366] Server: Based on location information and camera footage, calculates the display position of the AI ​​character and sends this to the device. The input is location information and video data, and the output is character display position data. Specifically, the server uses the AR engine to analyze the local geographic information and video data and calculate the optimal character display position.

[1367] Step 15:

[1368] Terminal: Based on the received data, an AI character is superimposed on the camera image. The input is character display position data and camera image data, and the output is the character displayed in AR. In concrete terms, the terminal superimposes a virtual character on the camera image based on the received position data. For example, when you stand in front of a tourist attraction, an AI guide character will appear in AR and begin to explain.

[1369] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[1370] (Application example 1)

[1371] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1372] Currently, there is a significant shortage of tourist guides and store guides in regional tourist destinations and shopping malls. Furthermore, tourists and visitors have limited access to information specific to their region, raising concerns about declining satisfaction. Furthermore, information about tourist destinations and stores is often provided in standard Japanese, which means that the characteristics and atmosphere of the region cannot be fully conveyed. This makes it difficult to maximize the appeal of tourist destinations and shopping malls and provide visitors with personalized, attractive guides. Therefore, there is a need to develop a system that can efficiently provide tourist guides and store guides and offer new experiences to visitors.

[1373] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1374] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving tourist attraction information provided by the user in a database and updating the AI ​​model, means for searching for store information within a shopping mall from the database, means for using an AI character that provides guidance in the local dialect and displaying it on a smartphone camera image, means for saving new store information and photos posted by the user in the database, and means for users to input questions in a chat format and for analyzing and answering them using natural language processing technology. This enables tourist spots and shopping malls to efficiently provide visitors with personalized guides and information using the local dialect.

[1375] "Location information acquisition means" refers to a system in which a user's terminal acquires the current location using GPS or other technology.

[1376] The "tourist attraction information search means" is a function that allows the server to search the database for information on the relevant tourist attraction based on the acquired location information.

[1377] The "tourist attraction information providing means" is a function that transmits information about the searched tourist attractions to the user terminal and displays it on the user interface.

[1378] "AI characters that speak in local dialects" refers to artificial intelligence characters that provide audio guidance using a dialect specific to the region.

[1379] The "question answering means" is a function that accepts questions from users, analyzes them, and generates and provides appropriate answers.

[1380] "Augmented reality display means" is a technology that displays virtual information superimposed on real-world images, and in this case refers to the function of displaying an AI character superimposed on a real-world scene.

[1381] The "user-provided information storage means" is a function that stores information and photos of tourist spots provided by users in a database.

[1382] "AI model updating means" refers to a method for retraining an artificial intelligence model based on stored user-provided information to improve its accuracy.

[1383] The "shopping mall information search means" is a function for searching for information on nearby stores based on the current location within the shopping mall.

[1384] "Natural language processing technology" is a technology that allows a computer to analyze a user's question, understand human language, and generate an appropriate answer.

[1385] This system complements the lack of tourist guides and store guides in shopping malls and local tourist destinations, providing visitors with a personalized new experience. This system uses an AI character that provides audio guidance in the local dialect and displays the information on the smartphone camera using augmented reality (AR). It also includes a function that uses natural language processing (NLP) to generate appropriate answers to user questions.

[1386] First, when a user visits a tourist spot or shopping mall, the location information acquisition means uses the smartphone's GPS to acquire current location information. This location information is sent to the server. Based on the received location information, the server searches a database for information on tourist attractions and stores in the vicinity of the location. For example, it provides information such as, "There is a recommended restaurant 200 meters from your current location."

[1387] Next, the tourist attraction information provider sends the searched tourist attraction and store information to the user's smartphone. The displayed information is guided by an AI character speaking the local dialect. For example, guidance such as "If you turn right at the next street, you will find a delicious Japanese restaurant" is provided.

[1388] The system also includes a chat-style question-and-answer mechanism. Users enter questions into a chat window on their smartphone, which is then sent to the server. The server uses natural language processing technology to analyze the question and generate an appropriate answer. For example, in response to the question, "What are some recommended cafes nearby?", the server will respond with, "The recommended nearby cafe is on the second floor of the mall."

[1389] Furthermore, as an augmented reality display method, when a user activates the smartphone camera and selects AR mode, an AI character is superimposed on the camera image, allowing the user to receive guidance while viewing the AI ​​character over the real-world scene.

[1390] The user-provided information storage means provides a mechanism that allows users to post new tourist spots and store information to the app. The posted information is sent to the server and stored in a database. This information is used as learning data for the AI ​​model, and the next time other users visit the same location, more up-to-date information will be provided.

[1391] As a concrete example, consider a smartphone app for tourists visiting a local shopping mall. As a user walks through the mall, their current location is acquired via GPS and sent to a server. The server uses this location information to search for nearby store information and provides information such as, "There's a new cafe 200 meters from your current location." When a user asks in a chat window, "What are some recommended restaurants nearby?", an answer analyzed using NLP technology is displayed. In addition, when using AR mode, an AI character appears on the camera image, providing specific directions.

[1392] The tour guide using the generative AI model generates a detailed scenario using the following prompt sentence:

[1393] ---

[1394] When a user visits a local shopping mall, provide an application that uses the smartphone's location information to provide voice guidance on recommended stores and tourist spots in the area. The application should include a feature that uses an AI character that speaks the local dialect, overlaying the AI ​​character on the screen in AR, and provide specific guidance. It should also have the ability to search a database based on GPS coordinates, analyze user questions using NLP technology, and allow users to post new store information. Please explain with specific examples.

[1395] ---

[1396] In this way, the present invention aims to make up for the lack of guides at tourist spots and shopping malls and provide visitors with a new experience.

[1397] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1398] Step 1:

[1399] The server receives GPS location information sent from the user's device. The input is the coordinate data of the current location obtained from the smartphone's GPS sensor. Based on this data, the server searches a database for information on relevant tourist attractions and stores in the shopping mall. As output, it generates information on the tourist attractions and stores found.

[1400] Step 2:

[1401] The server sends the searched tourist attraction and store information to the user terminal. The tourist attraction and store information generated in step 1 is the input. The server formats this and sends it to the user terminal in an appropriate format. The output is information data to be displayed on the user terminal.

[1402] Step 3:

[1403] The device uses an AI character to provide audio guidance in the local dialect based on the received information. The inputs include tourist attraction and store information sent from the server and the AI ​​character's audio data. The device combines this information to output audio and display on the screen. The output is audio guidance and visual guidance information presented to the user.

[1404] Step 4:

[1405] The user enters a question into the chat window on the smartphone. The input is a text question. The device sends this to the server. The output is the user question data.

[1406] Step 5:

[1407] The server analyzes the received question using natural language processing (NLP) technology and generates an appropriate answer. The input is the question data entered by the user. The server inputs this data into an NLP model and generates an answer based on the analysis results. The output is an appropriate answer text.

[1408] Step 6:

[1409] The server sends the generated answer to the user terminal. The input is the answer text generated in step 5. The server formats it and sends it to the user terminal. The output is the answer data to be displayed on the user terminal.

[1410] Step 7:

[1411] The device displays the received response in the user's chat window and responds audibly if necessary. The inputs are the response data sent from the server and the AI ​​character's voice data. The device combines these and outputs text and voice. The output is a visual and voice response to the user.

[1412] Step 8:

[1413] The user activates the smartphone camera and selects AR mode. The input is the camera image. The device captures it and sends it to the server. The output is real-time video data sent to the server.

[1414] Step 9:

[1415] The server calculates the display position of the AI ​​character based on the received camera image and location information, and sends this data to the device. The inputs are camera image data and GPS location information. The server processes this data and generates display position data for the AI ​​character. The output is sent to the device.

[1416] Step 10:

[1417] The device overlays the AI ​​character on the camera image based on the received AI character position data. The inputs are camera image data and AI character position data. The device combines these and displays them. The output is an augmented reality image from the user's perspective.

[1418] Step 11:

[1419] Users post information about new tourist spots and shops to the app. The input is text and image data. The device sends this to the server. The output is the posted data sent to the server.

[1420] Step 12:

[1421] The server stores the received posting data in a database and uses it as training data for the AI ​​model. The input includes tourist spot and store information data posted by users. The server stores this in the database and retrains the AI ​​model to improve its accuracy. As output, the new information data is stored in the database and an updated AI model is generated.

[1422] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1423] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR display, and an emotion engine to realize a tourist guide that responds to the user's emotions. Each component and its operation are explained in detail below.

[1424] Location information acquisition means

[1425] Device:

[1426] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[1427] Tourist attraction information search methods

[1428] server:

[1429] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[1430] Tourist attraction information provision method

[1431] server:

[1432] Tourist attraction information from search results is sent to the device.

[1433] Device:

[1434] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[1435] Tourist guide using AI characters in local dialects

[1436] Device:

[1437] The displayed tourist spot information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[1438] Chat-style question and answering tool

[1439] User:

[1440] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[1441] Device:

[1442] The entered question is sent to the server.

[1443] server:

[1444] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[1445] Device:

[1446] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[1447] AR display guide

[1448] User:

[1449] The user selects the AR mode on the device and activates the camera.

[1450] Device:

[1451] The camera captures the video and sends it to the server in real time.

[1452] server:

[1453] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[1454] Device:

[1455] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides sightseeing information.

[1456] A learning tool for user-provided information on famous places

[1457] User:

[1458] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[1459] Device:

[1460] The post content is sent to the server.

[1461] server:

[1462] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[1463] server:

[1464] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[1465] Response methods using emotion engines

[1466] Device:

[1467] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[1468] server:

[1469] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[1470] Device:

[1471] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[1472] Specific examples

[1473] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[1474] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In addition, by using AR mode, an AI character will appear on the camera image, providing visual guidance to the destination.

[1475] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas to help the user continue sightseeing in comfort.

[1476] As a result, the system of the present invention can provide a personalized tourist experience by addressing the lack of tourist guides and information specific to the region, as well as responding appropriately to the user's emotions.

[1477] The processing flow will be explained below.

[1478] Specific processing steps will be described below.

[1479] Step 1:

[1480] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[1481] Step 2:

[1482] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[1483] Step 3:

[1484] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and set up.

[1485] Step 4:

[1486] Server: Sends the data of the configured tourist guide AI character to the terminal.

[1487] Step 5:

[1488] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[1489] Step 6:

[1490] User: Visits a tourist spot and starts moving around with the device in hand.

[1491] Step 7:

[1492] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[1493] Step 8:

[1494] Server: Based on the received location information, search the database for nearby tourist attractions.

[1495] Step 9:

[1496] Server: Generates relevant tourist attraction information and sends it to the terminal.

[1497] Step 10:

[1498] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[1499] Step 11:

[1500] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[1501] Step 12:

[1502] Terminal: Sends the entered question to the server.

[1503] Step 13:

[1504] Server: Analyzes the question using natural language processing technology and generates an appropriate answer.

[1505] Step 14:

[1506] Server: Sends the generated answer to the device.

[1507] Step 15:

[1508] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[1509] Step 16:

[1510] User: Select the AR mode on the device and activate the camera.

[1511] Step 17:

[1512] Terminal: Captures camera images and sends them to the server in real time.

[1513] Step 18:

[1514] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the 3D model data and location information to the device.

[1515] Step 19:

[1516] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI character provides sightseeing information.

[1517] Step 20:

[1518] Users: When they discover a new tourist spot, they post a photo and description to the app.

[1519] Step 21:

[1520] Terminal: Sends the posted content to the server.

[1521] Step 22:

[1522] Server: Saves new tourist spot information in a database and uses it as training data for the AI ​​model.

[1523] Step 23:

[1524] Server: Updates the AI ​​model based on the learning results and provides the latest tourist spot information to other users from the next time onwards.

[1525] Step 24:

[1526] Terminal: Analyzes the user's facial expressions and voice, and recognizes their emotional state using an emotion engine.

[1527] Step 25:

[1528] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine.

[1529] Step 26:

[1530] Terminal: Providing information according to emotions.

[1531] Specific examples

[1532] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of a nearby tourist spot called "Mount XX," and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters away from here," and sends it to the device.

[1533] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In addition, by using AR mode, an AI character will appear on the camera image, providing visual guidance to the destination.

[1534] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas to help the user continue sightseeing in comfort.

[1535] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[1536] Example 2

[1537] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1538] In traditional tourist destinations, the number of guides is limited, making it difficult to provide individualized and satisfying tours to all visitors. There is also a lack of information that responds to visitors' emotions and personal preferences. Furthermore, there is a lack of technological means to provide advanced guidance and guidance at tourist destinations. This prevents visitors from fully enjoying the destinations, hindering the revitalization of the tourism industry as a whole.

[1539] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1540] In this invention, the server includes a means for acquiring location information, a means for searching for tourist attraction information, and a means for providing tourist attraction information to users, thereby making it possible to provide personalized tourist information to visitors and increase their satisfaction.

[1541] "Location information" is data that indicates a user's current geographic location.

[1542] "Tourist attraction" refers to a place or facility that is of interest to visitors for tourism purposes.

[1543] "Searching for information" is the act of locating relevant information from databases or sources based on specific criteria.

[1544] "Providing" is the act of giving or displaying information or services to a user.

[1545] An "artificial intelligence character" is a virtual guide character programmed using AI technology.

[1546] "Augmented reality display" is a technology that displays digital information and objects overlaid on real-world scenery.

[1547] A "database" is a collection of information for efficiently storing, searching, and managing large amounts of data.

[1548] An "artificial intelligence model" is an implementation of an AI algorithm that learns and makes inferences based on data.

[1549] "Emotional state" is information that indicates the psychological and physiological state of the user.

[1550] This invention is a system that compensates for the lack of guides in local tourist destinations and provides personalized tourist experiences. This system incorporates GPS functionality, natural language processing (NLP) technology, AR technology, and an emotion recognition engine. Below, we will explain in detail each component that makes up the system and its operation.

[1551] Location information acquisition means

[1552] Device: When a user visits a tourist spot, the device uses its GPS function to obtain current location information. This location information is periodically sent to the server. For example, the GPS module of a smartphone can be used to obtain the user's location and send it to the server as an HTTP request.

[1553] Tourist attraction information search methods

[1554] Server: Based on the location information received from the device, the server searches a database for information on nearby tourist attractions. This database contains detailed descriptions of tourist attractions, photos, user reviews, etc. The server executes SQL queries to search for this information.

[1555] Tourist attraction information provision method

[1556] Server: Sends tourist attraction information obtained as search results to the terminal.

[1557] Terminal: The terminal displays the information received from the server on a user interface. For example, it may display information such as, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[1558] Tourist guide using AI characters in local dialects

[1559] Device: The displayed tourist attraction information is provided to the user by an AI character that speaks the local dialect. For example, the device uses a speech synthesis engine to generate voice guidance in the local dialect, such as "If you go straight ahead, you will find a famous local shrine."

[1560] Chat-style question and answering tool

[1561] User: If the user has a question, they type it into the chat window on their device. For example, "What are some recommended restaurants nearby?"

[1562] Terminal: The terminal sends the entered question to the server.

[1563] Server: The server analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[1564] Device: The device will display the answers received from the server in a chat window, and the guide AI character will read the answers aloud in the local dialect.

[1565] AR display guide

[1566] User: The user selects the AR mode on the device and activates the camera.

[1567] Terminal: The terminal captures the camera image and transmits it to the server in real time.

[1568] Server: The server calculates the display position of the AI ​​character based on the location information and captured video, and sends the AI ​​character's 3D model data and location information to the device.

[1569] Terminal: Based on the received data, the terminal displays an AI character superimposed on the camera image. The guide AI character provides sightseeing information.

[1570] A learning tool for user-provided information on famous places

[1571] User: When a user finds a new tourist spot, they post a photo and description of it to the app. For example, they post something like, "I found a new photo spot."

[1572] Terminal: The terminal sends the posted content to the server.

[1573] Server: The server saves new tourist spot information in a database and uses it as learning data for the AI ​​model. It then updates the AI ​​model based on the learning results, enabling it to provide the latest tourist spot information to other users from the next time onwards.

[1574] Response methods using emotion engines

[1575] Device: The device analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[1576] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine.

[1577] Device: The device provides information based on the user's emotions. For example, if the device detects that the user is tired, it will suggest nearby rest spots or cafes.

[1578] Specific examples

[1579] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[1580] Based on this information, the device will have an AI guide character provide voice guidance in the local dialect. For example, it might say, "500 meters ahead there is a peak with a spectacular view." In addition, when using AR mode, the AI ​​character will appear on the camera image, providing visual guidance to the destination.

[1581] Furthermore, if the emotion engine detects that the user is tired, it will suggest nearby rest spots and relaxation areas. This allows the system of the present invention to address the lack of tourist guides and local information, and respond appropriately to the user's emotions, providing a personalized sightseeing experience.

[1582] Prompt Sentence Examples

[1583] The following are some examples of input prompts for a generative AI model:

[1584] "Please introduce us to a new tourist guide system. This system will address the lack of guides in local tourist destinations and provide a personalized tourist experience. Specifically, it will use location information acquisition, AI characters that speak the local dialect, AR displays, and an emotion engine to provide a tourist guide that responds to the user's emotions."

[1585] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1586] Step 1: Obtaining location information

[1587] Terminal: When a user visits a tourist spot, the terminal's GPS function is activated and acquires current location information. The input is geographic coordinate data from the GPS sensor, and the output is current location information. This location information is sent to the server at regular intervals.

[1588] Specific operation: The device uses the built-in GPS module to obtain geographic coordinates (latitude and longitude) and transmits them to the server as location information.

[1589] Step 2: Search for tourist attraction information

[1590] Server: Receives the acquired location information and searches for information on nearby tourist attractions. The input is the location information and the output is a list of nearby tourist attractions. The server searches the database using an SQL query and obtains the search results.

[1591] Specific operation: Based on the received location information, the server performs a range search on the tourist attraction database to obtain detailed information about related tourist attractions.

[1592] Step 3: Providing tourist attraction information

[1593] Server: Sends tourist attraction information obtained as search results to the terminal. The input is tourist attraction information, and the output is the transmission of information to the terminal.

[1594] Specific operation: The server composes the search results in JSON format and sends them to the terminal as an HTTP response.

[1595] Step 4: View tourist attraction information

[1596] Terminal: The terminal displays tourist attraction information received from the server on a user interface. The input is tourist attraction information from the server, and the output is the information displayed on the display.

[1597] Specific operation: The terminal parses the received JSON data, binds the information to the user interface, and displays it.

[1598] Step 5: Tourist guide in local dialect

[1599] Terminal: An AI character speaking the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is text information about the tourist attraction, and the output is synthesized speech.

[1600] Specific operation: The device passes the text information to a TTS (Text-to-Speech) engine, synthesizes it in the local dialect, and plays it back through the speaker.

[1601] Step 6: Chat-style Q&A

[1602] User: The user types a question into the chat window on the device. The input is the text of the user's question.

[1603] Terminal: The terminal sends the input question to the server. The output is the question data sent to the server.

[1604] Specific behavior: The user types a question into the chat interface, the device captures it, and sends it to the server via an HTTP POST request.

[1605] Server: The server analyzes the question content using NLP technology and generates an appropriate answer. The input is the user's question text, and the output is the generated answer text.

[1606] What it does: The server uses an NLP model to analyze the question and generate an appropriate answer from a database or other source.

[1607] Terminal: The terminal displays the answer received from the server in a chat window and has an AI character read it aloud in the local dialect. The input is the answer text from the server, and the output is the displayed answer and audio.

[1608] Specific operation: The device displays the received reply text in the chat window, synthesizes it using the TTS engine, and plays it back.

[1609] Step 7: Guided by AR display

[1610] User: The user selects the device's AR mode and activates the camera. The input is the AR mode selection action.

[1611] Terminal: The terminal captures camera images and transmits them to the server in real time. The input is the camera image and the output is the video data sent to the server.

[1612] Specific operation: The device captures camera footage and streams it to the server along with location information.

[1613] Server: The server calculates the display position of the AI ​​character based on the location information and video data, and sends the 3D model data to the terminal. The input is the camera image and location information, and the output is the display position of the AI ​​character and 3D model data.

[1614] Specific operation: The server uses the AR calculation engine to calculate the appropriate display position and send the 3D model data to the terminal.

[1615] Terminal: Based on the received data, the terminal displays an AI character superimposed on the camera image. The input is 3D model data and display position information, and the output is an AI character superimposed on the camera image.

[1616] What it does: The device uses an AR library to render a 3D model overlaid on the camera image.

[1617] Step 8: Learning user-provided points of interest information

[1618] User: When a user finds a new tourist spot, he / she posts a photo and description of the spot to the app. The input is a photo of the new tourist spot and text information.

[1619] Terminal: The terminal sends the posted content to the server. The output is the user posted data sent to the server.

[1620] Specific operation: The device collects posting data including information entered by the user and sends it to the server as an HTTP POST request.

[1621] Server: The server stores the received information in a database and uses it as training data for the AI ​​model. The input is new tourist attraction information, and the output is an updated database and AI model.

[1622] Specific operation: The server stores the received post data in a database and reflects it in the AI ​​model in the next learning phase.

[1623] Step 9: Respond with the Emotion Engine

[1624] Terminal: The terminal analyzes the user's facial expressions and voice and recognizes their emotional state using an emotion engine. The input is the user's facial expression data and voice data, and the output is the recognized emotional information.

[1625] How it works: The device uses a camera and microphone to capture the user's facial expressions and voice, and applies emotion analysis algorithms to identify their emotional state.

[1626] Server: Adjusts tourist attraction information and guidance content based on the user's emotions recognized by the emotion engine. The input is the recognized emotion information, and the output is adjusted tourist attraction information.

[1627] Specific operation: The server dynamically adjusts the method and content of providing tourist attraction information based on the emotion data.

[1628] Terminal: The terminal provides tailored information to the user. For example, if the terminal recognizes that the user is tired, it suggests nearby rest spots or cafes. The input is tailored tourist attraction information, and the output is the provided information and guidance.

[1629] Specific operation: Based on the adjusted information, the device updates the user interface and voice guidance to provide the user with the most appropriate information.

[1630] (Application example 2)

[1631] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1632] There is a shortage of tourist guides in local tourist destinations, and it is difficult to provide tourists with a personalized guided experience. Furthermore, there is a demand for dynamic information provision based on the user's emotions and location information. Furthermore, tourist guides using local dialects and visual guidance using augmented reality displays are insufficient. A means is needed to address these issues and provide a richer tourist experience.

[1633] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1634] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving the tourist attraction information provided by the user in a database and updating the AI ​​model, means for analyzing emotions from the user's facial expressions and voice and providing information based on the emotions, means for capturing camera footage and transmitting it to the server in real time, and means for calculating the display position of the AI ​​character based on the location information and the captured footage and superimposing the AI ​​character on the camera footage. This makes it possible to address the lack of tourist destination guides and local information, and to provide personalized tour guidance based on the user's emotions, as well as a visual guide experience.

[1635] Key Word Definitions

[1636] The "means for acquiring location information" is a means for acquiring the user's current location information using the GPS function and periodically transmitting it to the server.

[1637] The "means for searching tourist attraction information" is a means for searching a database for information on tourist attractions that exist in the vicinity of a point based on the acquired location information.

[1638] The "means for providing tourist attraction information to the user" refers to a means for displaying information about the searched tourist attraction to the user through a user interface.

[1639] "A means for displaying an AI character that speaks in the local dialect and provides tourist guidance" is a means for an AI character displayed to the user to provide tourist guidance using the local dialect.

[1640] The "means for generating an answer according to the question content and providing it to the user" is a means for analyzing the question entered by the user, generating an appropriate answer, and providing it to the user in a chat window or by voice.

[1641] "Means for displaying AI characters in augmented reality" refers to a means for displaying a 3D model of an AI character overlaid on camera images based on location information and captured images.

[1642] "Means for saving tourist spot information in a database and updating the AI ​​model" refers to a means for saving new tourist spot information provided by users in a database and using that information to learn and update the AI ​​model.

[1643] The "means for analyzing emotions and providing information based on emotions" refers to a means for analyzing a user's facial expressions and voice to recognize emotions and providing information according to those emotions.

[1644] "Means for capturing camera images and transmitting them to a server in real time" refers to means for capturing images using a camera on a user terminal and transmitting the images to a server in real time.

[1645] "Means for calculating the display position and superimposing and displaying an AI character on the camera image" refers to means for calculating the display position of an AI character based on the acquired position information and camera image, and superimposing and displaying the AI ​​character on the camera image at that position.

[1646] MODE FOR CARRYING OUT THE INVENTION

[1647] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR displays, an emotion engine, and more to provide a tourist guide that responds to the user's emotions. Each component and its operation are described in detail below.

[1648] Location information acquisition means

[1649] Device:

[1650] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[1651] Tourist attraction information search methods

[1652] server:

[1653] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[1654] Tourist attraction information provision method

[1655] server:

[1656] Tourist attraction information from search results is sent to the device.

[1657] Device:

[1658] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[1659] Tourist guide using AI characters in local dialects

[1660] Device:

[1661] The displayed tourist attraction information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[1662] Chat-style question and answering tool

[1663] User:

[1664] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[1665] Device:

[1666] The entered question is sent to the server.

[1667] server:

[1668] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[1669] Device:

[1670] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[1671] AR display guide

[1672] User:

[1673] The user selects the AR mode on the device and activates the camera.

[1674] Device:

[1675] The camera captures the video and sends it to the server in real time.

[1676] server:

[1677] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[1678] Device:

[1679] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides tourist information.

[1680] A learning tool for user-provided information on famous places

[1681] User:

[1682] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[1683] Device:

[1684] The post content is sent to the server.

[1685] server:

[1686] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[1687] server:

[1688] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[1689] Response methods using emotion engines

[1690] Device:

[1691] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[1692] server:

[1693] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[1694] Device:

[1695] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[1696] Specific examples

[1697] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant location, and generates a guide message, such as "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device. Based on this guide message, an AI guide character on the device provides audio guidance in the local dialect. In addition, by using AR mode, an AI character appears on the camera image, providing visual guidance to the destination. Furthermore, if the emotion engine detects that the user appears tired, it will suggest nearby rest spots and relaxation areas, helping the user continue sightseeing comfortably.

[1698] As a result, the system of the present invention can provide a personalized tourist experience by addressing the lack of tourist guides, the lack of local information, and the user's emotions.

[1699] Example prompt sentence:

[1700] A user visits a local tourist destination and wears smart glasses to travel to the nearest tourist attraction from their current location. The glasses provide appropriate information based on the user's emotions, and audio guidance is provided in the local dialect. Information about tourist attractions is displayed in AR.

[1701] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1702] Program processing flow

[1703] Step 1:

[1704] Obtaining location information

[1705] (input)

[1706] The terminal uses the GPS function to obtain the user's current location information.

[1707] (Data processing)

[1708] Location information is obtained in the form of latitude and longitude and is updated periodically.

[1709] (output)

[1710] The acquired location information is temporarily stored on the device and then sent to the server.

[1711] Step 2:

[1712] Search for tourist attraction information

[1713] (input)

[1714] Based on the location information received from the terminal, the server searches a database for information on tourist attractions in the vicinity of the relevant location.

[1715] (Data calculation)

[1716] Execute a database query to retrieve a list of attractions within a specified radius.

[1717] (output)

[1718] Search results include detailed information about tourist attractions (descriptions, photos, and user reviews).

[1719] Step 3:

[1720] Providing tourist attraction information

[1721] (input)

[1722] The server transmits the retrieved tourist attraction information to the terminal.

[1723] (Data calculation)

[1724] The acquired tourist attraction information is converted into an appropriate format and sent via the API.

[1725] (output)

[1726] The terminal displays the received tourist attraction information on a user interface.

[1727] Step 4:

[1728] Guided by an AI character who speaks the local dialect

[1729] (input)

[1730] Based on tourist attraction information received from the server, the device displays an AI character that speaks the local dialect and provides audio guidance.

[1731] (Data calculation)

[1732] Using speech synthesis technology, tourist attraction information is read out in the local dialect.

[1733] (output)

[1734] The user is provided with audio tourist information.

[1735] Step 5:

[1736] Chat-style question and answering

[1737] (input)

[1738] The user types a question into the chat window on the terminal.

[1739] (Data calculation)

[1740] The entered question is sent to the server, where it is analyzed using natural language processing technology and an appropriate answer is generated.

[1741] (output)

[1742] The generated answer is sent to the device, displayed in the chat window, and read aloud by an AI character.

[1743] Step 6:

[1744] AR display guide

[1745] (input)

[1746] The user selects the AR mode on the device and activates the camera.

[1747] (Data processing)

[1748] The device captures camera footage and transmits it to the server in real time.

[1749] (output)

[1750] The server calculates the display position of the AI ​​character based on the location information and captured video, and sends the AI ​​character's 3D model data and location information to the device. The device then displays the AI ​​character overlaid on the camera image based on the received data.

[1751] Step 7:

[1752] Learning user-provided points of interest information

[1753] (input)

[1754] When users discover a new tourist spot, they post a photo and description to the app.

[1755] (Data calculation)

[1756] The posted content is sent to the server and saved in the database as new tourist spot information.

[1757] (output)

[1758] Based on this, the AI ​​model is updated and the latest tourist spot information is provided to other users from the next time onwards.

[1759] Step 8:

[1760] Emotional engine response

[1761] (input)

[1762] The device analyzes the user's facial expressions and voice and uses an emotion engine to recognize the user's emotional state.

[1763] (Data calculation)

[1764] The emotion engine analyzes emotions based on the data it acquires and adjusts the information it provides accordingly.

[1765] (output)

[1766] The server selects appropriate tourist attraction information and guidance content based on the user's emotions and transmits them to the terminal, which then provides the adjusted information to the user.

[1767] Prompt Sentence Examples

[1768] A user visits a local tourist destination and wears smart glasses to travel to the nearest tourist attraction from their current location. The glasses provide appropriate information based on the user's emotions, and audio guidance is provided in the local dialect. Information about tourist attractions is displayed in AR.

[1769] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1770] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1771] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1772] [Fourth embodiment]

[1773] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1774] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1775] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1776] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1777] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1778] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1779] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1780] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1781] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1782] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1783] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1784] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1785] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1786] This invention is a system that complements the lack of tourist guides in local tourist destinations and provides tourists with a new experience. This system utilizes location information and features an AI character that speaks in the local dialect to provide tourist guidance. Below, we will explain each component of the system and its operation in detail.

[1787] Location information acquisition means

[1788] Device:

[1789] When a user visits a tourist spot, the device acquires the current location information using GPS, which is then sent to the server.

[1790] Tourist attraction information search methods

[1791] server:

[1792] Based on the received location information, the server searches a database for information on nearby tourist attractions, including detailed descriptions, photos, and reviews of the attractions.

[1793] Tourist attraction information provision method

[1794] server:

[1795] The searched tourist attraction information is transmitted to the user terminal.

[1796] Device:

[1797] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is displayed.

[1798] Tourist guide using AI characters in local dialects

[1799] Device:

[1800] The tourist attraction information displayed is provided to the user by an AI character speaking in the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[1801] Chat-style question and answering tool

[1802] User:

[1803] Users can type questions into a chat window on their device (e.g., "What are some recommended restaurants nearby?").

[1804] Device:

[1805] The entered question is sent to the server.

[1806] server:

[1807] The received question is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[1808] Device:

[1809] The generated answer is displayed in the chat window, and the AI ​​character also responds verbally in the local dialect.

[1810] AR display guide

[1811] User:

[1812] The user selects the AR mode on the device and activates the camera.

[1813] Device:

[1814] The camera captures the video and sends it to the server in real time.

[1815] server:

[1816] Based on the location information and captured video, the display position of the AI ​​character is calculated and sent to the device.

[1817] Device:

[1818] Based on the received data, an AI character is superimposed on the captured image. For example, if you stand in front of a tourist spot, an AI guide character will appear in AR and begin explaining.

[1819] A learning tool for user-provided information on famous places

[1820] User:

[1821] Users post photos and descriptions of new tourist spots to the app (e.g., when they discover a new photo spot).

[1822] Device:

[1823] The post content is sent to the server.

[1824] server:

[1825] The new information is stored in a database and used as training data for the AI ​​model.

[1826] server:

[1827] The learning results will be used to provide the latest tourist spot information to other users the next time they visit the same location.

[1828] Specific examples

[1829] As a concrete example, consider the case where a user visits a mountain in a local tourist area. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server obtains information about the peak of a nearby tourist attraction called "Mount XX," and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[1830] As a result, the system of the present invention can solve problems such as a lack of tourist guides and information specific to a particular region, and provide tourists with new experiences.

[1831] The processing flow will be explained below.

[1832] Specific processing steps will be described below.

[1833] Step 1:

[1834] User: Launches the app and begins the initial setup. Enters profile information (name, age, hobbies, sightseeing purpose, etc.).

[1835] Step 2:

[1836] Device: Sends user profile information to the server, and also uses GPS to obtain location information and sends it to the server.

[1837] Step 3:

[1838] Server: The received user profile information and location information is stored in a database, and an AI tourist guide character suited to the user is generated and configured.

[1839] Step 4:

[1840] Server: Sends the data of the configured tourist guide AI character to the terminal.

[1841] Step 5:

[1842] Terminal: The received tourist guide AI character is displayed on the user interface, and the initial settings are complete.

[1843] Step 6:

[1844] User: Visits a tourist spot and starts moving around with the device in hand.

[1845] Step 7:

[1846] Device: Periodically obtains location information using GPS and sends it to the server in real time.

[1847] Step 8:

[1848] Server: Based on the received location information, search the database for nearby tourist attractions.

[1849] Step 9:

[1850] Server: Generates relevant tourist attraction information and sends it to the terminal.

[1851] Step 10:

[1852] Terminal: The received tourist attraction information is displayed on the user interface, and the guide AI provides audio guidance in the local dialect.

[1853] Step 11:

[1854] User: Type a question into the chat window on the device and send it to the tourist guide AI.

[1855] Step 12:

[1856] Terminal: Sends the entered question to the server.

[1857] Step 13:

[1858] Server: Analyzes the question using natural language processing (NLP) technology and generates an appropriate answer.

[1859] Step 14:

[1860] Server: Sends the generated answer to the device.

[1861] Step 15:

[1862] Device: The received response is displayed in the chat window, and the guide AI responds verbally in the local dialect.

[1863] Step 16:

[1864] User: Select the AR mode on the device and activate the camera.

[1865] Step 17:

[1866] Terminal: Captures camera images and sends them to the server in real time.

[1867] Step 18:

[1868] Server: Based on the received video and location information, calculates the display position of the AI ​​character and sends the data to the device.

[1869] Step 19:

[1870] Terminal: Based on the received data, an AI character is superimposed on the camera image. The guide AI provides tourist information using AR display.

[1871] Step 20:

[1872] Users: When they discover a new tourist spot, they post a photo and description to the app.

[1873] Step 21:

[1874] Device: Sends the posted content (photos, text, location information, etc.) to the server.

[1875] Step 22:

[1876] Server: Stores new tourist spot information in a database and uses it as training data for the AI ​​model.

[1877] Step 23:

[1878] Server: Updates the AI ​​model based on the training data to improve the accuracy of the tourist guide AI.

[1879] This allows the user to understand the specific processing flow of each function provided by the system of the present invention.

[1880] Example 1

[1881] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1882] In regional tourist destinations, a lack of tourist guides means that tourists cannot obtain sufficient information. Furthermore, tourist information is often provided in standard Japanese, leaving few opportunities to experience the unique culture and dialects of the region. Furthermore, there is a lack of systems that can quickly and accurately respond to the diverse questions tourists ask. Another issue is how to efficiently incorporate new tourist information provided by users and utilize it in future tourism services.

[1883] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1884] In this invention, the server includes a means for acquiring location information, a means for searching for information on tourist spots based on the acquired location information, a means for providing the user with the searched information on tourist spots, a means for displaying an AI character that speaks in the local dialect and providing tourist information, a means for accepting questions from the user, generating answers based on the questions, and providing the answer to the user, a means for displaying the AI ​​character in augmented reality based on the user's location information, and a means for storing the tourist spot information provided by the user in a database and updating the AI ​​model. This solves the problem of a lack of tourist guides and allows users to have new experiences through local information and dialects. Furthermore, the server can respond quickly and accurately to user questions and provide the latest tourist information that is continuously updated.

[1885] "Location information" is information that indicates the user's current geographical location.

[1886] "Means of acquisition" refers to the method or apparatus by which a device or system collects location information or other data.

[1887] "Tourist destination" refers to a place or attraction that tourists are expected to visit.

[1888] A "search means" refers to a method or device for finding required information from a database or information source based on specific criteria.

[1889] "Means for providing" refers to the method or device for making information or services available to users.

[1890] "Local dialect" refers to linguistic expressions and ways of speaking that are unique to a region.

[1891] "Artificial intelligence characters" refer to virtual people or characters with specific roles that are generated using AI technology.

[1892] "Tourist information" refers to activities and services that provide information and guide tourists about tourist destinations and attractions.

[1893] "Means for accepting queries" refers to a method or device for accepting and processing inquiries or questions from users.

[1894] "Answer generation means" refers to a method or device that generates an appropriate response to a user's question.

[1895] "Augmented reality display" refers to a technology that overlays digital information on the real world.

[1896] A "database" refers to a collection of information that is systematically organized and made available for efficient search and use.

[1897] An "artificial intelligence model" refers to a conceptual model of an AI system that is trained based on large amounts of data to perform specific tasks.

[1898] The present invention is a system for supplementing the lack of tourist guides in tourist destinations and providing tourists with information and experiences unique to the region. This system utilizes location information and provides tourist guidance using an AI character that speaks in the local dialect. Specific embodiments of this system are described in detail below.

[1899] Hardware and Software Configuration

[1900] Device:

[1901] The device is a user's smartphone or tablet. This device is equipped with a GPS function, a camera, a microphone, and a speaker. The device also has network connectivity (Wi-Fi or mobile data).

[1902] server:

[1903] A cloud server is used to perform location analysis, database management, natural language processing (NLP), and augmented reality (AR) calculations. The following software and services run on the server.

[1904] DBMS (Database Management System): Manages tourist spot information, user-provided information, etc.

[1905] NLP engine: Analyzes user questions and generates appropriate answers

[1906] AR engine: Calculates the display position of the AI ​​character

[1907] Program processing overview

[1908] Device:

[1909] 1. When a user visits a tourist spot, the device uses GPS to obtain current location information, which is then sent to the server in real time.

[1910] 2. The received tourist spot information is displayed on the user interface. For example, it might say, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[1911] 3. The displayed tourist attraction information is then provided to the user by an AI character speaking in the local dialect. For example, the guide will say, "If you go straight ahead, you will come to a famous local shrine."

[1912] server:

[1913] 1. The server searches the database based on the received location information and obtains information about tourist attractions in the vicinity of the location (detailed descriptions, photos, reviews, etc.).

[1914] 2. Analyze the received question using natural language processing technology and generate an appropriate answer.

[1915] 3. Based on the received camera footage and location information, the display position of the AI ​​character is calculated and sent to the device.

[1916] Specific examples

[1917] As a concrete example, consider the case where a user visits a mountain called "Mount XX" in a local tourist spot. When the user starts climbing the mountain with their device, the device sends their current location information to the server. Based on this location information, the server retrieves information about the peak of nearby tourist spot "Mount XX" from a database and sends guidance to the device, such as "The view from the peak is spectacular, and it is located 500 meters from here." On the user's device, an AI character using the local dialect explains this information aloud, and by using AR mode, a guide AI character appears on the screen and provides specific guidance to the user.

[1918] Prompt Sentence Examples

[1919] The following prompt sentences could be input to the generative AI model:

[1920] Example: "Please explain in detail the system that displays an AI character that guides users in the local dialect when they visit a local tourist spot."

[1921] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[1922] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1923] Step 1:

[1924] Terminal: When a user visits a tourist spot, the terminal uses a GPS module to obtain current location information. The input is the user's physical location information, and the output is GPS coordinate data (latitude and longitude). Specifically, the GPS module determines the location of the spot, and the data is processed internally on the terminal to generate coordinate information. This coordinate information is then sent to the server.

[1925] Step 2:

[1926] Terminal: Sends the acquired location information to the server. The input is the GPS coordinate data acquired in step 1, and the output is the location data sent to the server. Specifically, the terminal uses a network communication module (e.g., Wi-Fi or mobile data communication) to send the location data to the specified endpoint on the server.

[1927] Step 3:

[1928] Server: Based on the received location information, the server searches a database. The input is location data, and the output is information about tourist attractions near the location. Specifically, the server performs a database query to obtain detailed information (e.g., name, description, photos, reviews, etc.) about the nearest tourist attraction based on the location information.

[1929] Step 4:

[1930] Server: Sends the searched tourist attraction information to the user's device. The input is the tourist attraction information retrieved from the database, and the output is the tourist attraction information data sent to the device. Specifically, the server packages the tourist attraction information in an appropriate data format (e.g., JSON format) and sends it to the user's device via the network.

[1931] Step 5:

[1932] Terminal: The terminal displays the information it receives on its user interface. The input is tourist attraction information data sent from the server, and the output is tourist attraction information that is visually displayed to the user. Specifically, the terminal parses the received JSON data and displays the information in an appropriate format on the application's user interface. For example, it might display, "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it."

[1933] Step 6:

[1934] Terminal: An AI character speaking in the local dialect provides audio guidance to the user about the displayed tourist attraction information. The input is the visually displayed tourist attraction information, and the output is audio guidance information. Specifically, the terminal inputs text data into a speech synthesis engine (e.g., a TTS engine), which generates and plays audio data in the local dialect.

[1935] Step 7:

[1936] User: The user types a question into the chat window of the terminal. The input is the user's text-based question, and the output is the question data stored in the terminal. In concrete terms, the user types a question into the chat window of the application and presses the send button.

[1937] Step 8:

[1938] Terminal: Sends the entered question to the server. The input is the question data from the user, and the output is the question data sent to the server. In concrete terms, the terminal sends the question data to the server via the network communication module.

[1939] Step 9:

[1940] Server: Analyzes the received question using natural language processing technology and generates an appropriate answer. The input is the user's question data, and the output is the generated answer data. Specifically, the server uses a natural language processing engine to analyze the question and generate the optimal answer from a database or knowledge base.

[1941] Step 10:

[1942] Server: Sends the generated answer to the terminal. The input is the answer data, and the output is the answer data sent to the terminal. Specifically, the server packages the answer data in an appropriate data format and sends it to the user's terminal via the network.

[1943] Step 11:

[1944] Terminal: The terminal displays the generated answer in the chat window, and the AI ​​character also answers aloud in the local dialect. The input is the answer data sent from the server, and the output is the answer information provided visually and aloud. Specifically, the terminal analyzes the received answer data and displays it as text in the chat window. It also uses a speech synthesis engine to generate and play back audio data in the local dialect.

[1945] Step 12:

[1946] User: The user selects the AR mode on the device and activates the camera. The input is the user's operation, and the output is the video data captured by the camera. Specifically, the user presses the AR mode button in the application to activate the device's camera.

[1947] Step 13:

[1948] Terminal: Captures camera images and sends them to the server in real time. The input is the camera's image data, and the output is the image data sent to the server. Specifically, the terminal acquires image data from the camera module, compresses it, and sends it to the server.

[1949] Step 14:

[1950] Server: Based on location information and camera footage, calculates the display position of the AI ​​character and sends this to the device. The input is location information and video data, and the output is character display position data. Specifically, the server uses the AR engine to analyze the local geographic information and video data and calculate the optimal character display position.

[1951] Step 15:

[1952] Terminal: Based on the received data, an AI character is superimposed on the camera image. The input is character display position data and camera image data, and the output is the character displayed in AR. In concrete terms, the terminal superimposes a virtual character on the camera image based on the received position data. For example, when you stand in front of a tourist attraction, an AI guide character will appear in AR and begin to explain.

[1953] As a result, the system of the present invention can solve the problem of a lack of tourist guides and information specific to a particular region, and provide tourists with a new experience.

[1954] (Application example 1)

[1955] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1956] Currently, there is a significant shortage of tourist guides and store guides in regional tourist destinations and shopping malls. Furthermore, tourists and visitors have limited access to information specific to their region, raising concerns about declining satisfaction. Furthermore, information about tourist destinations and stores is often provided in standard Japanese, which means that the characteristics and atmosphere of the region cannot be fully conveyed. This makes it difficult to maximize the appeal of tourist destinations and shopping malls and provide visitors with personalized, attractive guides. Therefore, there is a need to develop a system that can efficiently provide tourist guides and store guides and offer new experiences to visitors.

[1957] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1958] In this invention, the server includes means for acquiring location information, means for searching for tourist attraction information based on the acquired location information, means for providing the user with the searched tourist attraction information, means for displaying an AI character that speaks in the local dialect and providing a tour guide, means for accepting questions from the user and generating answers based on the questions and providing them to the user, means for displaying the AI ​​character in augmented reality based on the user's location information, means for saving tourist attraction information provided by the user in a database and updating the AI ​​model, means for searching for store information within a shopping mall from the database, means for using an AI character that provides guidance in the local dialect and displaying it on a smartphone camera image, means for saving new store information and photos posted by the user in the database, and means for users to input questions in a chat format and for analyzing and answering them using natural language processing technology. This enables tourist spots and shopping malls to efficiently provide visitors with personalized guides and information using the local dialect.

[1959] "Location information acquisition means" refers to a system in which a user's terminal acquires the current location using GPS or other technology.

[1960] The "tourist attraction information search means" is a function that allows the server to search the database for information on the relevant tourist attraction based on the acquired location information.

[1961] The "tourist attraction information providing means" is a function that transmits information about the searched tourist attractions to the user terminal and displays it on the user interface.

[1962] "AI characters that speak in local dialects" refers to artificial intelligence characters that provide audio guidance using a dialect specific to the region.

[1963] The "question answering means" is a function that accepts questions from users, analyzes them, and generates and provides appropriate answers.

[1964] "Augmented reality display means" is a technology that displays virtual information superimposed on real-world images, and in this case refers to the function of displaying an AI character superimposed on a real-world scene.

[1965] The "user-provided information storage means" is a function that stores information and photos of tourist spots provided by users in a database.

[1966] "AI model updating means" refers to a method for retraining an artificial intelligence model based on stored user-provided information to improve its accuracy.

[1967] The "shopping mall information search means" is a function for searching for information on nearby stores based on the current location within the shopping mall.

[1968] "Natural language processing technology" is a technology that allows a computer to analyze a user's question, understand human language, and generate an appropriate answer.

[1969] This system complements the lack of tourist guides and store guides in shopping malls and local tourist destinations, providing visitors with a personalized new experience. This system uses an AI character that provides audio guidance in the local dialect and displays the information on the smartphone camera using augmented reality (AR). It also includes a function that uses natural language processing (NLP) to generate appropriate answers to user questions.

[1970] First, when a user visits a tourist spot or shopping mall, the location information acquisition means uses the smartphone's GPS to acquire current location information. This location information is sent to the server. Based on the received location information, the server searches a database for information on tourist attractions and stores in the vicinity of the location. For example, it provides information such as, "There is a recommended restaurant 200 meters from your current location."

[1971] Next, the tourist attraction information provider sends the searched tourist attraction and store information to the user's smartphone. The displayed information is guided by an AI character speaking the local dialect. For example, guidance such as "If you turn right at the next street, you will find a delicious Japanese restaurant" is provided.

[1972] The system also includes a chat-style question-and-answer mechanism. Users enter questions into a chat window on their smartphone, which is then sent to the server. The server uses natural language processing technology to analyze the question and generate an appropriate answer. For example, in response to the question, "What are some recommended cafes nearby?", the server will respond with, "The recommended nearby cafe is on the second floor of the mall."

[1973] Furthermore, as an augmented reality display method, when a user activates the smartphone camera and selects AR mode, an AI character is superimposed on the camera image, allowing the user to receive guidance while viewing the AI ​​character over the real-world scene.

[1974] The user-provided information storage means provides a mechanism that allows users to post new tourist spots and store information to the app. The posted information is sent to the server and stored in a database. This information is used as learning data for the AI ​​model, and the next time other users visit the same location, more up-to-date information will be provided.

[1975] As a concrete example, consider a smartphone app for tourists visiting a local shopping mall. As a user walks through the mall, their current location is acquired via GPS and sent to a server. The server uses this location information to search for nearby store information and provides information such as, "There's a new cafe 200 meters from your current location." When a user asks in a chat window, "What are some recommended restaurants nearby?", an answer analyzed using NLP technology is displayed. In addition, when using AR mode, an AI character appears on the camera image, providing specific directions.

[1976] The tour guide using the generative AI model generates a detailed scenario using the following prompt sentence:

[1977] ---

[1978] When a user visits a local shopping mall, provide an application that uses the smartphone's location information to provide voice guidance on recommended stores and tourist spots in the area. The application should include a feature that uses an AI character that speaks the local dialect, overlaying the AI ​​character on the screen in AR, and provide specific guidance. It should also have the ability to search a database based on GPS coordinates, analyze user questions using NLP technology, and allow users to post new store information. Please explain with specific examples.

[1979] ---

[1980] In this way, the present invention aims to make up for the lack of guides at tourist spots and shopping malls and provide visitors with a new experience.

[1981] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1982] Step 1:

[1983] The server receives GPS location information sent from the user's device. The input is the coordinate data of the current location obtained from the smartphone's GPS sensor. Based on this data, the server searches a database for information on relevant tourist attractions and stores in the shopping mall. As output, it generates information on the tourist attractions and stores found.

[1984] Step 2:

[1985] The server sends the searched tourist attraction and store information to the user terminal. The tourist attraction and store information generated in step 1 is the input. The server formats this and sends it to the user terminal in an appropriate format. The output is information data to be displayed on the user terminal.

[1986] Step 3:

[1987] The device uses an AI character to provide audio guidance in the local dialect based on the received information. The inputs include tourist attraction and store information sent from the server and the AI ​​character's audio data. The device combines this information to output audio and display on the screen. The output is audio guidance and visual guidance information presented to the user.

[1988] Step 4:

[1989] The user enters a question into the chat window on the smartphone. The input is a text question. The device sends this to the server. The output is the user question data.

[1990] Step 5:

[1991] The server analyzes the received question using natural language processing (NLP) technology and generates an appropriate answer. The input is the question data entered by the user. The server inputs this data into an NLP model and generates an answer based on the analysis results. The output is an appropriate answer text.

[1992] Step 6:

[1993] The server sends the generated answer to the user terminal. The input is the answer text generated in step 5. The server formats it and sends it to the user terminal. The output is the answer data to be displayed on the user terminal.

[1994] Step 7:

[1995] The device displays the received response in the user's chat window and responds audibly if necessary. The inputs are the response data sent from the server and the AI ​​character's voice data. The device combines these and outputs text and voice. The output is a visual and voice response to the user.

[1996] Step 8:

[1997] The user activates the smartphone camera and selects AR mode. The input is the camera image. The device captures it and sends it to the server. The output is real-time video data sent to the server.

[1998] Step 9:

[1999] The server calculates the display position of the AI ​​character based on the received camera image and location information, and sends this data to the device. The inputs are camera image data and GPS location information. The server processes this data and generates display position data for the AI ​​character. The output is sent to the device.

[2000] Step 10:

[2001] The device overlays the AI ​​character on the camera image based on the received AI character position data. The inputs are camera image data and AI character position data. The device combines these and displays them. The output is an augmented reality image from the user's perspective.

[2002] Step 11:

[2003] Users post information about new tourist spots and shops to the app. The input is text and image data. The device sends this to the server. The output is the posted data sent to the server.

[2004] Step 12:

[2005] The server stores the received posting data in a database and uses it as training data for the AI ​​model. The input includes tourist spot and store information data posted by users. The server stores this in the database and retrains the AI ​​model to improve its accuracy. As output, the new information data is stored in the database and an updated AI model is generated.

[2006] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2007] This invention is a system that compensates for the lack of tourist guides in local tourist destinations and provides a personalized tourist experience. This system combines location information acquisition, AI characters that speak the local dialect, AR display, and an emotion engine to realize a tourist guide that responds to the user's emotions. Each component and its operation are explained in detail below.

[2008] Location information acquisition means

[2009] Device:

[2010] When a user visits a tourist spot, the device uses the GPS function to obtain current location information, which is then periodically sent to the server.

[2011] Tourist attraction information search methods

[2012] server:

[2013] Based on the received location information, the server searches a database for information on nearby tourist attractions, which includes detailed descriptions of the attractions, photos, and user reviews.

[2014] Tourist attraction information provision method

[2015] server:

[2016] Tourist attraction information from search results is sent to the device.

[2017] Device:

[2018] The received information is displayed on the user interface. For example, information such as "There is a historical shrine 500 meters from your current location. It is said that good things will happen if you visit it" is provided.

[2019] Tourist guide using AI characters in local dialects

[2020] Device:

[2021] The displayed tourist spot information is then provided to the user by an AI character speaking the local dialect, for example, "If you go straight ahead, you will come across a famous local shrine."

[2022] Chat-style question and answering tool

[2023] User:

[2024] The user types a question into the device's chat window (e.g., "What are some recommended restaurants nearby?").

[2025] Device:

[2026] The entered question is sent to the server.

[2027] server:

[2028] The question content is analyzed using natural language processing (NLP) technology to generate an appropriate answer.

[2029] Device:

[2030] The received answers are displayed in a chat window, and the guide AI reads them out loud in the local dialect.

[2031] AR display guide

[2032] User:

[2033] The user selects the AR mode on the device and activates the camera.

[2034] Device:

[2035] The camera captures the video and sends it to the server in real time.

[2036] server:

[2037] Based on the location information and captured video, the display position of the AI ​​character is calculated, and the AI ​​character's 3D model data and location information are sent to the device.

[2038] Device:

[2039] Based on the received data, an AI character is superimposed on the camera image and the guide AI character provides sightseeing information.

[2040] A learning tool for user-provided information on famous places

[2041] User:

[2042] When a user finds a new tourist spot, they post a photo and description to the app (e.g., new photo spot).

[2043] Device:

[2044] The post content is sent to the server.

[2045] server:

[2046] New tourist spot information is stored in a database and used as training data for the AI ​​model.

[2047] server:

[2048] The AI ​​model is updated based on the learning results, and the latest tourist spot information is provided to other users from the next time onwards.

[2049] Response methods using emotion engines

[2050] Device:

[2051] The system analyzes the user's facial expressions and voice and uses an emotion engine to recognize their emotional state.

[2052] server:

[2053] Tourist attraction information and guidance content are adjusted based on the user's emotions recognized by the emotion engine.

[2054] Device:

[2055] It provides information based on emotions. For example, if it detects that the user is tired, it will suggest nearby rest spots or cafes.

[2056] Specific examples

[2057] As a concrete example, consider the case where a user visits a mountain in a local tourist spot. When the user climbs the mountain with their device, the device sends their current location information to the server. Based on the location information, the server obtains information about the peak of "Mount XX," a tourist attraction in the vicinity of the relevant point, and generates a guide that says, "The view from the peak is spectacular, and it is located 500 meters from here," and sends it to the device.

[2058] Based on this information, an AI guide character will provide voice guidance in the local dialect on the device. In ...

Claims

1. A means for acquiring location information; A means for searching for information on tourist attractions based on the acquired location information; means for providing information about the searched tourist attractions to the user; A means to display an AI character that speaks in the local dialect and act as a tourist guide, means for accepting a question from a user, generating an answer according to the question, and providing the answer to the user; A means for displaying an AI character in AR based on the user's location information; and a means for storing the tourist spot information provided by the user in a database and updating the AI ​​model.

2. The system according to claim 1, further comprising means for the AI ​​character, which speaks in a local dialect, to provide information independently based on personal information.

3. 2. The system according to claim 1, further comprising means for learning tourist spot information provided by a user and providing better information to other users from the next time onwards.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A