system

The system uses generative AI to create personalized, multilingual AR experiences at tourist destinations, addressing integration challenges and enhancing cultural engagement, thereby revitalizing the tourism industry.

JP2026073515APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing systems struggle to integrate entertainment closely related to regional culture and history in tourist destinations, provide consistent services to multilingual tourists, and promote interaction among users with similar interests, while facing challenges in content generation cost and operational complexity.

Method used

A system utilizing generative AI to create entertainment content tailored to tourist destinations, providing augmented reality experiences based on user location, translating content into multiple languages, and facilitating user interaction to enhance cultural understanding and community engagement.

Benefits of technology

Enriches tourist experiences by offering personalized, multilingual, and interactive AR content, deepening cultural understanding, and promoting local culture, thus revitalizing the tourism industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073515000001_ABST
    Figure 2026073515000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of automatically generating entertainment content related to tourist destinations using generative artificial intelligence, A means for delivering relevant augmented reality content to a user's device based on the user's location information, A means of translating the generated content into multiple languages ​​and adapting it to international users, A means for users to virtually experience specific celebrities or historical events through augmented reality experiences at tourist destinations, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a tourist destination, it has been difficult to provide new experiences while integrating entertainment closely related to the culture and history of the region. Also, it is required to provide consistent services to tourists who speak various languages, and further to promote communication among people who are interested in the same tourist destination. In existing systems, it is difficult to integrate these elements, and furthermore, the content generation cost and the complexity of operation have been problems.

Means for Solving the Problems

[0005] This invention provides a system that automatically generates entertainment content related to tourist destinations using generative AI technology and provides relevant augmented reality content based on the user's location information. Furthermore, it translates the generated content into multiple languages, realizing a consistent experience that can be accommodated by international users. This provides a means to enrich the experience at tourist destinations and promote interaction among users with similar interests. In addition, by collaborating with local organizations, it realizes tourist experiences that have unique characteristics specific to the region.

[0006] "Generative artificial intelligence" is a technology that automatically creates content tailored to specific purposes, particularly generating entertaining content using information related to the history and culture of tourist destinations.

[0007] "Augmented reality content" refers to content that uses technology to overlay digital data onto real-world scenery, and includes virtual guides and messages that can be experienced at specific tourist destinations based on the user's location information.

[0008] "Multilingual translation means" refers to technologies that convert generated content into multiple languages, and is a system that enables consistent information provision to international users.

[0009] A "user profile" is a collection of information based on a user's personal preferences and interests, and is used to customize individual experiences related to tourist destinations or celebrities.

[0010] "Regional culture" refers to a collection of cultural elements such as traditions, history, customs, and art in a particular region, and is an important element that constitutes a unique experience in a tourist destination. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, the terms used in the following description will be explained.

[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] This invention provides a tourism support system that combines generative artificial intelligence (AI) and augmented reality (AR) technology. The system aims to enhance the tourist experience and provide tourists with rich information about local culture and history.

[0033] Specifically, users visit tourist destinations using a device with a dedicated application installed. Through the app, users can register an account and customize their profile according to their interests. This prepares users to individually optimize their experiences at their destinations.

[0034] When a user arrives at a tourist destination, the device uses GPS to determine its current location. This location information is sent to a server, where historical and cultural data related to that location is collected. The server then utilizes generative artificial intelligence to generate entertaining AR content based on this data. This content could include, for example, a virtual guide to a famous person associated with the area, a reenactment of a historical event, or an introduction to local specialties.

[0035] The generated content is sent from the server to the user's device, which then displays the content in its location using AR technology. Users can activate their camera and visually experience the AR content in accordance with the surrounding scenery. Furthermore, this system is equipped with a multilingual translation function, so the content is instantly translated according to the user's language settings, providing a consistent service to tourists who speak different languages.

[0036] Furthermore, users can use the app's social features to chat with other users visiting the same location and share their experiences. This transforms the tourist experience from an isolated one into a community experience.

[0037] For example, when a user visits a specific shrine, the device displays AR content incorporating myths and legends associated with that location, providing an experience as if they have traveled back in time. This feature allows users to deepen their understanding of the historical background of a place and further increase their affection for that tourist destination. This system also contributes to the promotion of local culture and promotes the revitalization of the entire tourism industry.

[0038] The following describes the processing flow.

[0039] Step 1:

[0040] Users install a dedicated application on their device and register an account upon first launch. They enter the necessary information and create a profile within the app. This profile can include information about tourist destinations and celebrities of interest.

[0041] Step 2:

[0042] When a user arrives at a tourist destination and launches the app, the device uses its GPS function to obtain its current location. It then sends this location information to a server and requests historical and cultural data related to that location from the server.

[0043] Step 3:

[0044] The server uses the user's location information to search relevant databases and collect information about famous people, historical events, and cultural properties associated with that location. Using generative artificial intelligence, it generates entertainment content based on the collected information.

[0045] Step 4:

[0046] The server sends the generated content to the user's terminal. This content may include virtual guides to celebrities, reenactments of historical events, and introductions to local specialties.

[0047] Step 5:

[0048] The user's device activates its camera to display the received content in AR format, overlaying virtual content onto the real-world scenery. Through this AR content, the user can enjoy a virtual sightseeing experience.

[0049] Step 6:

[0050] The server provides multilingual support by translating content into the user's chosen language as needed and returning it to the device. This ensures that users who speak different languages ​​can enjoy a consistent travel experience.

[0051] Step 7:

[0052] Users can use the app's social features to chat with other users visiting the same tourist destinations and share their experiences in real time. This promotes communication among users.

[0053] This is the processing flow of this system, which enables us to provide users with a personalized travel experience.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] Modern tourist destinations face the challenge of visitors having difficulty gaining a deep understanding of the local culture and history. Furthermore, it is difficult for tourists speaking diverse languages ​​to obtain consistent information, making it challenging to provide experiences tailored to their individual interests. Another challenge is the lack of sufficient means for visitors to virtually experience specific people or historical events.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes means for automatically generating visual effects based on information related to tourist destinations using a generative model, means for transmitting relevant display information to user devices based on the visitor's location information, and means for translating the generated information into multiple languages ​​to accommodate users who speak different languages. This makes it possible for visitors to receive customized information in multiple languages ​​based on their individual interests and to gain a deep understanding of the local culture and history while experiencing it realistically.

[0059] A "generative model" refers to an algorithm that automatically creates visual effects based on information related to tourist destinations.

[0060] "Visitors" refer to people who visit tourist destinations and seek information about the local culture and history.

[0061] "User equipment" refers to terminal devices used by visitors, which acquire location information and display received content.

[0062] "Visual effects" refer to virtual information that is overlaid onto the real world through generated content.

[0063] "Multilingual conversion" refers to the process of converting information into multiple languages ​​so that users who speak different languages ​​can understand the content.

[0064] "Local culture" refers to historical, social, and cultural information and background related to a particular tourist destination or its surrounding area.

[0065] This invention is a system that enhances the tourist experience, utilizing generative AI models and augmented reality technology. Specifically, users install a dedicated application on their device, register an account through that application, and customize their profile. This profile customization prepares the system for providing content tailored to the user's interests.

[0066] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location and sends this information to the server. The server collects relevant data from the received location information and uses an AI model to create AR content based on that data. An example of a prompt used here would be, "Please generate content that introduces historical events related to this region in an interesting way."

[0067] The generated AR content is sent from the server to the user's device, which then overlays it onto the real-world environment. Users can visually enjoy the AR content, adapted to the surrounding scenery, using their device's camera function. The device also features a multilingual translation function, translating the content according to the user's language settings. This feature makes it possible to provide the same experience to users who speak different languages.

[0068] Furthermore, users can communicate with other tourists using the app's interaction features. This feature facilitates information sharing with other users in the same location they are visiting, allowing the visit experience to be enjoyed as a community activity.

[0069] This invention enables visitors to gain a deeper understanding of the history and culture of tourist destinations through generated content, leading to more personal and engaging experiences. This is expected to contribute to the promotion of local culture and revitalize the tourism industry.

[0070] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0071] Step 1:

[0072] Users install a dedicated application on their device, register an account, and customize their profile. The input consists of personal information and interest information provided by the user, and the output is customized profile data. This profile configuration forms the basis for determining the content the user is interested in.

[0073] Step 2:

[0074] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location. The input is location data received from the GPS sensor, and the output is latitude and longitude location information. This location information determines the user's current location.

[0075] Step 3:

[0076] The device sends location information to the server. The input is the location information sent from the device, and the output is the location information stored on the server. This allows the server to recognize the user's current location and prepare to collect relevant data.

[0077] Step 4:

[0078] The server uses location information to collect historical and cultural data related to tourist destinations from a database. The input is the location information transmitted earlier, and the output is the cultural and historical data corresponding to that location. Through data collection, information is aggregated for provision to the user.

[0079] Step 5:

[0080] The server generates AR content using an AI model based on the collected data. The input consists of collected data and a prompt, such as "Generate content introducing historical events related to this region." The output is AR content generated according to the theme. This process ensures the creation of entertaining content.

[0081] Step 6:

[0082] The server sends the generated AR content to the user's device. The input is the generated AR content, and the output is the content received by the device. At this stage, the content for the user to actually experience is transferred to the device.

[0083] Step 7:

[0084] The device overlays the received content onto the real-world environment. The input is AR content transferred from the server, and the output is an augmented reality experience tailored to the user's camera view. When the user activates the camera, the AR content is displayed appropriately, providing the user with visual enjoyment.

[0085] Step 8:

[0086] The device translates content into multiple languages ​​based on the user's language settings. Input is the language setting and the original content data, while output is the translated content. This provides a smooth experience for multilingual tourists.

[0087] Step 9:

[0088] Users utilize the app's interaction features to chat with other visitors and share information. Input consists of messages and data shared by users, while output is communication and social interaction with other users. At this stage, user interaction is facilitated.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] To enhance the visitor experience in tourist destinations, it is necessary to provide abundant information related to the local culture and history. However, providing information individually optimized according to each user's interests and language is difficult. Furthermore, to improve the purchasing experience in physical stores, products and services need to be presented in a more intuitively understandable format. Effective means to solve these challenges are needed.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes means for automatically generating entertainment information related to geographical areas using generative artificial intelligence, means for transmitting relevant augmented virtual reality information to the user's mobile terminal based on the user's geographic coordinate information, and means for visualizing historical information and product information based on specific locations in physical stores. This enables users to individually optimize their experiences at their destinations and increase their purchasing intent through the information presented in physical stores.

[0094] "Generative artificial intelligence" is a technology that collects information related to the user's interests and geographical area, and automatically generates entertainment-oriented information based on this information.

[0095] Augmented virtual reality is a technology that overlays computer-generated information onto the real world through the user's vision and other senses.

[0096] "Geographic coordinate information" refers to numerical information used to indicate a specific location, and is usually expressed in terms of latitude and longitude.

[0097] A "physical store" is a form of sales that exists in a physical location and provides goods and services in person.

[0098] "Historical information" refers to information that includes facts and stories about past events and their background.

[0099] "Product information" refers to detailed information about a product or service, including its description, price, and manufacturing background.

[0100] "Diverse languages" refers to the types of languages ​​used to accommodate users who speak different languages.

[0101] This invention primarily relies on the mutual cooperation between a server system and a mobile terminal. The server utilizes generative artificial intelligence to analyze and automatically generate entertainment information related to the user's interests and geographical location. The server also receives the user's geographic coordinate information to generate augmented virtual reality information appropriate to the user's current location and transmits it to the mobile terminal. The mobile terminal uses this received information to provide a function that visualizes historical and product information based on specific locations within a physical store.

[0102] The device displays augmented virtual reality information using a dedicated application. Specifically, a smartphone is used as the hardware, and augmented reality libraries such as ARCore or Vuforia are employed. Furthermore, a multilingual service like Google® Translate API is used for information translation. This ensures that consistent information is provided to users who speak different languages.

[0103] For example, when a user visits a tea house in Kyoto, they can use a terminal to visualize the history of tea and the origins of the products displayed before them. In this way, users can deepen their understanding of the unique culture and products of the region and increase their purchasing intent.

[0104] A concrete example of a prompt for a generative AI model is: "Create a prompt for a generative AI model that intuitively displays historical information and product descriptions based on a specific location within a physical store, such as a tea house in Kyoto." The information generated by the generative artificial intelligence according to this prompt will provide users with high entertainment and educational value.

[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0106] Step 1:

[0107] The user accesses tourist destinations by operating a mobile device with a dedicated application installed. The device uses a GPS sensor to obtain the user's geographic coordinate information. This geographic coordinate information serves as input.

[0108] Step 2:

[0109] The terminal sends the acquired geographic coordinate information to the server. The server generates a prompt message based on this information and sends it to the generating artificial intelligence. The prompt message includes the instruction, "Generate entertainment information related to the current user's location." This generates information about the specified geographic area.

[0110] Step 3:

[0111] The server receives entertainment-oriented information generated using artificial intelligence and constructs augmented virtual reality (AVR) information. This information includes historical, cultural, and regional specialty information. The generated information becomes the output.

[0112] Step 4:

[0113] The server sends the generated augmented virtual reality information to a multilingual service to translate it into the user's language and retrieve the content in the appropriate language. The translated information is then output.

[0114] Step 5:

[0115] The device receives augmented virtual reality information from the server and visualizes it using an augmented reality library. Specifically, it overlays generated historical and product information onto the scenery displayed through the device's camera. This visualized information is then displayed on the device as output.

[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0117] This invention is a system that combines generative artificial intelligence, augmented reality technology, and an emotion recognition engine to improve the user experience in tourist destinations. In this embodiment, the user installs a dedicated application on their device, creates an account, and sets up a profile. The profile can register information about tourist destinations and celebrities that reflect the user's interests.

[0118] When a user arrives at a specific tourist destination, the device uses GPS to determine its current location. This information is sent to a server, where historical and cultural data related to the location is collected. The server utilizes generative artificial intelligence to generate entertainment content based on this data. This content may include, for example, reenactments of historical events or introductions to the region's unique culture.

[0119] The generated content is translated into multiple languages ​​by a generation AI and sent from the server to the device so that international users can experience it without any problems. The device then uses AR technology to display the content overlaid on the real-world scenery. Users can enjoy an augmented reality experience by viewing the scenery through the device's camera.

[0120] A notable feature of this system is its ability to evaluate the user's emotions in real time using an emotion recognition engine. It analyzes the user's facial expressions and voice using the device's camera and microphone to determine their emotional state. This allows the server to dynamically adjust content according to the user's emotions, providing a more personalized experience.

[0121] For example, if the emotion engine detects a user's heightened interest while they are visiting a historical site, the server can further engage the user by providing relevant additional information and images. Similarly, if the user appears tired, the server can deliver content to help them relax, such as suggestions for places to unwind or simple quizzes.

[0122] Furthermore, the server collects feedback from an emotion engine and uses this to continuously improve the quality of the content. It also features a user-to-user information sharing function, allowing users to exchange opinions with other users visiting the site at the same time.

[0123] Thus, a key feature of this invention is that, by utilizing an emotion recognition engine, it enables flexible tourism experiences tailored to the individual needs of each user. This embodiment maximizes the use of local tourism resources and contributes to the revitalization of the tourism industry.

[0124] The following describes the processing flow.

[0125] Step 1:

[0126] Users install a dedicated application on their device and create an account by entering the required information on the account registration screen. Users customize their profile based on their personal preferences and set information about tourist destinations and celebrities.

[0127] Step 2:

[0128] After the user arrives at a tourist destination, the device uses its GPS function to obtain its current location. This location information is sent to a server, and cultural and historical data related to that location is requested.

[0129] Step 3:

[0130] The server searches relevant databases based on the received location information and extracts information related to tourist destinations. It also utilizes artificial intelligence to automatically generate entertainment content based on the user's interests.

[0131] Step 4:

[0132] The generated content is translated into multiple languages ​​on the server and then delivered to the user's device. This ensures that content is provided in a language tailored to the user's settings.

[0133] Step 5:

[0134] The device displays the delivered content using AR technology. Users can activate the camera and look around tourist destinations, experiencing virtual guides and historical reenactments superimposed on the real-world scenery.

[0135] Step 6:

[0136] The device uses its camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine, and evaluates the user's emotional state in real time. This emotional data is sent to a server.

[0137] Step 7:

[0138] The server dynamically adjusts content based on emotional data. For example, if a user is excited, it updates the content to provide more detailed information and deepen their interest.

[0139] Step 8:

[0140] Users can use the app's social features to share their impressions with other users visiting the same tourist destination. This feature stimulates communication among users and promotes information sharing.

[0141] Step 9:

[0142] The server tracks user sentiment data and usage patterns to evaluate the popularity and effectiveness of content. This data will be used for future content creation and system improvements.

[0143] (Example 2)

[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0145] Providing timely and engaging information to tourists is challenging. Given the diverse language and cultural backgrounds of visitors, appropriate content and translation are essential. Furthermore, there is a desire for personalized experiences tailored to individual interests and emotions, rather than standardized information. To meet these needs, establishing an effective and flexible tourism information system is a critical challenge.

[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0147] In this invention, the server includes means for automatically generating information presentations related to tourist spots using generative artificial intelligence, means for delivering relevant augmented reality information presentations to the user's device based on the user's location data, and means for recognizing the user's emotions in real time and dynamically adjusting the information presentation. This makes it possible to provide each user with a flexible and personalized tourist experience based on their interests and emotions.

[0148] "Generative artificial intelligence" is a form of artificial intelligence that has the ability to automatically generate information and content related to tourism and entertainment from data.

[0149] A "tourist spot" is a geographical or cultural place that users intend to visit.

[0150] "Information presentation" refers to digital data and content provided in a way that is easy for users to understand.

[0151] "Location data" refers to information indicating the user's current geographical location, and is obtained using technologies such as GPS.

[0152] Augmented reality is a technology that overlays digital information onto images of the real world.

[0153] "User device" refers to a device carried by the user, including smartphones and tablets.

[0154] "Recognizing emotions in real time" refers to analyzing the user's facial expressions and voice to evaluate their emotional state at that moment.

[0155] "Translating into another language" refers to the process of converting information or content written in one language into a different language.

[0156] A "local organization" is an organization or group that represents the interests or culture of a specific region.

[0157] An "administrative organization" is an organization that is responsible for providing administrative services through public institutions and local governments.

[0158] "Personalized" refers to experiences and services that are tailored based on the individual user's characteristics and interests.

[0159] This invention is a system that combines various technologies to provide personalized tourism experiences to individual users in tourist destinations. Key components include generative AI, augmented reality technology, and an emotion recognition engine.

[0160] When a user visits a tourist destination, they access the system through an application installed on their device. First, the user installs the application on their smart device, creates an account, and sets up a profile based on their interests. This profile allows them to register information about tourist destinations they want to visit and people they are interested in.

[0161] The device uses GPS functionality to accurately determine the user's current location. This location information is transmitted to a server, which then uses this information to collect historical and cultural data related to the tourist destination from the internet and internal databases. The server processes the collected data using artificial intelligence to generate entertainment content tailored to the user. Specific examples of generated content include virtual experiences of historical events and introductions to local culture.

[0162] The generated content is translated into multiple languages ​​using generation AI and provided in a format suitable for international users. The server sends this translated content to the device, which then uses augmented reality technology to overlay it onto the real-world scenery. Through the device, users can have a culturally immersive experience.

[0163] Furthermore, the device uses its camera and microphone to analyze the user's facial expressions and voice in real time. An emotion recognition engine uses this to determine the user's emotional state, and this information is sent to the server. The server dynamically adjusts the content generated based on the emotion data, providing information that is appropriate for the user's emotions. For example, if the user shows particular interest, additional information is presented; if they show signs of fatigue, information that allows them to rest or lighter content is presented.

[0164] Examples of specific prompts to input into a generative AI model:

[0165] "To enhance the user experience at tourist destinations, generate individually tailored AR content using real-time sentiment data. Provide specific entertainment experiences based on the user's interests and emotional state."

[0166] In this way, this invention makes it possible to realize flexible and immersive tourism experiences that are tailored to individual needs.

[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0168] Step 1:

[0169] The device uses GPS functionality to obtain the user's current location. This location data is collected and sent to the server. The input is location information from the device's GPS sensor, and the output is the location data sent to the server. The device periodically updates this location information to maintain accurate positioning.

[0170] Step 2:

[0171] The server collects historical and cultural data for the relevant tourist spot from a database based on the received location data. The input is the location data received from the terminal, and the output is a set of related data. The server quickly searches for information associated with each tourist spot and performs operations to analyze the collected data.

[0172] Step 3:

[0173] The server uses generative artificial intelligence to generate entertainment content based on collected data. The input is collected historical and cultural data, and the output is the generated entertainment content. The server sends prompts to the generative AI model, which processes the data and generates stories and visual effects.

[0174] Step 4:

[0175] The server translates the generated content into multiple languages. The input is the content generated in step 3, and the output is the translated multilingual content. The generation AI performs translation into different languages, making it accessible to international users.

[0176] Step 5:

[0177] The server sends the translated content to the terminal. The input is multilingual content, and the output is the content delivered to the terminal. The server uses a secure communication protocol to send the content to the terminal.

[0178] Step 6:

[0179] The device overlays the received content onto the real world using augmented reality technology. The input is translated content received from the server, and the output is a visual experience utilizing augmented reality. The device integrates the content into the real world's image through its camera and displays it to the user.

[0180] Step 7:

[0181] The device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine. The input is real-time data from the user, and the output is analyzed emotion data. Based on the collected data, the device performs specific actions to determine the emotional state.

[0182] Step 8:

[0183] The server dynamically adjusts the generated content based on emotional data. The input is emotional data received from the terminal, and the output is the adjusted entertainment content. Based on the user's emotions, the server performs specific actions to process the data in order to provide additional information or new content.

[0184] (Application Example 2)

[0185] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0186] Problems faced by factory workers include decreased work efficiency and ensuring safety. In particular, if the work environment is monotonous and fatigue easily accumulates, work performance can be significantly affected. Furthermore, while support tailored to the different needs of each worker is necessary, providing this support in real time is difficult.

[0187] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for automatically generating operation support information related to the work environment using generative artificial intelligence, means for analyzing the emotional state of the worker and dynamically distributing support information for improving work efficiency, and means for visualizing the generated support information and displaying it as augmented reality on the user terminal. This enables immediate and effective support tailored to the individual needs of the worker.

[0188] "Generative artificial intelligence" is an AI technology that automatically generates new information and content based on given data and conditions.

[0189] "Work environment" refers to the physical and electronic environment within a factory or production facility where workers perform their duties.

[0190] "Operation support information" refers to information such as guidelines and advice provided to improve work efficiency and safety.

[0191] "Emotional state" refers to the psychological and emotional condition of a worker, and the state in which individual support is needed based on this condition.

[0192] "Dynamic delivery" means providing information appropriately in real time or as needed.

[0193] "Visualization" is a technique that displays abstract information in a way that is easy for users to understand, such as through images or videos.

[0194] A "user terminal" refers to an electronic device capable of displaying or inputting information, and includes smart glasses and personal digital assistants (PDAs).

[0195] Augmented reality is a technology that overlays digital information onto the real world.

[0196] To implement this invention, a system is constructed in which a server and a user's mobile device work in cooperation. The server is equipped with generative artificial intelligence and automatically generates operation support information related to the work environment. This information includes important advice and guidelines for the tasks performed by the worker. The generative artificial intelligence processes pre-collected data and provides optimal support information according to the actual work situation.

[0197] User terminals, particularly devices such as smart glasses, monitor the worker's emotional state in real time. Specifically, they utilize an emotion recognition engine to analyze the user's facial expressions and voice. This information is transmitted to a server, and support information is dynamically updated and delivered based on the worker's emotional state.

[0198] This generated support information is visualized as augmented reality through the user's terminal. Information can be overlaid in real time onto the work environment via the terminal's display, enabling workers to perform their tasks efficiently. AR technology intuitively displays information and guidelines within the worker's field of view, enhancing work safety.

[0199] For example, if the emotion recognition engine assesses that a worker is lacking concentration, the server will provide information on relaxation techniques and ways to simplify work procedures. In this way, workers can receive appropriate support when they need it.

[0200] An example of a prompt for a generative AI model is: "Create a guide to maximize work efficiency in the factory based on emotional states. Provide solutions for when workers are not concentrating."

[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0202] Step 1:

[0203] The server collects relevant data from the work environment. Input data is required, including information about the work location and the worker's job duties. Generative artificial intelligence analyzes this data to generate optimal support information. This support information includes safety guides for work procedures and suggestions for efficiency improvements.

[0204] Step 2:

[0205] The terminal analyzes the worker's emotional state in real time. The input consists of facial expression data and voice data collected by the terminal's camera and microphone. An emotion recognition engine processes this information to determine the worker's psychological state. The output provides the worker's stress and fatigue levels.

[0206] Step 3:

[0207] The server dynamically updates work support information based on the acquired emotional state. The input is the emotional state data obtained in step 2. The generating AI model receives prompts and creates support content tailored to the worker's state. The output is customized support information according to the situation.

[0208] Step 4:

[0209] The terminal visualizes support information received from the server as augmented reality. The input is the support information sent from the server. AR technology overlays the information onto the real-world work environment. As output, intuitively understandable operating guidelines appear in the worker's field of vision.

[0210] Step 5:

[0211] The user proceeds with the task based on the support information presented through the terminal. Input consists of the displayed operational support information. The user utilizes the presented information to perform the task safely and efficiently. This is expected to improve work performance.

[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0215] [Second Embodiment]

[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0228] This invention provides a tourism support system that combines generative artificial intelligence (AI) and augmented reality (AR) technology. The system aims to enhance the tourist experience and provide tourists with rich information about local culture and history.

[0229] Specifically, users visit tourist destinations using a device with a dedicated application installed. Through the app, users can register an account and customize their profile according to their interests. This prepares users to individually optimize their experiences at their destinations.

[0230] When a user arrives at a tourist destination, the device uses GPS to determine its current location. This location information is sent to a server, where historical and cultural data related to that location is collected. The server then utilizes generative artificial intelligence to generate entertaining AR content based on this data. This content could include, for example, a virtual guide to a famous person associated with the area, a reenactment of a historical event, or an introduction to local specialties.

[0231] The generated content is sent from the server to the user's device, which then displays the content in its location using AR technology. Users can activate their camera and visually experience the AR content in accordance with the surrounding scenery. Furthermore, this system is equipped with a multilingual translation function, so the content is instantly translated according to the user's language settings, providing a consistent service to tourists who speak different languages.

[0232] Furthermore, users can use the app's social features to chat with other users visiting the same location and share their experiences. This transforms the tourist experience from an isolated one into a community experience.

[0233] For example, when a user visits a specific shrine, the device displays AR content incorporating myths and legends associated with that location, providing an experience as if they have traveled back in time. This feature allows users to deepen their understanding of the historical background of a place and further increase their affection for that tourist destination. This system also contributes to the promotion of local culture and promotes the revitalization of the entire tourism industry.

[0234] The following describes the processing flow.

[0235] Step 1:

[0236] Users install a dedicated application on their device and register an account upon first launch. They enter the necessary information and create a profile within the app. This profile can include information about tourist destinations and celebrities of interest.

[0237] Step 2:

[0238] When a user arrives at a tourist destination and launches the app, the device uses its GPS function to obtain its current location. It then sends this location information to a server and requests historical and cultural data related to that location from the server.

[0239] Step 3:

[0240] The server uses the user's location information to search relevant databases and collect information about famous people, historical events, and cultural properties associated with that location. Using generative artificial intelligence, it generates entertainment content based on the collected information.

[0241] Step 4:

[0242] The server sends the generated content to the user's terminal. This content may include virtual guides to celebrities, reenactments of historical events, and introductions to local specialties.

[0243] Step 5:

[0244] The user's device activates its camera to display the received content in AR format, overlaying virtual content onto the real-world scenery. Through this AR content, the user can enjoy a virtual sightseeing experience.

[0245] Step 6:

[0246] The server provides multilingual support by translating content into the user's chosen language as needed and returning it to the device. This ensures that users who speak different languages ​​can enjoy a consistent travel experience.

[0247] Step 7:

[0248] Users can use the app's social features to chat with other users visiting the same tourist destinations and share their experiences in real time. This promotes communication among users.

[0249] This is the processing flow of this system, which enables us to provide users with a personalized travel experience.

[0250] (Example 1)

[0251] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0252] Modern tourist destinations face the challenge of visitors having difficulty gaining a deep understanding of the local culture and history. Furthermore, it is difficult for tourists speaking diverse languages ​​to obtain consistent information, making it challenging to provide experiences tailored to their individual interests. Another challenge is the lack of sufficient means for visitors to virtually experience specific people or historical events.

[0253] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0254] In this invention, the server includes means for automatically generating visual effects based on information related to tourist destinations using a generative model, means for transmitting relevant display information to user devices based on the visitor's location information, and means for translating the generated information into multiple languages ​​to accommodate users who speak different languages. This makes it possible for visitors to receive customized information in multiple languages ​​based on their individual interests and to gain a deep understanding of the local culture and history while experiencing it realistically.

[0255] A "generative model" refers to an algorithm that automatically creates visual effects based on information related to tourist destinations.

[0256] "Visitors" refer to people who visit tourist destinations and seek information about the local culture and history.

[0257] "User equipment" refers to terminal devices used by visitors, which acquire location information and display received content.

[0258] "Visual effects" refer to virtual information that is overlaid onto the real world through generated content.

[0259] "Multilingual conversion" refers to the process of converting information into multiple languages ​​so that users who speak different languages ​​can understand the content.

[0260] "Local culture" refers to historical, social, and cultural information and background related to a particular tourist destination or its surrounding area.

[0261] This invention is a system that enhances the tourist experience, utilizing generative AI models and augmented reality technology. Specifically, users install a dedicated application on their device, register an account through that application, and customize their profile. This profile customization prepares the system for providing content tailored to the user's interests.

[0262] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location and sends this information to the server. The server collects relevant data from the received location information and uses an AI model to create AR content based on that data. An example of a prompt used here would be, "Please generate content that introduces historical events related to this region in an interesting way."

[0263] The generated AR content is sent from the server to the user's device, which then overlays it onto the real-world environment. Users can visually enjoy the AR content, adapted to the surrounding scenery, using their device's camera function. The device also features a multilingual translation function, translating the content according to the user's language settings. This feature makes it possible to provide the same experience to users who speak different languages.

[0264] Furthermore, users can communicate with other tourists using the app's interaction features. This feature facilitates information sharing with other users in the same location they are visiting, allowing the visit experience to be enjoyed as a community activity.

[0265] This invention enables visitors to gain a deeper understanding of the history and culture of tourist destinations through generated content, leading to more personal and engaging experiences. This is expected to contribute to the promotion of local culture and revitalize the tourism industry.

[0266] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0267] Step 1:

[0268] Users install a dedicated application on their device, register an account, and customize their profile. The input consists of personal information and interest information provided by the user, and the output is customized profile data. This profile configuration forms the basis for determining the content the user is interested in.

[0269] Step 2:

[0270] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location. The input is location data received from the GPS sensor, and the output is latitude and longitude location information. This location information determines the user's current location.

[0271] Step 3:

[0272] The device sends location information to the server. The input is the location information sent from the device, and the output is the location information stored on the server. This allows the server to recognize the user's current location and prepare to collect relevant data.

[0273] Step 4:

[0274] The server uses location information to collect historical and cultural data related to tourist destinations from a database. The input is the location information transmitted earlier, and the output is the cultural and historical data corresponding to that location. Through data collection, information is aggregated for provision to the user.

[0275] Step 5:

[0276] The server generates AR content using an AI model based on the collected data. The input consists of collected data and a prompt, such as "Generate content introducing historical events related to this region." The output is AR content generated according to the theme. This process ensures the creation of entertaining content.

[0277] Step 6:

[0278] The server sends the generated AR content to the user's device. The input is the generated AR content, and the output is the content received by the device. At this stage, the content for the user to actually experience is transferred to the device.

[0279] Step 7:

[0280] The terminal overlays and displays the received content in the real environment. The input is the AR content transferred from the server, and the output is an augmented reality experience adapted to the user's camera view. When the user activates the camera, the AR content is properly displayed, providing visual enjoyment to the user.

[0281] Step 8:

[0282] The terminal translates the content into multiple languages based on the user's language settings. The input is the language setting for display and the original content data, and the output is the translated content. This provides a smooth experience for tourists who speak multiple languages.

[0283] Step 9:

[0284] The user uses the communication function within the app to chat with other visitors and share information. The input is the messages sent by the user and the data shared, and the output is communication and social interaction with other users. At this stage, communication between users is promoted.

[0285] (Application Example 1)

[0286] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0287] In order to enhance the experience of visitors at tourist destinations, it is required to provide rich information related to the local culture and history. However, it is difficult to provide information optimized individually according to each user's interests and language. Also, in order to improve the purchasing experience at physical stores, it is necessary to present products and services in a more intuitive form. An effective means to solve these problems is required.

[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0289] In this invention, the server includes means for automatically generating entertainment information related to geographical areas using generative artificial intelligence, means for transmitting relevant augmented virtual reality information to the user's mobile terminal based on the user's geographic coordinate information, and means for visualizing historical information and product information based on specific locations in physical stores. This enables users to individually optimize their experiences at their destinations and increase their purchasing intent through the information presented in physical stores.

[0290] "Generative artificial intelligence" is a technology that collects information related to the user's interests and geographical area, and automatically generates entertainment-oriented information based on this information.

[0291] Augmented virtual reality is a technology that overlays computer-generated information onto the real world through the user's vision and other senses.

[0292] "Geographic coordinate information" refers to numerical information used to indicate a specific location, and is usually expressed in terms of latitude and longitude.

[0293] A "physical store" is a form of sales that exists in a physical location and provides goods and services in person.

[0294] "Historical information" refers to information that includes facts and stories about past events and their background.

[0295] "Product information" refers to detailed information about a product or service, including its description, price, and manufacturing background.

[0296] "Diverse languages" refers to the types of languages ​​used to accommodate users who speak different languages.

[0297] This invention primarily relies on the mutual cooperation between a server system and a mobile terminal. The server utilizes generative artificial intelligence to analyze and automatically generate entertainment information related to the user's interests and geographical location. The server also receives the user's geographic coordinate information to generate augmented virtual reality information appropriate to the user's current location and transmits it to the mobile terminal. The mobile terminal uses this received information to provide a function that visualizes historical and product information based on specific locations within a physical store.

[0298] The device displays augmented virtual reality information using a dedicated application. A smartphone is used as the specific hardware, and augmented reality libraries such as ARCore or Vuforia are employed. Furthermore, a multilingual service like the Google Translate API is used for information translation. This ensures that consistent information is provided to users who speak different languages.

[0299] For example, when a user visits a tea house in Kyoto, they can use a terminal to visualize the history of tea and the origins of the products displayed before them. In this way, users can deepen their understanding of the unique culture and products of the region and increase their purchasing intent.

[0300] A concrete example of a prompt for a generative AI model is: "Create a prompt for a generative AI model that intuitively displays historical information and product descriptions based on a specific location within a physical store, such as a tea house in Kyoto." The information generated by the generative artificial intelligence according to this prompt will provide users with high entertainment and educational value.

[0301] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0302] Step 1:

[0303] The user operates a mobile terminal installed with a dedicated application to access a tourist destination. The terminal uses a GPS sensor to obtain the user's geographical coordinate information. This geographical coordinate information serves as the input.

[0304] Step 2:

[0305] The terminal transmits the obtained geographical coordinate information to the server. The server generates a prompt sentence based on this information and transmits it to the generative artificial intelligence. The prompt sentence includes an instruction such as "Please generate information on entertainment uses related to the current user's location." As a result, information regarding the specified geographical area is generated.

[0306] Step 3:

[0307] The server receives the generated entertainment use information using the generative artificial intelligence and constructs augmented virtual reality information. This information includes information on history, culture, and local specialties. The generated information serves as the output.

[0308] Step 4:

[0309] The server translates the generated augmented virtual reality information into the user's language. To do this, it transmits the information to a multilingual support service and obtains the content in the appropriate language. The translated information also serves as the output.

[0310] Step 5:

[0311] The terminal incorporates the augmented virtual reality information received from the server and visualizes the information using an augmented reality library. Specifically, it superimposes the generated historical information and product information on the landscape displayed through the terminal's camera. This visualized information is the output displayed on the terminal.

[0312] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.

[0313] This invention is a system that combines generative artificial intelligence, augmented reality technology, and an emotion recognition engine to improve the user experience in tourist destinations. In this embodiment, the user installs a dedicated application on their device, creates an account, and sets up a profile. The profile can register information about tourist destinations and celebrities that reflect the user's interests.

[0314] When a user arrives at a specific tourist destination, the device uses GPS to determine its current location. This information is sent to a server, where historical and cultural data related to the location is collected. The server utilizes generative artificial intelligence to generate entertainment content based on this data. This content may include, for example, reenactments of historical events or introductions to the region's unique culture.

[0315] The generated content is translated into multiple languages ​​by a generation AI and sent from the server to the device so that international users can experience it without any problems. The device then uses AR technology to display the content overlaid on the real-world scenery. Users can enjoy an augmented reality experience by viewing the scenery through the device's camera.

[0316] A notable feature of this system is its ability to evaluate the user's emotions in real time using an emotion recognition engine. It analyzes the user's facial expressions and voice using the device's camera and microphone to determine their emotional state. This allows the server to dynamically adjust content according to the user's emotions, providing a more personalized experience.

[0317] For example, if the emotion engine detects a user's heightened interest while they are visiting a historical site, the server can further engage the user by providing relevant additional information and images. Similarly, if the user appears tired, the server can deliver content to help them relax, such as suggestions for places to unwind or simple quizzes.

[0318] Furthermore, the server collects feedback from an emotion engine and uses this to continuously improve the quality of the content. It also features a user-to-user information sharing function, allowing users to exchange opinions with other users visiting the site at the same time.

[0319] Thus, a key feature of this invention is that, by utilizing an emotion recognition engine, it enables flexible tourism experiences tailored to the individual needs of each user. This embodiment maximizes the use of local tourism resources and contributes to the revitalization of the tourism industry.

[0320] The following describes the processing flow.

[0321] Step 1:

[0322] Users install a dedicated application on their device and create an account by entering the required information on the account registration screen. Users customize their profile based on their personal preferences and set information about tourist destinations and celebrities.

[0323] Step 2:

[0324] After the user arrives at a tourist destination, the device uses its GPS function to obtain its current location. This location information is sent to a server, and cultural and historical data related to that location is requested.

[0325] Step 3:

[0326] The server searches relevant databases based on the received location information and extracts information related to tourist destinations. It also utilizes artificial intelligence to automatically generate entertainment content based on the user's interests.

[0327] Step 4:

[0328] The generated content is translated into multiple languages ​​on the server and then delivered to the user's device. This ensures that content is provided in a language tailored to the user's settings.

[0329] Step 5:

[0330] The device displays the delivered content using AR technology. Users can activate the camera and look around tourist destinations, experiencing virtual guides and historical reenactments superimposed on the real-world scenery.

[0331] Step 6:

[0332] The device uses its camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine, and evaluates the user's emotional state in real time. This emotional data is sent to a server.

[0333] Step 7:

[0334] The server dynamically adjusts content based on emotional data. For example, if a user is excited, it updates the content to provide more detailed information and deepen their interest.

[0335] Step 8:

[0336] Users can use the app's social features to share their impressions with other users visiting the same tourist destination. This feature stimulates communication among users and promotes information sharing.

[0337] Step 9:

[0338] The server tracks user sentiment data and usage patterns to evaluate the popularity and effectiveness of content. This data will be used for future content creation and system improvements.

[0339] (Example 2)

[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0341] Providing timely and engaging information to tourists is challenging. Given the diverse language and cultural backgrounds of visitors, appropriate content and translation are essential. Furthermore, there is a desire for personalized experiences tailored to individual interests and emotions, rather than standardized information. To meet these needs, establishing an effective and flexible tourism information system is a critical challenge.

[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0343] In this invention, the server includes means for automatically generating information presentations related to tourist spots using generative artificial intelligence, means for delivering relevant augmented reality information presentations to the user's device based on the user's location data, and means for recognizing the user's emotions in real time and dynamically adjusting the information presentation. This makes it possible to provide each user with a flexible and personalized tourist experience based on their interests and emotions.

[0344] "Generative artificial intelligence" is a form of artificial intelligence that has the ability to automatically generate information and content related to tourism and entertainment from data.

[0345] A "tourist spot" is a geographical or cultural place that users intend to visit.

[0346] "Information presentation" refers to digital data and content provided in a way that is easy for users to understand.

[0347] "Location data" refers to information indicating the user's current geographical location, and is obtained using technologies such as GPS.

[0348] Augmented reality is a technology that overlays digital information onto images of the real world.

[0349] "User device" refers to a device carried by the user, including smartphones and tablets.

[0350] "Recognizing emotions in real time" refers to analyzing the user's facial expressions and voice to evaluate their emotional state at that moment.

[0351] "Translating into another language" refers to the process of converting information or content written in one language into a different language.

[0352] A "local organization" is an organization or group that represents the interests or culture of a specific region.

[0353] An "administrative organization" is an organization that is responsible for providing administrative services through public institutions and local governments.

[0354] "Personalized" refers to experiences and services that are tailored based on the individual user's characteristics and interests.

[0355] This invention is a system that combines various technologies to provide personalized tourism experiences to individual users in tourist destinations. Key components include generative AI, augmented reality technology, and an emotion recognition engine.

[0356] When a user visits a tourist destination, they access the system through an application installed on their device. First, the user installs the application on their smart device, creates an account, and sets up a profile based on their interests. This profile allows them to register information about tourist destinations they want to visit and people they are interested in.

[0357] The device uses GPS functionality to accurately determine the user's current location. This location information is transmitted to a server, which then uses this information to collect historical and cultural data related to the tourist destination from the internet and internal databases. The server processes the collected data using artificial intelligence to generate entertainment content tailored to the user. Specific examples of generated content include virtual experiences of historical events and introductions to local culture.

[0358] The generated content is translated into multiple languages ​​using generation AI and provided in a format suitable for international users. The server sends this translated content to the device, which then uses augmented reality technology to overlay it onto the real-world scenery. Through the device, users can have a culturally immersive experience.

[0359] Furthermore, the device uses its camera and microphone to analyze the user's facial expressions and voice in real time. An emotion recognition engine uses this to determine the user's emotional state, and this information is sent to the server. The server dynamically adjusts the content generated based on the emotion data, providing information that is appropriate for the user's emotions. For example, if the user shows particular interest, additional information is presented; if they show signs of fatigue, information that allows them to rest or lighter content is presented.

[0360] Examples of specific prompts to input into a generative AI model:

[0361] "To enhance the user experience at tourist destinations, generate individually tailored AR content using real-time sentiment data. Provide specific entertainment experiences based on the user's interests and emotional state."

[0362] In this way, this invention makes it possible to realize flexible and immersive tourism experiences that are tailored to individual needs.

[0363] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0364] Step 1:

[0365] The device uses GPS functionality to obtain the user's current location. This location data is collected and sent to the server. The input is location information from the device's GPS sensor, and the output is the location data sent to the server. The device periodically updates this location information to maintain accurate positioning.

[0366] Step 2:

[0367] The server collects historical and cultural data for the relevant tourist spot from a database based on the received location data. The input is the location data received from the terminal, and the output is a set of related data. The server quickly searches for information associated with each tourist spot and performs operations to analyze the collected data.

[0368] Step 3:

[0369] The server uses generative artificial intelligence to generate entertainment content based on collected data. The input is collected historical and cultural data, and the output is the generated entertainment content. The server sends prompts to the generative AI model, which processes the data and generates stories and visual effects.

[0370] Step 4:

[0371] The server translates the generated content into multiple languages. The input is the content generated in step 3, and the output is the translated multilingual content. The generation AI performs translation into different languages, making it accessible to international users.

[0372] Step 5:

[0373] The server sends the translated content to the terminal. The input is multilingual content, and the output is the content delivered to the terminal. The server uses a secure communication protocol to send the content to the terminal.

[0374] Step 6:

[0375] The device overlays the received content onto the real world using augmented reality technology. The input is translated content received from the server, and the output is a visual experience utilizing augmented reality. The device integrates the content into the real world's image through its camera and displays it to the user.

[0376] Step 7:

[0377] The device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine. The input is real-time data from the user, and the output is analyzed emotion data. Based on the collected data, the device performs specific actions to determine the emotional state.

[0378] Step 8:

[0379] The server dynamically adjusts the generated content based on emotional data. The input is emotional data received from the terminal, and the output is the adjusted entertainment content. Based on the user's emotions, the server performs specific actions to process the data in order to provide additional information or new content.

[0380] (Application Example 2)

[0381] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0382] Problems faced by factory workers include decreased work efficiency and ensuring safety. In particular, if the work environment is monotonous and fatigue easily accumulates, work performance can be significantly affected. Furthermore, while support tailored to the different needs of each worker is necessary, providing this support in real time is difficult.

[0383] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for automatically generating operation support information related to the work environment using generative artificial intelligence, means for analyzing the emotional state of the worker and dynamically distributing support information for improving work efficiency, and means for visualizing the generated support information and displaying it as augmented reality on the user terminal. This enables immediate and effective support tailored to the individual needs of the worker.

[0384] "Generative artificial intelligence" is an AI technology that automatically generates new information and content based on given data and conditions.

[0385] "Work environment" refers to the physical and electronic environment within a factory or production facility where workers perform their duties.

[0386] "Operation support information" refers to information such as guidelines and advice provided to improve work efficiency and safety.

[0387] "Emotional state" refers to the psychological and emotional condition of a worker, and the state in which individual support is needed based on this condition.

[0388] "Dynamic delivery" means providing information appropriately in real time or as needed.

[0389] "Visualization" is a technique that displays abstract information in a way that is easy for users to understand, such as through images or videos.

[0390] A "user terminal" refers to an electronic device capable of displaying or inputting information, and includes smart glasses and personal digital assistants (PDAs).

[0391] Augmented reality is a technology that overlays digital information onto the real world.

[0392] To implement this invention, a system is constructed in which a server and a user's mobile device work in cooperation. The server is equipped with generative artificial intelligence and automatically generates operation support information related to the work environment. This information includes important advice and guidelines for the tasks performed by the worker. The generative artificial intelligence processes pre-collected data and provides optimal support information according to the actual work situation.

[0393] User terminals, particularly devices such as smart glasses, monitor the worker's emotional state in real time. Specifically, they utilize an emotion recognition engine to analyze the user's facial expressions and voice. This information is transmitted to a server, and support information is dynamically updated and delivered based on the worker's emotional state.

[0394] This generated support information is visualized as augmented reality through the user's terminal. Information can be overlaid in real time onto the work environment via the terminal's display, enabling workers to perform their tasks efficiently. AR technology intuitively displays information and guidelines within the worker's field of view, enhancing work safety.

[0395] For example, if the emotion recognition engine assesses that a worker is lacking concentration, the server will provide information on relaxation techniques and ways to simplify work procedures. In this way, workers can receive appropriate support when they need it.

[0396] An example of a prompt for a generative AI model is: "Create a guide to maximize work efficiency in the factory based on emotional states. Provide solutions for when workers are not concentrating."

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] The server collects relevant data from the work environment. Input data is required, including information about the work location and the worker's job duties. Generative artificial intelligence analyzes this data to generate optimal support information. This support information includes safety guides for work procedures and suggestions for efficiency improvements.

[0400] Step 2:

[0401] The terminal analyzes the worker's emotional state in real time. The input consists of facial expression data and voice data collected by the terminal's camera and microphone. An emotion recognition engine processes this information to determine the worker's psychological state. The output provides the worker's stress and fatigue levels.

[0402] Step 3:

[0403] The server dynamically updates work support information based on the acquired emotional state. The input is the emotional state data obtained in step 2. The generating AI model receives prompts and creates support content tailored to the worker's state. The output is customized support information according to the situation.

[0404] Step 4:

[0405] The terminal visualizes support information received from the server as augmented reality. The input is the support information sent from the server. AR technology overlays the information onto the real-world work environment. As output, intuitively understandable operating guidelines appear in the worker's field of vision.

[0406] Step 5:

[0407] The user proceeds with the task based on the support information presented through the terminal. Input consists of the displayed operational support information. The user utilizes the presented information to perform the task safely and efficiently. This is expected to improve work performance.

[0408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0409] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0411] [Third Embodiment]

[0412] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0413] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0415] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0419] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0422] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0423] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0424] This invention provides a tourism support system that combines generative artificial intelligence (AI) and augmented reality (AR) technology. The system aims to enhance the tourist experience and provide tourists with rich information about local culture and history.

[0425] Specifically, users visit tourist destinations using a device with a dedicated application installed. Through the app, users can register an account and customize their profile according to their interests. This prepares users to individually optimize their experiences at their destinations.

[0426] When a user arrives at a tourist destination, the device uses GPS to determine its current location. This location information is sent to a server, where historical and cultural data related to that location is collected. The server then utilizes generative artificial intelligence to generate entertaining AR content based on this data. This content could include, for example, a virtual guide to a famous person associated with the area, a reenactment of a historical event, or an introduction to local specialties.

[0427] The generated content is sent from the server to the user's device, which then displays the content in its location using AR technology. Users can activate their camera and visually experience the AR content in accordance with the surrounding scenery. Furthermore, this system is equipped with a multilingual translation function, so the content is instantly translated according to the user's language settings, providing a consistent service to tourists who speak different languages.

[0428] Furthermore, users can use the app's social features to chat with other users visiting the same location and share their experiences. This transforms the tourist experience from an isolated one into a community experience.

[0429] For example, when a user visits a specific shrine, the device displays AR content incorporating myths and legends associated with that location, providing an experience as if they have traveled back in time. This feature allows users to deepen their understanding of the historical background of a place and further increase their affection for that tourist destination. This system also contributes to the promotion of local culture and promotes the revitalization of the entire tourism industry.

[0430] The following describes the processing flow.

[0431] Step 1:

[0432] Users install a dedicated application on their device and register an account upon first launch. They enter the necessary information and create a profile within the app. This profile can include information about tourist destinations and celebrities of interest.

[0433] Step 2:

[0434] When a user arrives at a tourist destination and launches the app, the device uses its GPS function to obtain its current location. It then sends this location information to a server and requests historical and cultural data related to that location from the server.

[0435] Step 3:

[0436] The server uses the user's location information to search relevant databases and collect information about famous people, historical events, and cultural properties associated with that location. Using generative artificial intelligence, it generates entertainment content based on the collected information.

[0437] Step 4:

[0438] The server sends the generated content to the user's terminal. This content may include virtual guides to celebrities, reenactments of historical events, and introductions to local specialties.

[0439] Step 5:

[0440] The user's device activates its camera to display the received content in AR format, overlaying virtual content onto the real-world scenery. Through this AR content, the user can enjoy a virtual sightseeing experience.

[0441] Step 6:

[0442] The server provides multilingual support by translating content into the user's chosen language as needed and returning it to the device. This ensures that users who speak different languages ​​can enjoy a consistent travel experience.

[0443] Step 7:

[0444] Users can use the app's social features to chat with other users visiting the same tourist destinations and share their experiences in real time. This promotes communication among users.

[0445] This is the processing flow of this system, which enables us to provide users with a personalized travel experience.

[0446] (Example 1)

[0447] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0448] Modern tourist destinations face the challenge of visitors having difficulty gaining a deep understanding of the local culture and history. Furthermore, it is difficult for tourists speaking diverse languages ​​to obtain consistent information, making it challenging to provide experiences tailored to their individual interests. Another challenge is the lack of sufficient means for visitors to virtually experience specific people or historical events.

[0449] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0450] In this invention, the server includes means for automatically generating visual effects based on information related to tourist destinations using a generative model, means for transmitting relevant display information to user devices based on the visitor's location information, and means for translating the generated information into multiple languages ​​to accommodate users who speak different languages. This makes it possible for visitors to receive customized information in multiple languages ​​based on their individual interests and to gain a deep understanding of the local culture and history while experiencing it realistically.

[0451] A "generative model" refers to an algorithm that automatically creates visual effects based on information related to tourist destinations.

[0452] "Visitors" refer to people who visit tourist destinations and seek information about the local culture and history.

[0453] "User equipment" refers to terminal devices used by visitors, which acquire location information and display received content.

[0454] "Visual effects" refer to virtual information that is overlaid onto the real world through generated content.

[0455] "Multilingual conversion" refers to the process of converting information into multiple languages ​​so that users who speak different languages ​​can understand the content.

[0456] "Local culture" refers to historical, social, and cultural information and background related to a particular tourist destination or its surrounding area.

[0457] This invention is a system that enhances the tourist experience, utilizing generative AI models and augmented reality technology. Specifically, users install a dedicated application on their device, register an account through that application, and customize their profile. This profile customization prepares the system for providing content tailored to the user's interests.

[0458] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location and sends this information to the server. The server collects relevant data from the received location information and uses an AI model to create AR content based on that data. An example of a prompt used here would be, "Please generate content that introduces historical events related to this region in an interesting way."

[0459] The generated AR content is sent from the server to the user's device, which then overlays it onto the real-world environment. Users can visually enjoy the AR content, adapted to the surrounding scenery, using their device's camera function. The device also features a multilingual translation function, translating the content according to the user's language settings. This feature makes it possible to provide the same experience to users who speak different languages.

[0460] Furthermore, users can communicate with other tourists using the app's interaction features. This feature facilitates information sharing with other users in the same location they are visiting, allowing the visit experience to be enjoyed as a community activity.

[0461] This invention enables visitors to gain a deeper understanding of the history and culture of tourist destinations through generated content, leading to more personal and engaging experiences. This is expected to contribute to the promotion of local culture and revitalize the tourism industry.

[0462] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0463] Step 1:

[0464] Users install a dedicated application on their device, register an account, and customize their profile. The input consists of personal information and interest information provided by the user, and the output is customized profile data. This profile configuration forms the basis for determining the content the user is interested in.

[0465] Step 2:

[0466] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location. The input is location data received from the GPS sensor, and the output is latitude and longitude location information. This location information determines the user's current location.

[0467] Step 3:

[0468] The device sends location information to the server. The input is the location information sent from the device, and the output is the location information stored on the server. This allows the server to recognize the user's current location and prepare to collect relevant data.

[0469] Step 4:

[0470] The server uses location information to collect historical and cultural data related to tourist destinations from a database. The input is the location information transmitted earlier, and the output is the cultural and historical data corresponding to that location. Through data collection, information is aggregated for provision to the user.

[0471] Step 5:

[0472] The server generates AR content using an AI model based on the collected data. The input consists of collected data and a prompt, such as "Generate content introducing historical events related to this region." The output is AR content generated according to the theme. This process ensures the creation of entertaining content.

[0473] Step 6:

[0474] The server sends the generated AR content to the user's device. The input is the generated AR content, and the output is the content received by the device. At this stage, the content for the user to actually experience is transferred to the device.

[0475] Step 7:

[0476] The device overlays the received content onto the real-world environment. The input is AR content transferred from the server, and the output is an augmented reality experience tailored to the user's camera view. When the user activates the camera, the AR content is displayed appropriately, providing the user with visual enjoyment.

[0477] Step 8:

[0478] The device translates content into multiple languages ​​based on the user's language settings. Input is the language setting and the original content data, while output is the translated content. This provides a smooth experience for multilingual tourists.

[0479] Step 9:

[0480] Users utilize the app's interaction features to chat with other visitors and share information. Input consists of messages and data shared by users, while output is communication and social interaction with other users. At this stage, user interaction is facilitated.

[0481] (Application Example 1)

[0482] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0483] To enhance the visitor experience in tourist destinations, it is necessary to provide abundant information related to the local culture and history. However, providing information individually optimized according to each user's interests and language is difficult. Furthermore, to improve the purchasing experience in physical stores, products and services need to be presented in a more intuitively understandable format. Effective means to solve these challenges are needed.

[0484] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0485] In this invention, the server includes means for automatically generating entertainment information related to geographical areas using generative artificial intelligence, means for transmitting relevant augmented virtual reality information to the user's mobile terminal based on the user's geographic coordinate information, and means for visualizing historical information and product information based on specific locations in physical stores. This enables users to individually optimize their experiences at their destinations and increase their purchasing intent through the information presented in physical stores.

[0486] "Generative artificial intelligence" is a technology that collects information related to the user's interests and geographical area, and automatically generates entertainment-oriented information based on this information.

[0487] Augmented virtual reality is a technology that overlays computer-generated information onto the real world through the user's vision and other senses.

[0488] "Geographic coordinate information" refers to numerical information used to indicate a specific location, and is usually expressed in terms of latitude and longitude.

[0489] A "physical store" is a form of sales that exists in a physical location and provides goods and services in person.

[0490] "Historical information" refers to information that includes facts and stories about past events and their background.

[0491] "Product information" refers to detailed information about a product or service, including its description, price, and manufacturing background.

[0492] "Diverse languages" refers to the types of languages ​​used to accommodate users who speak different languages.

[0493] This invention primarily relies on the mutual cooperation between a server system and a mobile terminal. The server utilizes generative artificial intelligence to analyze and automatically generate entertainment information related to the user's interests and geographical location. The server also receives the user's geographic coordinate information to generate augmented virtual reality information appropriate to the user's current location and transmits it to the mobile terminal. The mobile terminal uses this received information to provide a function that visualizes historical and product information based on specific locations within a physical store.

[0494] The device displays augmented virtual reality information using a dedicated application. A smartphone is used as the specific hardware, and augmented reality libraries such as ARCore or Vuforia are employed. Furthermore, a multilingual service like the Google Translate API is used for information translation. This ensures that consistent information is provided to users who speak different languages.

[0495] For example, when a user visits a tea house in Kyoto, they can use a terminal to visualize the history of tea and the origins of the products displayed before them. In this way, users can deepen their understanding of the unique culture and products of the region and increase their purchasing intent.

[0496] A concrete example of a prompt for a generative AI model is: "Create a prompt for a generative AI model that intuitively displays historical information and product descriptions based on a specific location within a physical store, such as a tea house in Kyoto." The information generated by the generative artificial intelligence according to this prompt will provide users with high entertainment and educational value.

[0497] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0498] Step 1:

[0499] The user accesses tourist destinations by operating a mobile device with a dedicated application installed. The device uses a GPS sensor to obtain the user's geographic coordinate information. This geographic coordinate information serves as input.

[0500] Step 2:

[0501] The terminal sends the acquired geographic coordinate information to the server. The server generates a prompt message based on this information and sends it to the generating artificial intelligence. The prompt message includes the instruction, "Generate entertainment information related to the current user's location." This generates information about the specified geographic area.

[0502] Step 3:

[0503] The server receives entertainment-oriented information generated using artificial intelligence and constructs augmented virtual reality (AVR) information. This information includes historical, cultural, and regional specialty information. The generated information becomes the output.

[0504] Step 4:

[0505] The server sends the generated augmented virtual reality information to a multilingual service to translate it into the user's language and retrieve the content in the appropriate language. The translated information is then output.

[0506] Step 5:

[0507] The device receives augmented virtual reality information from the server and visualizes it using an augmented reality library. Specifically, it overlays generated historical and product information onto the scenery displayed through the device's camera. This visualized information is then displayed on the device as output.

[0508] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0509] This invention is a system that combines generative artificial intelligence, augmented reality technology, and an emotion recognition engine to improve the user experience in tourist destinations. In this embodiment, the user installs a dedicated application on their device, creates an account, and sets up a profile. The profile can register information about tourist destinations and celebrities that reflect the user's interests.

[0510] When a user arrives at a specific tourist destination, the device uses GPS to determine its current location. This information is sent to a server, where historical and cultural data related to the location is collected. The server utilizes generative artificial intelligence to generate entertainment content based on this data. This content may include, for example, reenactments of historical events or introductions to the region's unique culture.

[0511] The generated content is translated into multiple languages ​​by a generation AI and sent from the server to the device so that international users can experience it without any problems. The device then uses AR technology to display the content overlaid on the real-world scenery. Users can enjoy an augmented reality experience by viewing the scenery through the device's camera.

[0512] A notable feature of this system is its ability to evaluate the user's emotions in real time using an emotion recognition engine. It analyzes the user's facial expressions and voice using the device's camera and microphone to determine their emotional state. This allows the server to dynamically adjust content according to the user's emotions, providing a more personalized experience.

[0513] For example, if the emotion engine detects a user's heightened interest while they are visiting a historical site, the server can further engage the user by providing relevant additional information and images. Similarly, if the user appears tired, the server can deliver content to help them relax, such as suggestions for places to unwind or simple quizzes.

[0514] Furthermore, the server collects feedback from an emotion engine and uses this to continuously improve the quality of the content. It also features a user-to-user information sharing function, allowing users to exchange opinions with other users visiting the site at the same time.

[0515] Thus, a key feature of this invention is that, by utilizing an emotion recognition engine, it enables flexible tourism experiences tailored to the individual needs of each user. This embodiment maximizes the use of local tourism resources and contributes to the revitalization of the tourism industry.

[0516] The following describes the processing flow.

[0517] Step 1:

[0518] Users install a dedicated application on their device and create an account by entering the required information on the account registration screen. Users customize their profile based on their personal preferences and set information about tourist destinations and celebrities.

[0519] Step 2:

[0520] After the user arrives at a tourist destination, the device uses its GPS function to obtain its current location. This location information is sent to a server, and cultural and historical data related to that location is requested.

[0521] Step 3:

[0522] The server searches relevant databases based on the received location information and extracts information related to tourist destinations. It also utilizes artificial intelligence to automatically generate entertainment content based on the user's interests.

[0523] Step 4:

[0524] The generated content is translated into multiple languages ​​on the server and then delivered to the user's device. This ensures that content is provided in a language tailored to the user's settings.

[0525] Step 5:

[0526] The device displays the delivered content using AR technology. Users can activate the camera and look around tourist destinations, experiencing virtual guides and historical reenactments superimposed on the real-world scenery.

[0527] Step 6:

[0528] The device uses its camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine, and evaluates the user's emotional state in real time. This emotional data is sent to a server.

[0529] Step 7:

[0530] The server dynamically adjusts content based on emotional data. For example, if a user is excited, it updates the content to provide more detailed information and deepen their interest.

[0531] Step 8:

[0532] Users can use the app's social features to share their impressions with other users visiting the same tourist destination. This feature stimulates communication among users and promotes information sharing.

[0533] Step 9:

[0534] The server tracks user sentiment data and usage patterns to evaluate the popularity and effectiveness of content. This data will be used for future content creation and system improvements.

[0535] (Example 2)

[0536] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0537] Providing timely and engaging information to tourists is challenging. Given the diverse language and cultural backgrounds of visitors, appropriate content and translation are essential. Furthermore, there is a desire for personalized experiences tailored to individual interests and emotions, rather than standardized information. To meet these needs, establishing an effective and flexible tourism information system is a critical challenge.

[0538] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0539] In this invention, the server includes means for automatically generating information presentations related to tourist spots using generative artificial intelligence, means for delivering relevant augmented reality information presentations to the user's device based on the user's location data, and means for recognizing the user's emotions in real time and dynamically adjusting the information presentation. This makes it possible to provide each user with a flexible and personalized tourist experience based on their interests and emotions.

[0540] "Generative artificial intelligence" is a form of artificial intelligence that has the ability to automatically generate information and content related to tourism and entertainment from data.

[0541] A "tourist spot" is a geographical or cultural place that users intend to visit.

[0542] "Information presentation" refers to digital data and content provided in a way that is easy for users to understand.

[0543] "Location data" refers to information indicating the user's current geographical location, and is obtained using technologies such as GPS.

[0544] Augmented reality is a technology that overlays digital information onto images of the real world.

[0545] "User device" refers to a device carried by the user, including smartphones and tablets.

[0546] "Recognizing emotions in real time" refers to analyzing the user's facial expressions and voice to evaluate their emotional state at that moment.

[0547] "Translating into another language" refers to the process of converting information or content written in one language into a different language.

[0548] A "local organization" is an organization or group that represents the interests or culture of a specific region.

[0549] An "administrative organization" is an organization that is responsible for providing administrative services through public institutions and local governments.

[0550] "Personalized" refers to experiences and services that are tailored based on the individual user's characteristics and interests.

[0551] This invention is a system that combines various technologies to provide personalized tourism experiences to individual users in tourist destinations. Key components include generative AI, augmented reality technology, and an emotion recognition engine.

[0552] When a user visits a tourist destination, they access the system through an application installed on their device. First, the user installs the application on their smart device, creates an account, and sets up a profile based on their interests. This profile allows them to register information about tourist destinations they want to visit and people they are interested in.

[0553] The device uses GPS functionality to accurately determine the user's current location. This location information is transmitted to a server, which then uses this information to collect historical and cultural data related to the tourist destination from the internet and internal databases. The server processes the collected data using artificial intelligence to generate entertainment content tailored to the user. Specific examples of generated content include virtual experiences of historical events and introductions to local culture.

[0554] The generated content is translated into multiple languages ​​using generation AI and provided in a format suitable for international users. The server sends this translated content to the device, which then uses augmented reality technology to overlay it onto the real-world scenery. Through the device, users can have a culturally immersive experience.

[0555] Furthermore, the device uses its camera and microphone to analyze the user's facial expressions and voice in real time. An emotion recognition engine uses this to determine the user's emotional state, and this information is sent to the server. The server dynamically adjusts the content generated based on the emotion data, providing information that is appropriate for the user's emotions. For example, if the user shows particular interest, additional information is presented; if they show signs of fatigue, information that allows them to rest or lighter content is presented.

[0556] Examples of specific prompts to input into a generative AI model:

[0557] "To enhance the user experience at tourist destinations, generate individually tailored AR content using real-time sentiment data. Provide specific entertainment experiences based on the user's interests and emotional state."

[0558] In this way, this invention makes it possible to realize flexible and immersive tourism experiences that are tailored to individual needs.

[0559] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0560] Step 1:

[0561] The device uses GPS functionality to obtain the user's current location. This location data is collected and sent to the server. The input is location information from the device's GPS sensor, and the output is the location data sent to the server. The device periodically updates this location information to maintain accurate positioning.

[0562] Step 2:

[0563] The server collects historical and cultural data for the relevant tourist spot from a database based on the received location data. The input is the location data received from the terminal, and the output is a set of related data. The server quickly searches for information associated with each tourist spot and performs operations to analyze the collected data.

[0564] Step 3:

[0565] The server uses generative artificial intelligence to generate entertainment content based on collected data. The input is collected historical and cultural data, and the output is the generated entertainment content. The server sends prompts to the generative AI model, which processes the data and generates stories and visual effects.

[0566] Step 4:

[0567] The server translates the generated content into multiple languages. The input is the content generated in step 3, and the output is the translated multilingual content. The generation AI performs translation into different languages, making it accessible to international users.

[0568] Step 5:

[0569] The server sends the translated content to the terminal. The input is multilingual content, and the output is the content delivered to the terminal. The server uses a secure communication protocol to send the content to the terminal.

[0570] Step 6:

[0571] The device overlays the received content onto the real world using augmented reality technology. The input is translated content received from the server, and the output is a visual experience utilizing augmented reality. The device integrates the content into the real world's image through its camera and displays it to the user.

[0572] Step 7:

[0573] The device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine. The input is real-time data from the user, and the output is analyzed emotion data. Based on the collected data, the device performs specific actions to determine the emotional state.

[0574] Step 8:

[0575] The server dynamically adjusts the generated content based on emotional data. The input is emotional data received from the terminal, and the output is the adjusted entertainment content. Based on the user's emotions, the server performs specific actions to process the data in order to provide additional information or new content.

[0576] (Application Example 2)

[0577] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0578] Problems faced by factory workers include decreased work efficiency and ensuring safety. In particular, if the work environment is monotonous and fatigue easily accumulates, work performance can be significantly affected. Furthermore, while support tailored to the different needs of each worker is necessary, providing this support in real time is difficult.

[0579] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for automatically generating operation support information related to the work environment using generative artificial intelligence, means for analyzing the emotional state of the worker and dynamically distributing support information for improving work efficiency, and means for visualizing the generated support information and displaying it as augmented reality on the user terminal. This enables immediate and effective support tailored to the individual needs of the worker.

[0580] "Generative artificial intelligence" is an AI technology that automatically generates new information and content based on given data and conditions.

[0581] "Work environment" refers to the physical and electronic environment within a factory or production facility where workers perform their duties.

[0582] "Operation support information" refers to information such as guidelines and advice provided to improve work efficiency and safety.

[0583] "Emotional state" refers to the psychological and emotional condition of a worker, and the state in which individual support is needed based on this condition.

[0584] "Dynamic delivery" means providing information appropriately in real time or as needed.

[0585] "Visualization" is a technique that displays abstract information in a way that is easy for users to understand, such as through images or videos.

[0586] A "user terminal" refers to an electronic device capable of displaying or inputting information, and includes smart glasses and personal digital assistants (PDAs).

[0587] Augmented reality is a technology that overlays digital information onto the real world.

[0588] To implement this invention, a system is constructed in which a server and a user's mobile device work in cooperation. The server is equipped with generative artificial intelligence and automatically generates operation support information related to the work environment. This information includes important advice and guidelines for the tasks performed by the worker. The generative artificial intelligence processes pre-collected data and provides optimal support information according to the actual work situation.

[0589] User terminals, particularly devices such as smart glasses, monitor the worker's emotional state in real time. Specifically, they utilize an emotion recognition engine to analyze the user's facial expressions and voice. This information is transmitted to a server, and support information is dynamically updated and delivered based on the worker's emotional state.

[0590] This generated support information is visualized as augmented reality through the user's terminal. Information can be overlaid in real time onto the work environment via the terminal's display, enabling workers to perform their tasks efficiently. AR technology intuitively displays information and guidelines within the worker's field of view, enhancing work safety.

[0591] For example, if the emotion recognition engine assesses that a worker is lacking concentration, the server will provide information on relaxation techniques and ways to simplify work procedures. In this way, workers can receive appropriate support when they need it.

[0592] An example of a prompt for a generative AI model is: "Create a guide to maximize work efficiency in the factory based on emotional states. Provide solutions for when workers are not concentrating."

[0593] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0594] Step 1:

[0595] The server collects relevant data from the work environment. Input data is required, including information about the work location and the worker's job duties. Generative artificial intelligence analyzes this data to generate optimal support information. This support information includes safety guides for work procedures and suggestions for efficiency improvements.

[0596] Step 2:

[0597] The terminal analyzes the worker's emotional state in real time. The input consists of facial expression data and voice data collected by the terminal's camera and microphone. An emotion recognition engine processes this information to determine the worker's psychological state. The output provides the worker's stress and fatigue levels.

[0598] Step 3:

[0599] The server dynamically updates work support information based on the acquired emotional state. The input is the emotional state data obtained in step 2. The generating AI model receives prompts and creates support content tailored to the worker's state. The output is customized support information according to the situation.

[0600] Step 4:

[0601] The terminal visualizes support information received from the server as augmented reality. The input is the support information sent from the server. AR technology overlays the information onto the real-world work environment. As output, intuitively understandable operating guidelines appear in the worker's field of vision.

[0602] Step 5:

[0603] The user proceeds with the task based on the support information presented through the terminal. Input consists of the displayed operational support information. The user utilizes the presented information to perform the task safely and efficiently. This is expected to improve work performance.

[0604] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0605] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0606] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0607] [Fourth Embodiment]

[0608] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0609] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0610] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0611] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0612] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0613] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0614] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0615] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0616] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0617] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0618] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0619] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0620] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0621] This invention provides a tourism support system that combines generative artificial intelligence (AI) and augmented reality (AR) technology. The system aims to enhance the tourist experience and provide tourists with rich information about local culture and history.

[0622] Specifically, users visit tourist destinations using a device with a dedicated application installed. Through the app, users can register an account and customize their profile according to their interests. This prepares users to individually optimize their experiences at their destinations.

[0623] When a user arrives at a tourist destination, the device uses GPS to determine its current location. This location information is sent to a server, where historical and cultural data related to that location is collected. The server then utilizes generative artificial intelligence to generate entertaining AR content based on this data. This content could include, for example, a virtual guide to a famous person associated with the area, a reenactment of a historical event, or an introduction to local specialties.

[0624] The generated content is sent from the server to the user's device, which then displays the content in its location using AR technology. Users can activate their camera and visually experience the AR content in accordance with the surrounding scenery. Furthermore, this system is equipped with a multilingual translation function, so the content is instantly translated according to the user's language settings, providing a consistent service to tourists who speak different languages.

[0625] Furthermore, users can use the app's social features to chat with other users visiting the same location and share their experiences. This transforms the tourist experience from an isolated one into a community experience.

[0626] For example, when a user visits a specific shrine, the device displays AR content incorporating myths and legends associated with that location, providing an experience as if they have traveled back in time. This feature allows users to deepen their understanding of the historical background of a place and further increase their affection for that tourist destination. This system also contributes to the promotion of local culture and promotes the revitalization of the entire tourism industry.

[0627] The following describes the processing flow.

[0628] Step 1:

[0629] Users install a dedicated application on their device and register an account upon first launch. They enter the necessary information and create a profile within the app. This profile can include information about tourist destinations and celebrities of interest.

[0630] Step 2:

[0631] When a user arrives at a tourist destination and launches the app, the device uses its GPS function to obtain its current location. It then sends this location information to a server and requests historical and cultural data related to that location from the server.

[0632] Step 3:

[0633] The server uses the user's location information to search relevant databases and collect information about famous people, historical events, and cultural properties associated with that location. Using generative artificial intelligence, it generates entertainment content based on the collected information.

[0634] Step 4:

[0635] The server sends the generated content to the user's terminal. This content may include virtual guides to celebrities, reenactments of historical events, and introductions to local specialties.

[0636] Step 5:

[0637] The user's device activates its camera to display the received content in AR format, overlaying virtual content onto the real-world scenery. Through this AR content, the user can enjoy a virtual sightseeing experience.

[0638] Step 6:

[0639] The server provides multilingual support by translating content into the user's chosen language as needed and returning it to the device. This ensures that users who speak different languages ​​can enjoy a consistent travel experience.

[0640] Step 7:

[0641] Users can use the app's social features to chat with other users visiting the same tourist destinations and share their experiences in real time. This promotes communication among users.

[0642] This is the processing flow of this system, which enables us to provide users with a personalized travel experience.

[0643] (Example 1)

[0644] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0645] Modern tourist destinations face the challenge of visitors having difficulty gaining a deep understanding of the local culture and history. Furthermore, it is difficult for tourists speaking diverse languages ​​to obtain consistent information, making it challenging to provide experiences tailored to their individual interests. Another challenge is the lack of sufficient means for visitors to virtually experience specific people or historical events.

[0646] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0647] In this invention, the server includes means for automatically generating visual effects based on information related to tourist destinations using a generative model, means for transmitting relevant display information to user devices based on the visitor's location information, and means for translating the generated information into multiple languages ​​to accommodate users who speak different languages. This makes it possible for visitors to receive customized information in multiple languages ​​based on their individual interests and to gain a deep understanding of the local culture and history while experiencing it realistically.

[0648] A "generative model" refers to an algorithm that automatically creates visual effects based on information related to tourist destinations.

[0649] "Visitors" refer to people who visit tourist destinations and seek information about the local culture and history.

[0650] "User equipment" refers to terminal devices used by visitors, which acquire location information and display received content.

[0651] "Visual effects" refer to virtual information that is overlaid onto the real world through generated content.

[0652] "Multilingual conversion" refers to the process of converting information into multiple languages ​​so that users who speak different languages ​​can understand the content.

[0653] "Local culture" refers to historical, social, and cultural information and background related to a particular tourist destination or its surrounding area.

[0654] This invention is a system that enhances the tourist experience, utilizing generative AI models and augmented reality technology. Specifically, users install a dedicated application on their device, register an account through that application, and customize their profile. This profile customization prepares the system for providing content tailored to the user's interests.

[0655] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location and sends this information to the server. The server collects relevant data from the received location information and uses an AI model to create AR content based on that data. An example of a prompt used here would be, "Please generate content that introduces historical events related to this region in an interesting way."

[0656] The generated AR content is sent from the server to the user's device, which then overlays it onto the real-world environment. Users can visually enjoy the AR content, adapted to the surrounding scenery, using their device's camera function. The device also features a multilingual translation function, translating the content according to the user's language settings. This feature makes it possible to provide the same experience to users who speak different languages.

[0657] Furthermore, users can communicate with other tourists using the app's interaction features. This feature facilitates information sharing with other users in the same location they are visiting, allowing the visit experience to be enjoyed as a community activity.

[0658] This invention enables visitors to gain a deeper understanding of the history and culture of tourist destinations through generated content, leading to more personal and engaging experiences. This is expected to contribute to the promotion of local culture and revitalize the tourism industry.

[0659] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0660] Step 1:

[0661] Users install a dedicated application on their device, register an account, and customize their profile. The input consists of personal information and interest information provided by the user, and the output is customized profile data. This profile configuration forms the basis for determining the content the user is interested in.

[0662] Step 2:

[0663] When a user arrives at a tourist destination, the device uses its built-in GPS function to determine its current location. The input is location data received from the GPS sensor, and the output is latitude and longitude location information. This location information determines the user's current location.

[0664] Step 3:

[0665] The device sends location information to the server. The input is the location information sent from the device, and the output is the location information stored on the server. This allows the server to recognize the user's current location and prepare to collect relevant data.

[0666] Step 4:

[0667] The server uses location information to collect historical and cultural data related to tourist destinations from a database. The input is the location information transmitted earlier, and the output is the cultural and historical data corresponding to that location. Through data collection, information is aggregated for provision to the user.

[0668] Step 5:

[0669] The server generates AR content using an AI model based on the collected data. The input consists of collected data and a prompt, such as "Generate content introducing historical events related to this region." The output is AR content generated according to the theme. This process ensures the creation of entertaining content.

[0670] Step 6:

[0671] The server sends the generated AR content to the user's device. The input is the generated AR content, and the output is the content received by the device. At this stage, the content for the user to actually experience is transferred to the device.

[0672] Step 7:

[0673] The device overlays the received content onto the real-world environment. The input is AR content transferred from the server, and the output is an augmented reality experience tailored to the user's camera view. When the user activates the camera, the AR content is displayed appropriately, providing the user with visual enjoyment.

[0674] Step 8:

[0675] The device translates content into multiple languages ​​based on the user's language settings. Input is the language setting and the original content data, while output is the translated content. This provides a smooth experience for multilingual tourists.

[0676] Step 9:

[0677] Users utilize the app's interaction features to chat with other visitors and share information. Input consists of messages and data shared by users, while output is communication and social interaction with other users. At this stage, user interaction is facilitated.

[0678] (Application Example 1)

[0679] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0680] To enhance the visitor experience in tourist destinations, it is necessary to provide abundant information related to the local culture and history. However, providing information individually optimized according to each user's interests and language is difficult. Furthermore, to improve the purchasing experience in physical stores, products and services need to be presented in a more intuitively understandable format. Effective means to solve these challenges are needed.

[0681] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0682] In this invention, the server includes means for automatically generating entertainment information related to geographical areas using generative artificial intelligence, means for transmitting relevant augmented virtual reality information to the user's mobile terminal based on the user's geographic coordinate information, and means for visualizing historical information and product information based on specific locations in physical stores. This enables users to individually optimize their experiences at their destinations and increase their purchasing intent through the information presented in physical stores.

[0683] "Generative artificial intelligence" is a technology that collects information related to the user's interests and geographical area, and automatically generates entertainment-oriented information based on this information.

[0684] Augmented virtual reality is a technology that overlays computer-generated information onto the real world through the user's vision and other senses.

[0685] "Geographic coordinate information" refers to numerical information used to indicate a specific location, and is usually expressed in terms of latitude and longitude.

[0686] A "physical store" is a form of sales that exists in a physical location and provides goods and services in person.

[0687] "Historical information" refers to information that includes facts and stories about past events and their background.

[0688] "Product information" refers to detailed information about a product or service, including its description, price, and manufacturing background.

[0689] "Diverse languages" refers to the types of languages ​​used to accommodate users who speak different languages.

[0690] This invention primarily relies on the mutual cooperation between a server system and a mobile terminal. The server utilizes generative artificial intelligence to analyze and automatically generate entertainment information related to the user's interests and geographical location. The server also receives the user's geographic coordinate information to generate augmented virtual reality information appropriate to the user's current location and transmits it to the mobile terminal. The mobile terminal uses this received information to provide a function that visualizes historical and product information based on specific locations within a physical store.

[0691] The device displays augmented virtual reality information using a dedicated application. A smartphone is used as the specific hardware, and augmented reality libraries such as ARCore or Vuforia are employed. Furthermore, a multilingual service like the Google Translate API is used for information translation. This ensures that consistent information is provided to users who speak different languages.

[0692] For example, when a user visits a tea house in Kyoto, they can use a terminal to visualize the history of tea and the origins of the products displayed before them. In this way, users can deepen their understanding of the unique culture and products of the region and increase their purchasing intent.

[0693] A concrete example of a prompt for a generative AI model is: "Create a prompt for a generative AI model that intuitively displays historical information and product descriptions based on a specific location within a physical store, such as a tea house in Kyoto." The information generated by the generative artificial intelligence according to this prompt will provide users with high entertainment and educational value.

[0694] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0695] Step 1:

[0696] The user accesses tourist destinations by operating a mobile device with a dedicated application installed. The device uses a GPS sensor to obtain the user's geographic coordinate information. This geographic coordinate information serves as input.

[0697] Step 2:

[0698] The terminal sends the acquired geographic coordinate information to the server. The server generates a prompt message based on this information and sends it to the generating artificial intelligence. The prompt message includes the instruction, "Generate entertainment information related to the current user's location." This generates information about the specified geographic area.

[0699] Step 3:

[0700] The server receives entertainment-oriented information generated using artificial intelligence and constructs augmented virtual reality (AVR) information. This information includes historical, cultural, and regional specialty information. The generated information becomes the output.

[0701] Step 4:

[0702] The server sends the generated augmented virtual reality information to a multilingual service to translate it into the user's language and retrieve the content in the appropriate language. The translated information is then output.

[0703] Step 5:

[0704] The device receives augmented virtual reality information from the server and visualizes it using an augmented reality library. Specifically, it overlays generated historical and product information onto the scenery displayed through the device's camera. This visualized information is then displayed on the device as output.

[0705] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0706] This invention is a system that combines generative artificial intelligence, augmented reality technology, and an emotion recognition engine to improve the user experience in tourist destinations. In this embodiment, the user installs a dedicated application on their device, creates an account, and sets up a profile. The profile can register information about tourist destinations and celebrities that reflect the user's interests.

[0707] When a user arrives at a specific tourist destination, the device uses GPS to determine its current location. This information is sent to a server, where historical and cultural data related to the location is collected. The server utilizes generative artificial intelligence to generate entertainment content based on this data. This content may include, for example, reenactments of historical events or introductions to the region's unique culture.

[0708] The generated content is translated into multiple languages ​​by a generation AI and sent from the server to the device so that international users can experience it without any problems. The device then uses AR technology to display the content overlaid on the real-world scenery. Users can enjoy an augmented reality experience by viewing the scenery through the device's camera.

[0709] A notable feature of this system is its ability to evaluate the user's emotions in real time using an emotion recognition engine. It analyzes the user's facial expressions and voice using the device's camera and microphone to determine their emotional state. This allows the server to dynamically adjust content according to the user's emotions, providing a more personalized experience.

[0710] For example, if the emotion engine detects a user's heightened interest while they are visiting a historical site, the server can further engage the user by providing relevant additional information and images. Similarly, if the user appears tired, the server can deliver content to help them relax, such as suggestions for places to unwind or simple quizzes.

[0711] Furthermore, the server collects feedback from an emotion engine and uses this to continuously improve the quality of the content. It also features a user-to-user information sharing function, allowing users to exchange opinions with other users visiting the site at the same time.

[0712] Thus, a key feature of this invention is that, by utilizing an emotion recognition engine, it enables flexible tourism experiences tailored to the individual needs of each user. This embodiment maximizes the use of local tourism resources and contributes to the revitalization of the tourism industry.

[0713] The following describes the processing flow.

[0714] Step 1:

[0715] Users install a dedicated application on their device and create an account by entering the required information on the account registration screen. Users customize their profile based on their personal preferences and set information about tourist destinations and celebrities.

[0716] Step 2:

[0717] After the user arrives at a tourist destination, the device uses its GPS function to obtain its current location. This location information is sent to a server, and cultural and historical data related to that location is requested.

[0718] Step 3:

[0719] The server searches relevant databases based on the received location information and extracts information related to tourist destinations. It also utilizes artificial intelligence to automatically generate entertainment content based on the user's interests.

[0720] Step 4:

[0721] The generated content is translated into multiple languages ​​on the server and then delivered to the user's device. This ensures that content is provided in a language tailored to the user's settings.

[0722] Step 5:

[0723] The device displays the delivered content using AR technology. Users can activate the camera and look around tourist destinations, experiencing virtual guides and historical reenactments superimposed on the real-world scenery.

[0724] Step 6:

[0725] The device uses its camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine, and evaluates the user's emotional state in real time. This emotional data is sent to a server.

[0726] Step 7:

[0727] The server dynamically adjusts content based on emotional data. For example, if a user is excited, it updates the content to provide more detailed information and deepen their interest.

[0728] Step 8:

[0729] Users can use the app's social features to share their impressions with other users visiting the same tourist destination. This feature stimulates communication among users and promotes information sharing.

[0730] Step 9:

[0731] The server tracks user sentiment data and usage patterns to evaluate the popularity and effectiveness of content. This data will be used for future content creation and system improvements.

[0732] (Example 2)

[0733] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0734] Providing timely and engaging information to tourists is challenging. Given the diverse language and cultural backgrounds of visitors, appropriate content and translation are essential. Furthermore, there is a desire for personalized experiences tailored to individual interests and emotions, rather than standardized information. To meet these needs, establishing an effective and flexible tourism information system is a critical challenge.

[0735] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0736] In this invention, the server includes means for automatically generating information presentations related to tourist spots using generative artificial intelligence, means for delivering relevant augmented reality information presentations to the user's device based on the user's location data, and means for recognizing the user's emotions in real time and dynamically adjusting the information presentation. This makes it possible to provide each user with a flexible and personalized tourist experience based on their interests and emotions.

[0737] "Generative artificial intelligence" is a form of artificial intelligence that has the ability to automatically generate information and content related to tourism and entertainment from data.

[0738] A "tourist spot" is a geographical or cultural place that users intend to visit.

[0739] "Information presentation" refers to digital data and content provided in a way that is easy for users to understand.

[0740] "Location data" refers to information indicating the user's current geographical location, and is obtained using technologies such as GPS.

[0741] Augmented reality is a technology that overlays digital information onto images of the real world.

[0742] "User device" refers to a device carried by the user, including smartphones and tablets.

[0743] "Recognizing emotions in real time" refers to analyzing the user's facial expressions and voice to evaluate their emotional state at that moment.

[0744] "Translating into another language" refers to the process of converting information or content written in one language into a different language.

[0745] A "local organization" is an organization or group that represents the interests or culture of a specific region.

[0746] An "administrative organization" is an organization that is responsible for providing administrative services through public institutions and local governments.

[0747] "Personalized" refers to experiences and services that are tailored based on the individual user's characteristics and interests.

[0748] This invention is a system that combines various technologies to provide personalized tourism experiences to individual users in tourist destinations. Key components include generative AI, augmented reality technology, and an emotion recognition engine.

[0749] When a user visits a tourist destination, they access the system through an application installed on their device. First, the user installs the application on their smart device, creates an account, and sets up a profile based on their interests. This profile allows them to register information about tourist destinations they want to visit and people they are interested in.

[0750] The device uses GPS functionality to accurately determine the user's current location. This location information is transmitted to a server, which then uses this information to collect historical and cultural data related to the tourist destination from the internet and internal databases. The server processes the collected data using artificial intelligence to generate entertainment content tailored to the user. Specific examples of generated content include virtual experiences of historical events and introductions to local culture.

[0751] The generated content is translated into multiple languages ​​using generation AI and provided in a format suitable for international users. The server sends this translated content to the device, which then uses augmented reality technology to overlay it onto the real-world scenery. Through the device, users can have a culturally immersive experience.

[0752] Furthermore, the device uses its camera and microphone to analyze the user's facial expressions and voice in real time. An emotion recognition engine uses this to determine the user's emotional state, and this information is sent to the server. The server dynamically adjusts the content generated based on the emotion data, providing information that is appropriate for the user's emotions. For example, if the user shows particular interest, additional information is presented; if they show signs of fatigue, information that allows them to rest or lighter content is presented.

[0753] Examples of specific prompts to input into a generative AI model:

[0754] "To enhance the user experience at tourist destinations, generate individually tailored AR content using real-time sentiment data. Provide specific entertainment experiences based on the user's interests and emotional state."

[0755] In this way, this invention makes it possible to realize flexible and immersive tourism experiences that are tailored to individual needs.

[0756] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0757] Step 1:

[0758] The device uses GPS functionality to obtain the user's current location. This location data is collected and sent to the server. The input is location information from the device's GPS sensor, and the output is the location data sent to the server. The device periodically updates this location information to maintain accurate positioning.

[0759] Step 2:

[0760] The server collects historical and cultural data for the relevant tourist spot from a database based on the received location data. The input is the location data received from the terminal, and the output is a set of related data. The server quickly searches for information associated with each tourist spot and performs operations to analyze the collected data.

[0761] Step 3:

[0762] The server uses generative artificial intelligence to generate entertainment content based on collected data. The input is collected historical and cultural data, and the output is the generated entertainment content. The server sends prompts to the generative AI model, which processes the data and generates stories and visual effects.

[0763] Step 4:

[0764] The server translates the generated content into multiple languages. The input is the content generated in step 3, and the output is the translated multilingual content. The generation AI performs translation into different languages, making it accessible to international users.

[0765] Step 5:

[0766] The server sends the translated content to the terminal. The input is multilingual content, and the output is the content delivered to the terminal. The server uses a secure communication protocol to send the content to the terminal.

[0767] Step 6:

[0768] The device overlays the received content onto the real world using augmented reality technology. The input is translated content received from the server, and the output is a visual experience utilizing augmented reality. The device integrates the content into the real world's image through its camera and displays it to the user.

[0769] Step 7:

[0770] The device uses a camera and microphone to analyze the user's facial expressions and voice using an emotion recognition engine. The input is real-time data from the user, and the output is analyzed emotion data. Based on the collected data, the device performs specific actions to determine the emotional state.

[0771] Step 8:

[0772] The server dynamically adjusts the generated content based on emotional data. The input is emotional data received from the terminal, and the output is the adjusted entertainment content. Based on the user's emotions, the server performs specific actions to process the data in order to provide additional information or new content.

[0773] (Application Example 2)

[0774] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0775] Problems faced by factory workers include decreased work efficiency and ensuring safety. In particular, if the work environment is monotonous and fatigue easily accumulates, work performance can be significantly affected. Furthermore, while support tailored to the different needs of each worker is necessary, providing this support in real time is difficult.

[0776] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for automatically generating operation support information related to the work environment using generative artificial intelligence, means for analyzing the emotional state of the worker and dynamically distributing support information for improving work efficiency, and means for visualizing the generated support information and displaying it as augmented reality on the user terminal. This enables immediate and effective support tailored to the individual needs of the worker.

[0777] "Generative artificial intelligence" is an AI technology that automatically generates new information and content based on given data and conditions.

[0778] "Work environment" refers to the physical and electronic environment within a factory or production facility where workers perform their duties.

[0779] "Operation support information" refers to information such as guidelines and advice provided to improve work efficiency and safety.

[0780] "Emotional state" refers to the psychological and emotional condition of a worker, and the state in which individual support is needed based on this condition.

[0781] "Dynamic delivery" means providing information appropriately in real time or as needed.

[0782] "Visualization" is a technique that displays abstract information in a way that is easy for users to understand, such as through images or videos.

[0783] A "user terminal" refers to an electronic device capable of displaying or inputting information, and includes smart glasses and personal digital assistants (PDAs).

[0784] Augmented reality is a technology that overlays digital information onto the real world.

[0785] To implement this invention, a system is constructed in which a server and a user's mobile device work in cooperation. The server is equipped with generative artificial intelligence and automatically generates operation support information related to the work environment. This information includes important advice and guidelines for the tasks performed by the worker. The generative artificial intelligence processes pre-collected data and provides optimal support information according to the actual work situation.

[0786] User terminals, particularly devices such as smart glasses, monitor the worker's emotional state in real time. Specifically, they utilize an emotion recognition engine to analyze the user's facial expressions and voice. This information is transmitted to a server, and support information is dynamically updated and delivered based on the worker's emotional state.

[0787] This generated support information is visualized as augmented reality through the user's terminal. Information can be overlaid in real time onto the work environment via the terminal's display, enabling workers to perform their tasks efficiently. AR technology intuitively displays information and guidelines within the worker's field of view, enhancing work safety.

[0788] For example, if the emotion recognition engine assesses that a worker is lacking concentration, the server will provide information on relaxation techniques and ways to simplify work procedures. In this way, workers can receive appropriate support when they need it.

[0789] An example of a prompt for a generative AI model is: "Create a guide to maximize work efficiency in the factory based on emotional states. Provide solutions for when workers are not concentrating."

[0790] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0791] Step 1:

[0792] The server collects relevant data from the work environment. Input data is required, including information about the work location and the worker's job duties. Generative artificial intelligence analyzes this data to generate optimal support information. This support information includes safety guides for work procedures and suggestions for efficiency improvements.

[0793] Step 2:

[0794] The terminal analyzes the worker's emotional state in real time. The input consists of facial expression data and voice data collected by the terminal's camera and microphone. An emotion recognition engine processes this information to determine the worker's psychological state. The output provides the worker's stress and fatigue levels.

[0795] Step 3:

[0796] The server dynamically updates work support information based on the acquired emotional state. The input is the emotional state data obtained in step 2. The generating AI model receives prompts and creates support content tailored to the worker's state. The output is customized support information according to the situation.

[0797] Step 4:

[0798] The terminal visualizes support information received from the server as augmented reality. The input is the support information sent from the server. AR technology overlays the information onto the real-world work environment. As output, intuitively understandable operating guidelines appear in the worker's field of vision.

[0799] Step 5:

[0800] The user proceeds with the task based on the support information presented through the terminal. Input consists of the displayed operational support information. The user utilizes the presented information to perform the task safely and efficiently. This is expected to improve work performance.

[0801] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0802] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0803] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0804] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0805] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0806] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0807] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0808] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0809] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0810] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0811] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0812] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0813] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0814] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0815] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0816] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0817] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0818] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0819] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0820] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0821] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0822] The following is further disclosed regarding the embodiments described above.

[0823] (Claim 1)

[0824] A means of automatically generating entertainment content related to tourist destinations using generative artificial intelligence,

[0825] A means for delivering relevant augmented reality content to a user's device based on the user's location information,

[0826] A means of translating the generated content into multiple languages ​​and adapting it to international users,

[0827] A means for users to virtually experience specific celebrities or historical events through augmented reality experiences at tourist destinations,

[0828] A system that includes this.

[0829] (Claim 2)

[0830] A means for users to create an account and customize their profile based on their interests in tourist destinations and celebrities,

[0831] It provides features to facilitate interaction among users, and a means to share information with other users who share the same hobbies.

[0832] The system according to claim 1, including the following:

[0833] (Claim 3)

[0834] A means of tracking and evaluating the popularity and usage of content generated using generative artificial intelligence,

[0835] In collaboration with local organizations and government agencies, we offer distinctive tourism experiences that reflect local culture.

[0836] The system according to claim 1, including the following:

[0837] "Example 1"

[0838] (Claim 1)

[0839] A method for automatically generating visual effects based on information related to tourist destinations using a generative model,

[0840] A means for transmitting relevant display information to a user's device based on the visitor's location information,

[0841] A means of converting generated information into multiple languages ​​to accommodate users who speak different languages,

[0842] A means for users to virtually experience specific famous people or historical events through a sense of reality,

[0843] A system that includes this.

[0844] (Claim 2)

[0845] The system according to claim 1, in which the user sets identification information and adjusts the profile based on regions and people of interest.

[0846] (Claim 3)

[0847] The system according to claim 1 for evaluating and tracking the usage of information created using a generative model.

[0848] "Application Example 1"

[0849] (Claim 1)

[0850] A means for automatically generating information for entertainment purposes related to geographical areas using generative artificial intelligence,

[0851] A means for transmitting relevant augmented virtual reality information to the user's mobile device based on the user's geographic coordinate information,

[0852] A means to convert the generated information into various languages ​​and make it accessible to international users,

[0853] A means for users to virtually experience specific famous people or historical events through augmented reality experiences in a geographical area,

[0854] A means of visualizing historical and product information based on specific locations in physical stores,

[0855] A system that includes this.

[0856] (Claim 2)

[0857] A means for users to create individual accounts and set personal information based on geographical areas of interest or notable figures,

[0858] It provides functions to facilitate information exchange among users, and a means of sharing information with other users who share common interests.

[0859] The system according to claim 1, including the following:

[0860] (Claim 3)

[0861] A means of tracking and evaluating the popularity and usage of information generated using generative artificial intelligence,

[0862] In cooperation with local organizations and public institutions, means of providing distinctive tourism experiences that reflect local culture,

[0863] This refers to methods aimed at increasing the user's purchasing intent through historical and product information presented within a physical store,

[0864] The system according to claim 1, including the following:

[0865] "Example 2 of combining an emotion engine"

[0866] (Claim 1)

[0867] A means for automatically generating information related to tourist destinations using generative artificial intelligence,

[0868] A means for delivering relevant augmented reality information to the user's device based on the user's location data,

[0869] A means of translating the generated information presentation into multiple languages ​​and adapting it to international users,

[0870] A means for users to virtually experience specific historical events through augmented reality experiences at tourist destinations,

[0871] A means of recognizing the user's emotions in real time and dynamically adjusting the information presented,

[0872] A system that includes this.

[0873] (Claim 2)

[0874] A means for users to create an account and customize their profile based on their interests in tourist destinations and people,

[0875] It provides features to facilitate interaction among users, and a means to share information with other users who have similar interests.

[0876] The system according to claim 1, including the following:

[0877] (Claim 3)

[0878] A means of tracking and evaluating the popularity and usage of information presented using generative artificial intelligence,

[0879] In cooperation with local organizations and administrative bodies, means of providing distinctive tourism experiences that reflect local culture,

[0880] A means of continuously improving the quality of information presentation using collected emotional data,

[0881] The system according to claim 1, including the following:

[0882] "Application example 2 when combining with an emotional engine"

[0883] (Claim 1)

[0884] A means for automatically generating operation support information related to the work environment using generative artificial intelligence,

[0885] A means of analyzing the emotional state of workers and dynamically delivering support information to improve work efficiency,

[0886] A means for visualizing the generated support information and displaying it as augmented reality on the user's terminal,

[0887] A system that includes this.

[0888] (Claim 2)

[0889] The system according to claim 1, wherein the worker creates a profile and customizes operational support information based on individual needs in the work environment.

[0890] (Claim 3)

[0891] The system according to claim 1, which evaluates the effectiveness of support information generated using generative artificial intelligence and continuously improves it. [Explanation of Symbols]

[0892] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of automatically generating entertainment content related to tourist destinations using generative artificial intelligence, A means for delivering relevant augmented reality content to a user's device based on the user's location information, A means of translating the generated content into multiple languages ​​and adapting it to international users, A means for users to virtually experience specific celebrities or historical events through augmented reality experiences at tourist destinations, A system that includes this.

2. A means for users to create an account and customize their profile based on their interests in tourist destinations and celebrities, It provides features to facilitate interaction among users, and a means to share information with other users who share the same hobbies. The system according to claim 1, including the following:

3. A means of tracking and evaluating the popularity and usage of content generated using generative artificial intelligence, In collaboration with local organizations and government agencies, we offer distinctive tourism experiences that reflect local culture. The system according to claim 1, including the following:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A