system

The system uses augmented reality to recreate animal behavior and provide information, addressing visitor disappointment by offering an engaging and educational experience at zoos.

JP2026073384APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Visitors to zoos often experience disappointment when animals are inactive, leading to a lack of engagement and understanding of animal behavior.

Method used

A system utilizing augmented reality technology to virtually reproduce animal behavior, provide characteristic information, and dynamically switch content based on user location and animal recognition, enhancing the zoo experience.

Benefits of technology

Enables visitors to have a realistic and educational experience with animals, even when they are not active, deepening their knowledge and interest.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073384000001_ABST
    Figure 2026073384000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A visual display means using augmented reality technology to virtually reproduce animal behavior, An information distribution method that provides users with characteristic information about animals, A control means for switching content based on the user's location and the results of animal recognition, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present invention is to solve the problem that when visiting a zoo, the expectations of visitors are damaged because animals are sleeping or not showing themselves. At the same time, it also aims to solve the problem that it is difficult to deepen a detailed understanding and interest in animals.

Means for Solving the Problems

[0005] To solve this problem, the present invention proposes a system comprising a visual display means using augmented reality technology to virtually reproduce animal behavior, an information distribution means that provides characteristic information about animals, and a control means that dynamically switches content based on the user's location and the results of animal recognition. As a result, even if the animals are not actually active, visitors can have a realistic experience and deepen their knowledge about animals.

[0006] "Visual display means" refers to a device or method that virtually reproduces animal behavior and displays it to a user using augmented reality technology.

[0007] "Information distribution means" refers to a device or method that provides users with information about animals in audio and text formats.

[0008] "Control means" refers to a device or system that dynamically switches the displayed content according to the user's location information and the results of animal recognition.

[0009] "Generation means" refers to a method or technique for generating virtual animals based on past animal behavior data.

[0010] "Animal behavioral data" refers to information about the activities and movements that animals have exhibited in the past. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] One embodiment of this invention is to provide an augmented reality experience system for visitors to a zoo to enhance their enjoyment of animals. The system mainly consists of a visual device (terminal) worn by the user, sensing devices installed in the zoo, and a database and processing server in the cloud.

[0033] When a user arrives at the zoo, the device first authenticates their login information. This information is sent to the server, which also retrieves data from the user's past visits. This prepares the device to provide information optimized for the visitor.

[0034] When the user approaches the animal area, the device uses GPS and Bluetooth beacons to pinpoint the user's location and confirms the presence of animals through its camera. Based on this information, the server references the animals' past movements and behavioral patterns and uses generative AI to create virtual animal movement data. The device then uses AR technology to project the virtual animal into the user's field of view based on this data. At this time, the realistic movements and behaviors of the animals are visually reproduced.

[0035] Furthermore, the server sends descriptions and learning information about the animals to the device in audio and text formats. Users can continue observing the animals while listening to or reading this information, gaining a deeper understanding of interesting characteristics and the animals themselves.

[0036] As a concrete example, users in the lion area can virtually experience the lions moving around through ZOO Glasses, and learn about their ecology, characteristics, and conservation information. This system can provide visitors with a consistent educational and entertaining experience even when animals are not actually visible, enhancing the value of visiting the zoo.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The user activates the ZOO glasses and enters their authentication information on the login screen. The device sends this information to the server to authenticate the user. The server confirms the user's authentication and sends data such as past visit history and personal settings to the device.

[0040] Step 2:

[0041] When a user approaches an animal area, the device uses its built-in GPS or Bluetooth beacon to determine the user's current location. The device activates its camera and performs image analysis on the captured video data to confirm the presence of animals.

[0042] Step 3:

[0043] The server receives the animal recognition results sent from the terminal and searches for past behavioral data related to that animal. The server uses generative AI to generate virtual animal data to reproduce the animal's natural movements and behaviors, and sends it to the terminal.

[0044] Step 4:

[0045] The device uses virtual animal data received from the server to create AR content, which is then displayed in the user's field of view in real time. Users can experience the virtualized animals' movements as if they were actually moving around right in front of them.

[0046] Step 5:

[0047] The server sends detailed information and entertainment information about animals to the device via voice and text. The device prioritizes and provides information of interest based on the user's actions. Users can request additional information via touch controls or voice commands.

[0048] Step 6:

[0049] When a user ends their visit, the device sends their browsing history and preferences to the server. The server stores this data and uses it to provide a personalized experience on their next visit. The user then logs out and exits the system.

[0050] (Example 1)

[0051] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0052] In recent years, zoos and museums have seen a growing demand for experiences that go beyond simple viewing, offering visitors deeper knowledge while simultaneously providing entertainment. However, conventional technology has struggled to effectively maintain visitors' interest during times when animals are inactive or when they are far away. Therefore, new methods are needed to provide visitors with consistent engagement.

[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0054] In this invention, the server includes means for authenticating information entered by the user, means for determining the user's location using location acquisition technology, and means for sending prompts to a generating AI model to generate movement data of a virtual object. This allows visitors to experience the movements and behaviors of virtual animals through augmented reality technology, even when animals are not actually visible, while simultaneously gaining deep knowledge and entertainment about animals in audio and text formats.

[0055] "Means of authenticating user-entered information" refers to a function that verifies the legitimacy of a user based on the authentication information provided by the user.

[0056] "Means of determining the user's location using location acquisition technology" refers to a function that accurately measures the user's current location using technologies such as GPS and Bluetooth beacons.

[0057] "Means for recognizing target objects" refers to functions that use cameras and sensors to identify surrounding objects and acquire related information.

[0058] "A means of sending prompts to a generating AI model to generate movement data for a virtual target" refers to a function that gives instructions to an AI model and generates new movement data based on past data and current information.

[0059] "Means of visual display using augmented reality technology" refers to a technology that overlays virtual information created by computer graphics onto images of the real world.

[0060] "Means of providing information in audio and text format" refers to a function that transmits necessary information to users through audio or text data.

[0061] This invention provides an augmented reality system that allows visitors to virtually experience the movements and information of animals and exhibits that they cannot see in the real world, in zoos and museums. The system mainly consists of a visual display device worn by the user, a location acquisition device installed in the zoo, and a server in the cloud.

[0062] The device receives the user's authentication information upon their arrival at the zoo. The user's login information is sent to the server via the device, verifying their legitimacy. This prepares the system to provide a personalized experience based on past visit data.

[0063] When the user approaches the animal area, the device accurately measures their location using its built-in GPS and Bluetooth beacons. It also recognizes surrounding animals using its camera and sensors. This information is transmitted to a server via the device for data processing. The server then sends prompts to a generating AI model to generate movement data for virtual animals. An example of a specific prompt might be, "Generate the movements of a virtual animal based on the behavior patterns of a lion."

[0064] The device receives movement data of the generated virtual animal and projects the virtual animal into the user's field of view using augmented reality technology. This allows the user to experience the animal's movements in real time. Furthermore, the server transmits detailed information about the animal to the device in audio and text formats, providing it to the user. This allows the user to enjoy a highly entertaining experience while learning about the animal's ecology and characteristics.

[0065] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0066] Step 1:

[0067] When a user arrives at the zoo, the terminal receives login information from the user. This login information is sent to the server, which uses this information to authenticate the user. The user is entered with a user ID and password, and after authentication, the server retrieves the user's past visit data from the database. This process prepares the terminal to provide visitors with personalized experience information.

[0068] Step 2:

[0069] When a user approaches an animal area, the device uses GPS and Bluetooth beacons to determine the user's location. It then activates its camera to recognize animals within its field of view. The acquired location information and camera image data are used, and the device sends this data to a server. Based on this information, the server processes the virtual movement of animals, taking into account the user's current location and the surrounding environment.

[0070] Step 3:

[0071] The server uses the received location information and animal recognition data to send prompt messages to the generating AI model to generate movement data for a virtual animal. Specifically, prompts such as "Generate the movements of a virtual animal based on the behavior patterns of a lion" are used. The AI ​​model combines past movement data with real-time location information to output the movements of the virtual animal.

[0072] Step 4:

[0073] The device receives virtual animal movement data transmitted from the server. Based on this data, it uses augmented reality technology to project the virtual animal into the user's field of view. The input is virtual animal movement data, which the device converts into visual information and displays on the screen in real time.

[0074] Step 5:

[0075] The server transmits descriptive information about animals to the device in audio and text formats. The generated information includes details about the animals' ecology, characteristics, and conservation status. The device plays this information as an audio guide or displays it on the screen as text. Through this information, users can deepen their knowledge about animals and enjoy a sustained entertainment experience.

[0076] (Application Example 1)

[0077] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0078] The aim is to solve the problem that consumers find it difficult to have an efficient and personalized purchasing experience due to the time-consuming process of trying on clothes and searching for product information.

[0079] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0080] In this invention, the server includes a visual display means using augmented reality technology to virtually recreate a product display, an information distribution means to provide product information to the user, and a control means to switch content based on the user's location and product selection results. This enables consumers to quickly and individually optimize product selection and information acquisition.

[0081] "Product display" refers to a means of presenting products in a physical or virtual way, creating a situation where consumers can visually examine the products.

[0082] Augmented reality technology is a technique that overlays computer graphics and digital information onto the real world, and is used to enrich the user's visual experience.

[0083] "Visual display means" refers to a method or apparatus provided for a user to obtain information visually via a device, and may utilize augmented reality.

[0084] "Product information" refers to information about a product's features, specifications, price, usage instructions, and design, and is provided to assist users in making purchasing decisions.

[0085] "Information distribution means" refers to methods or devices used to transmit predetermined information to specific recipients, and these methods may be carried out using audio or text format.

[0086] "User location" refers to the geographical or spatial location where the user is currently located, as identified using technologies such as GPS and beacons.

[0087] "Product selection results" refer to information about the products selected by consumers and the decisions they made regarding those selections. This data is used by the system to control its next actions.

[0088] "Control means" refers to a device or method provided for a system to switch or adjust its operation based on predetermined standards or conditions.

[0089] A "virtual product" is a product that does not exist in reality but is reproduced on a computer screen using digital technology, and can be visually confirmed by consumers.

[0090] "Generative means" refers to processes and techniques used to create new data or information based on input data.

[0091] The invention will now be described in terms of embodiments. The system of this invention provides consumers with a real-time experience in which they can virtually try on products and obtain product information using augmented reality technology.

[0092] The main components of the system are a server, user terminals (e.g., smart glasses or smartphones), and beacon devices installed in stores. The server resides in the cloud and is responsible for generating and delivering personalized product information to users within the store using a generative AI model.

[0093] The server generates virtual products using a generative AI model based on the user's past purchase data. The generated data is sent to the user's terminal via a visual display device, allowing the user to experience virtual try-on using augmented reality technology.

[0094] When a user arrives at the store, a beacon device identifies the user's location and transmits this information to a server. Based on this information, the server controls the content displayed to the user based on their location and the products they have selected.

[0095] As a concrete example, when a user requests a virtual try-on of jeans at a store, the server sends a prompt to a generation AI model to generate data on jeans of the appropriate size and style. For instance, a prompt might say, "Generate jeans recommended for him / her based on their past purchase history." This allows the user to efficiently select the most suitable jeans.

[0096] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0097] Step 1:

[0098] The server detects when a user arrives at a store using signals from beacon devices. Based on this input, the server identifies the user's location and retrieves their past purchase history from the database. This completes the preparation of the user information.

[0099] Step 2:

[0100] The server analyzes the user's purchase history and creates prompts for the generating AI model. These prompts are sent to the generating AI model to generate product information suitable for the user. An example of a prompt is, "Based on the user's past purchase history, please generate jeans that are recommended for him / her."

[0101] Step 3:

[0102] The generative AI model performs data calculations based on prompt messages received from the server to generate virtual product data. The generated product data includes size, style, and color information. This data is then ready for use in subsequent processing.

[0103] Step 4:

[0104] The server transmits information about the generated products to the user's terminal via a visual display. The terminal uses AR technology to provide the user with a virtual try-on interface. The user can then check the appearance and fit of the virtually tried-on products.

[0105] Step 5:

[0106] As the user operates the terminal and selects or adjusts virtual products, the server updates the information and product suggestions to be provided next based on the user's selection. This control mechanism allows the content to dynamically change according to the user's needs.

[0107] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0108] This invention proposes a system that combines augmented reality technology and emotion recognition technology to make the zoo experience more interactive and personalized. The system consists of a visual device (terminal) worn by the user, various sensing devices, and a processing server in the cloud.

[0109] When a user visits the zoo, they first put on a device and log into the system after going through an authentication process. The device uses a built-in camera and microphone to understand the user's emotional state, activating an emotion recognition engine. To eliminate ambiguity, multiple indicators (such as facial expressions and tone of voice) are used to analyze emotions.

[0110] The results of the emotion analysis are sent to the server, which selects the most appropriate content based on the user's current emotional state. For example, if the user is relaxed, videos that recreate the leisurely behavior of animals will be selected. On the other hand, if the user is feeling happy or excited, content that emphasizes the active movements of animals will be provided.

[0111] The device uses augmented reality (AR) technology to display virtual animals in the user's field of view based on data received from the server. Users can experience the animals' behavior in real time, and receive emotionally responsive audio guides and text information, making their zoo experience even more engaging.

[0112] For example, if the emotion engine detects that a user in the lion area is feeling tense, the device can recreate the image of a lion lying calmly and play soothing background music. This allows the user to relax and enjoy their time at the zoo. This system, with its emotion recognition engine, can enhance educational value by providing diverse information about animals while also being attentive to the visitor's emotions.

[0113] The following describes the processing flow.

[0114] Step 1:

[0115] The user activates the ZOO glasses and enters their login information. The device sends the entered login information to the server for user authentication. If authentication is successful, the server sends past visit history and personal settings information to the device.

[0116] Step 2:

[0117] When a user approaches the animal area, the device uses its built-in GPS or Bluetooth beacon to pinpoint the user's location. Simultaneously, it activates its built-in camera to capture the user's facial expressions and uses its microphone to record their voice tone.

[0118] Step 3:

[0119] The device inputs captured facial expression and audio data into an emotion recognition engine. The emotion recognition engine analyzes this data to identify the user's emotional state. The analysis results are sent to the server in real time.

[0120] Step 4:

[0121] The server selects appropriate animal behavior content based on the user's emotional state. For example, if the server determines that the user is stressed, it will select calm animal movements and relaxing background music. The server then sends the selected content to the device.

[0122] Step 5:

[0123] The device uses content data received from the server to display virtual animals overlaid on real-world scenery using augmented reality (AR) technology. Users can experience animal behavior optimized in response to their emotions in real time.

[0124] Step 6:

[0125] Furthermore, the device uses audio and text information to provide users with interesting information about animals. This allows users to gain a deeper understanding of animals while simultaneously enjoying a visual experience.

[0126] Step 7:

[0127] When a user finishes their zoo visit, the device sends the user's activity history and emotional data to a server. The server stores this data and uses it to further customize the experience for their next visit. The user then logs out of the system and shuts down the device.

[0128] (Example 2)

[0129] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0130] Zoos and similar facilities face the challenge of providing visitors with interactive and personalized experiences. In particular, there is a need to offer experiences that foster deeper understanding and satisfaction by providing information tailored to the visitor's emotional state.

[0131] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0132] In this invention, the server includes means for analyzing the user's emotional state and selecting optimal content based on that state, means for displaying the selected content using augmented reality technology, and means for providing information in audio and text formats. This makes it possible to provide visitors with an interactive and personalized experience that responds to their emotions.

[0133] "Means for analyzing a user's emotional state" refers to a technological system that acquires data such as a user's facial expressions and tone of voice, analyzes it, and identifies the user's current emotional state.

[0134] "Means for selecting optimal content based on emotional state" refers to an algorithm or system that selects the most appropriate content to provide based on the analyzed emotional state of the user.

[0135] "Means of displaying using augmented reality technology" refers to a technology or device that uses AR technology to display selected content superimposed on the real world within the user's field of view.

[0136] "Means of providing information in audio and text format" refers to systems that provide users with detailed information about animals through audio guides or text messages.

[0137] A "processing device" is a computer-based device designed to process large amounts of data quickly, and cloud servers are an example of such devices.

[0138] A "generative AI model" refers to a set of technologies, including generative battle networks and natural language processing, that have algorithms to provide optimal content based on user sentiment data.

[0139] This invention is a technological system for personalizing and making the zoo experience interactive. Users wear a visual device, a terminal, when visiting the zoo. This terminal has a built-in camera and microphone that captures the user's facial expressions and voice in real time.

[0140] The device's emotion recognition engine analyzes acquired data to identify the user's emotional state. This analysis utilizes image processing and voice analysis technologies. The analysis results are sent to a server, where a generative AI model is used to select the most appropriate content. The server operates on the cloud, enabling fast and secure data processing.

[0141] The selected content is sent to the device. The device uses augmented reality technology to display virtual animals and environments in the user's field of view. This display allows the user to experience the animals' behavior in real time and receive emotionally responsive audio guides and text information. For example, if the emotion recognition engine detects the user's tension, the device can display relaxed animals and play calming music.

[0142] An example of a prompt used in a generative AI model is the phrase, "Please specify animal content recommended for when the user is in a relaxed state." This prompt allows the server to generate and provide appropriate content to the user.

[0143] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0144] Step 1:

[0145] When a user puts on the visual device, the terminal first performs authentication. The inputs used are the user's facial image captured by the built-in camera and their voice recorded by the microphone. For data processing, a facial recognition algorithm and voice recognition software are used to log the user in. The output is the authentication success or failure status. Specifically, authentication is completed within a few seconds of the user putting on the device.

[0146] Step 2:

[0147] The device collects data using its built-in camera and microphone to analyze the user's emotional state. The input consists of image data showing the user's facial expressions and audio data showing their voice tone. For data processing, an emotion recognition engine analyzes this data to identify the user's current emotional state. The output is information about the analyzed emotional state, specifically states such as "relaxed," "excited," and "stressed."

[0148] Step 3:

[0149] The device sends the analyzed emotional state to the server. The emotional state information obtained in the previous step is used as input. A data transmission protocol is used to transmit the data to the cloud server in real time. The output is a confirmation notification that the server has received the emotional state information. Specifically, the data is transmitted encrypted, and confirmation is received within a few milliseconds.

[0150] Step 4:

[0151] The server runs a generative AI model based on the received emotional state to select the most suitable content. The inputs used are the user's emotional state information and the prompt "Please specify animal content recommended when the user is relaxed." During data calculation, the AI ​​algorithm compares the data against a historical database to select appropriate animal content. The output is the data of the selected content. In practice, the selection process is completed on the server within a few seconds.

[0152] Step 5:

[0153] The server sends the selected content to the terminal. Optimized content data is used as input. The data is delivered quickly and securely to the terminal via a transmission protocol. The output is the content data received by the terminal. Specifically, the content is compressed before transmission and decompressed after reception.

[0154] Step 6:

[0155] The device displays received content using augmented reality technology. It uses content data sent from a server as input. AR software processes the data, and a virtual animal appears in the user's field of view. The output is interactive content that the user can experience visually and aurally. Specifically, the virtual animal appears in the user's field of view, and an audio guide plays in accordance with its movements.

[0156] (Application Example 2)

[0157] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0158] Traditional zoo experiences have struggled to provide personalized content tailored to the individual emotions and interests of visitors. As a result, the information and visual enjoyment visitors experience are limited, creating a need for greater engagement. In particular, a system is needed that dynamically selects and delivers content based on the emotional state of each visitor.

[0159] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0160] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, information provision means for providing optimized animal content based on the analysis results, and visual display means for displaying virtual animals in the user's field of view using augmented reality technology. This makes it possible to provide a personalized zoo experience tailored to the emotional state of visitors.

[0161] An "emotion analysis device" is a technological device that analyzes a user's facial expressions and voice data to identify their emotional state in real time.

[0162] "Information provision means" refers to a technological device that provides animal-related content in audio or text format, tailored to the user's needs, based on the results of sentiment analysis.

[0163] A "visual display means" is a technological device that utilizes augmented reality technology to display the movements of animals in a virtual environment in real time within the user's field of view.

[0164] The embodiment of this invention consists mainly of a user-worn device, a cloud server for emotion analysis, and a data display system utilizing augmented reality (AR) technology. The system operates according to the following procedure.

[0165] When a user wears the device, it uses its camera and microphone to capture the user's facial expressions and voice in real time. This captured data is analyzed through a cloud-based emotion recognition API (e.g., Google® Cloud Vision or Microsoft® Azure® Emotion API) as a means of emotion analysis. The results of the emotion analysis are sent to a server.

[0166] The server operates an information delivery system that selects animal content appropriate for the user based on the results of sentiment analysis. This content includes videos demonstrating animal behavior, related audio guides, and text information, and is optimized for the user's emotional state.

[0167] The selected content is displayed on the device via augmented reality technology. The visual display means shows virtual animals to the user's vision, enabling them to interactively experience the content.

[0168] As a concrete example, consider a scenario where a user is experiencing a virtual tour of a wildlife park. If the emotion analysis system detects that the user is relaxed, the server can select a video of an elephant leisurely walking across a vast grassland and present the user with calming music.

[0169] Examples of prompts include, "If the user is excited, select content showing active animal behaviors that are appropriate for them," and "For a relaxed user, present calming animal behaviors."

[0170] This allows users to enjoy a unique and immersive zoo experience based on their own emotional state.

[0171] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0172] Step 1:

[0173] The device captures the user's facial expressions and voice using its camera and microphone. The input is real-time video and audio, which is collected as digital data and prepared to be sent to an emotion analysis API.

[0174] Step 2:

[0175] Facial expression and audio data transmitted from the device are received by a server in the cloud. The server uses an emotion recognition API to analyze this data and determine the user's emotional state. The output is an indicator of the user's emotional state.

[0176] Step 3:

[0177] The server selects appropriate animal content based on the user's emotional state. This process searches a database to select the most suitable video, audio guide, and text information for that emotion. The input is an indicator of the emotional state, and the output is the selected content data.

[0178] Step 4:

[0179] The server transmits the selected animal content to the terminal. This includes appropriate augmented reality video data and associated audio data. The terminal receives this information and prepares it for use with its visual display device.

[0180] Step 5:

[0181] The device displays received animal content using augmented reality technology. Virtual animals are projected into the user's field of view, allowing for an interactive experience of the content. The input is content data from the server, and the output is an AR display for the user.

[0182] Step 6:

[0183] Users experience the displayed augmented reality animal content and gain further knowledge about the animals based on audio guides and text information. This process allows users to enjoy a personalized zoo experience tailored to their own emotions.

[0184] This process allows the entire system to work together to provide the user with the best possible experience.

[0185] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0186] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0187] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0188] [Second Embodiment]

[0189] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0190] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0191] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0192] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0193] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0194] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0195] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0196] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0197] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0198] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0199] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0200] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0201] One embodiment of this invention is to provide an augmented reality experience system for visitors to a zoo to enhance their enjoyment of animals. The system mainly consists of a visual device (terminal) worn by the user, sensing devices installed in the zoo, and a database and processing server in the cloud.

[0202] When a user arrives at the zoo, the device first authenticates their login information. This information is sent to the server, which also retrieves data from the user's past visits. This prepares the device to provide information optimized for the visitor.

[0203] When the user approaches the animal area, the device uses GPS and Bluetooth beacons to pinpoint the user's location and confirms the presence of animals through its camera. Based on this information, the server references the animals' past movements and behavioral patterns and uses generative AI to create virtual animal movement data. The device then uses AR technology to project the virtual animal into the user's field of view based on this data. At this time, the realistic movements and behaviors of the animals are visually reproduced.

[0204] Furthermore, the server sends descriptions and learning information about the animals to the device in audio and text formats. Users can continue observing the animals while listening to or reading this information, gaining a deeper understanding of interesting characteristics and the animals themselves.

[0205] As a concrete example, users in the lion area can virtually experience the lions moving around through ZOO Glasses, and learn about their ecology, characteristics, and conservation information. This system can provide visitors with a consistent educational and entertaining experience even when animals are not actually visible, enhancing the value of visiting the zoo.

[0206] The following describes the processing flow.

[0207] Step 1:

[0208] The user activates the ZOO glasses and enters their authentication information on the login screen. The device sends this information to the server to authenticate the user. The server confirms the user's authentication and sends data such as past visit history and personal settings to the device.

[0209] Step 2:

[0210] When a user approaches an animal area, the device uses its built-in GPS or Bluetooth beacon to determine the user's current location. The device activates its camera and performs image analysis on the captured video data to confirm the presence of animals.

[0211] Step 3:

[0212] The server receives the animal recognition results sent from the terminal and searches for past behavioral data related to that animal. The server uses generative AI to generate virtual animal data to reproduce the animal's natural movements and behaviors, and sends it to the terminal.

[0213] Step 4:

[0214] The device uses virtual animal data received from the server to create AR content, which is then displayed in the user's field of view in real time. Users can experience the virtualized animals' movements as if they were actually moving around right in front of them.

[0215] Step 5:

[0216] The server sends detailed information and entertainment information about animals to the device via voice and text. The device prioritizes and provides information of interest based on the user's actions. Users can request additional information via touch controls or voice commands.

[0217] Step 6:

[0218] When a user ends their visit, the device sends their browsing history and preferences to the server. The server stores this data and uses it to provide a personalized experience on their next visit. The user then logs out and exits the system.

[0219] (Example 1)

[0220] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0221] In recent years, zoos and museums have seen a growing demand for experiences that go beyond simple viewing, offering visitors deeper knowledge while simultaneously providing entertainment. However, conventional technology has struggled to effectively maintain visitors' interest during times when animals are inactive or when they are far away. Therefore, new methods are needed to provide visitors with consistent engagement.

[0222] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0223] In this invention, the server includes means for authenticating information entered by the user, means for determining the user's location using location acquisition technology, and means for sending prompts to a generating AI model to generate movement data of a virtual object. This allows visitors to experience the movements and behaviors of virtual animals through augmented reality technology, even when animals are not actually visible, while simultaneously gaining deep knowledge and entertainment about animals in audio and text formats.

[0224] "Means of authenticating user-entered information" refers to a function that verifies the legitimacy of a user based on the authentication information provided by the user.

[0225] "Means of determining the user's location using location acquisition technology" refers to a function that accurately measures the user's current location using technologies such as GPS and Bluetooth beacons.

[0226] "Means for recognizing target objects" refers to functions that use cameras and sensors to identify surrounding objects and acquire related information.

[0227] "A means of sending prompts to a generating AI model to generate movement data for a virtual target" refers to a function that gives instructions to an AI model and generates new movement data based on past data and current information.

[0228] "Means of visual display using augmented reality technology" refers to a technology that overlays virtual information created by computer graphics onto images of the real world.

[0229] "Means of providing information in audio and text format" refers to a function that transmits necessary information to users through audio or text data.

[0230] This invention provides an augmented reality system that allows visitors to virtually experience the movements and information of animals and exhibits that they cannot see in the real world, in zoos and museums. The system mainly consists of a visual display device worn by the user, a location acquisition device installed in the zoo, and a server in the cloud.

[0231] The device receives the user's authentication information upon their arrival at the zoo. The user's login information is sent to the server via the device, verifying their legitimacy. This prepares the system to provide a personalized experience based on past visit data.

[0232] When the user approaches the animal area, the device accurately measures their location using its built-in GPS and Bluetooth beacons. It also recognizes surrounding animals using its camera and sensors. This information is transmitted to a server via the device for data processing. The server then sends prompts to a generating AI model to generate movement data for virtual animals. An example of a specific prompt might be, "Generate the movements of a virtual animal based on the behavior patterns of a lion."

[0233] The device receives movement data of the generated virtual animal and projects the virtual animal into the user's field of view using augmented reality technology. This allows the user to experience the animal's movements in real time. Furthermore, the server transmits detailed information about the animal to the device in audio and text formats, providing it to the user. This allows the user to enjoy a highly entertaining experience while learning about the animal's ecology and characteristics.

[0234] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0235] Step 1:

[0236] When a user arrives at the zoo, the terminal receives login information from the user. This login information is sent to the server, which uses this information to authenticate the user. The user is entered with a user ID and password, and after authentication, the server retrieves the user's past visit data from the database. This process prepares the terminal to provide visitors with personalized experience information.

[0237] Step 2:

[0238] When a user approaches an animal area, the device uses GPS and Bluetooth beacons to determine the user's location. It then activates its camera to recognize animals within its field of view. The acquired location information and camera image data are used, and the device sends this data to a server. Based on this information, the server processes the virtual movement of animals, taking into account the user's current location and the surrounding environment.

[0239] Step 3:

[0240] The server uses the received location information and animal recognition data to send prompt messages to the generating AI model to generate movement data for a virtual animal. Specifically, prompts such as "Generate the movements of a virtual animal based on the behavior patterns of a lion" are used. The AI ​​model combines past movement data with real-time location information to output the movements of the virtual animal.

[0241] Step 4:

[0242] The device receives virtual animal movement data transmitted from the server. Based on this data, it uses augmented reality technology to project the virtual animal into the user's field of view. The input is virtual animal movement data, which the device converts into visual information and displays on the screen in real time.

[0243] Step 5:

[0244] The server transmits descriptive information about animals to the device in audio and text formats. The generated information includes details about the animals' ecology, characteristics, and conservation status. The device plays this information as an audio guide or displays it on the screen as text. Through this information, users can deepen their knowledge about animals and enjoy a sustained entertainment experience.

[0245] (Application Example 1)

[0246] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0247] The aim is to solve the problem that consumers find it difficult to have an efficient and personalized purchasing experience due to the time-consuming process of trying on clothes and searching for product information.

[0248] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0249] In this invention, the server includes a visual display means using augmented reality technology to virtually recreate a product display, an information distribution means to provide product information to the user, and a control means to switch content based on the user's location and product selection results. This enables consumers to quickly and individually optimize product selection and information acquisition.

[0250] "Product display" refers to a means of presenting products in a physical or virtual way, creating a situation where consumers can visually examine the products.

[0251] Augmented reality technology is a technique that overlays computer graphics and digital information onto the real world, and is used to enrich the user's visual experience.

[0252] "Visual display means" refers to a method or apparatus provided for a user to obtain information visually via a device, and may utilize augmented reality.

[0253] "Product information" refers to information about a product's features, specifications, price, usage instructions, and design, and is provided to assist users in making purchasing decisions.

[0254] "Information distribution means" refers to methods or devices used to transmit predetermined information to specific recipients, and these methods may be carried out using audio or text format.

[0255] "User location" refers to the geographical or spatial location where the user is currently located, as identified using technologies such as GPS and beacons.

[0256] "Product selection results" refer to information about the products selected by consumers and the decisions they made regarding those selections. This data is used by the system to control its next actions.

[0257] "Control means" refers to a device or method provided for a system to switch or adjust its operation based on predetermined standards or conditions.

[0258] A "virtual product" is a product that does not exist in reality but is reproduced on a computer screen using digital technology, and can be visually confirmed by consumers.

[0259] "Generative means" refers to processes and techniques used to create new data or information based on input data.

[0260] The invention will now be described in terms of embodiments. The system of this invention provides consumers with a real-time experience in which they can virtually try on products and obtain product information using augmented reality technology.

[0261] The main components of the system are a server, user terminals (e.g., smart glasses or smartphones), and beacon devices installed in stores. The server resides in the cloud and is responsible for generating and delivering personalized product information to users within the store using a generative AI model.

[0262] The server generates virtual products using a generative AI model based on the user's past purchase data. The generated data is sent to the user's terminal via a visual display device, allowing the user to experience virtual try-on using augmented reality technology.

[0263] When a user arrives at the store, a beacon device identifies the user's location and transmits this information to a server. Based on this information, the server controls the content displayed to the user based on their location and the products they have selected.

[0264] As a concrete example, when a user requests a virtual try-on of jeans at a store, the server sends a prompt to a generation AI model to generate data on jeans of the appropriate size and style. For instance, a prompt might say, "Generate jeans recommended for him / her based on their past purchase history." This allows the user to efficiently select the most suitable jeans.

[0265] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0266] Step 1:

[0267] The server detects when a user arrives at a store using signals from beacon devices. Based on this input, the server identifies the user's location and retrieves their past purchase history from the database. This completes the preparation of the user information.

[0268] Step 2:

[0269] The server analyzes the user's purchase history and creates prompts for the generating AI model. These prompts are sent to the generating AI model to generate product information suitable for the user. An example of a prompt is, "Based on the user's past purchase history, please generate jeans that are recommended for him / her."

[0270] Step 3:

[0271] The generative AI model performs data calculations based on prompt messages received from the server to generate virtual product data. The generated product data includes size, style, and color information. This data is then ready for use in subsequent processing.

[0272] Step 4:

[0273] The server transmits information about the generated products to the user's terminal via a visual display. The terminal uses AR technology to provide the user with a virtual try-on interface. The user can then check the appearance and fit of the virtually tried-on products.

[0274] Step 5:

[0275] As the user operates the terminal and selects or adjusts virtual products, the server updates the information and product suggestions to be provided next based on the user's selection. This control mechanism allows the content to dynamically change according to the user's needs.

[0276] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0277] This invention proposes a system that combines augmented reality technology and emotion recognition technology to make the zoo experience more interactive and personalized. The system consists of a visual device (terminal) worn by the user, various sensing devices, and a processing server in the cloud.

[0278] When a user visits the zoo, they first put on a device and log into the system after going through an authentication process. The device uses a built-in camera and microphone to understand the user's emotional state, activating an emotion recognition engine. To eliminate ambiguity, multiple indicators (such as facial expressions and tone of voice) are used to analyze emotions.

[0279] The results of the emotion analysis are sent to the server, which selects the most appropriate content based on the user's current emotional state. For example, if the user is relaxed, videos that recreate the leisurely behavior of animals will be selected. On the other hand, if the user is feeling happy or excited, content that emphasizes the active movements of animals will be provided.

[0280] The device uses augmented reality (AR) technology to display virtual animals in the user's field of view based on data received from the server. Users can experience the animals' behavior in real time, and receive emotionally responsive audio guides and text information, making their zoo experience even more engaging.

[0281] For example, if the emotion engine detects that a user in the lion area is feeling tense, the device can recreate the image of a lion lying calmly and play soothing background music. This allows the user to relax and enjoy their time at the zoo. This system, with its emotion recognition engine, can enhance educational value by providing diverse information about animals while also being attentive to the visitor's emotions.

[0282] The following describes the processing flow.

[0283] Step 1:

[0284] The user launches the ZOO Glass and enters the login information. The terminal sends the entered login information to the server for user authentication. When the authentication is successful, the server sends the past visit history and personal setting information to the terminal.

[0285] Step 2:

[0286] When the user heads towards the animal area, the terminal uses the built-in GPS or Bluetooth beacon to identify the user's location. At the same time, it activates the built-in camera to capture the user's expression and uses the microphone to record the voice tone.

[0287] Step 3:

[0288] The terminal inputs the captured expression data and voice data into the emotion recognition engine. The emotion recognition engine analyzes these data to identify the user's emotional state. The analysis result is sent to the server in real time.

[0289] Step 4:

[0290] Based on the user's emotional state, the server selects appropriate animal movement content. For example, if the user is judged to be nervous, it selects the calm movements of animals and relaxing BGM. The server sends the selected content to the terminal.

[0291] Step 5:

[0292] The terminal uses the content data received from the server to display virtual animals superimposed on the real scenery using AR technology. The user can experience the optimized actions of animals in real time according to their emotions.

[0293] Step 6:

[0294] Furthermore, the device uses audio and text information to provide users with interesting information about animals. This allows users to gain a deeper understanding of animals while simultaneously enjoying a visual experience.

[0295] Step 7:

[0296] When a user finishes their zoo visit, the device sends the user's activity history and emotional data to a server. The server stores this data and uses it to further customize the experience for their next visit. The user then logs out of the system and shuts down the device.

[0297] (Example 2)

[0298] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0299] Zoos and similar facilities face the challenge of providing visitors with interactive and personalized experiences. In particular, there is a need to offer experiences that foster deeper understanding and satisfaction by providing information tailored to the visitor's emotional state.

[0300] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0301] In this invention, the server includes means for analyzing the user's emotional state and selecting optimal content based on that state, means for displaying the selected content using augmented reality technology, and means for providing information in audio and text formats. This makes it possible to provide visitors with an interactive and personalized experience that responds to their emotions.

[0302] "Means for analyzing a user's emotional state" refers to a technological system that acquires data such as a user's facial expressions and tone of voice, analyzes it, and identifies the user's current emotional state.

[0303] The means for selecting optimal content based on the emotional state is an algorithm or system that selects the most appropriate content to be provided according to the analyzed emotional state of the user.

[0304] The means for displaying using augmented reality technology is a technology or device that uses AR technology to display the selected content by overlapping it with the real world in the user's field of vision.

[0305] The means for providing information in voice and text formats is a system for providing detailed information about animals to the user via voice guides and text messages.

[0306] The processing device is a computer-based device for quickly processing a large amount of data, such as a cloud server.

[0307] The generative AI model has an algorithm for providing optimal content based on the user's emotional data and refers to a series of technologies including generative adversarial networks and natural language processing.

[0308] This invention is a technical system for personalizing and making the experience in a zoo interactive. When the user visits the zoo, they wear a terminal which is a visual device. This terminal has a built-in camera and microphone, and acquires the user's expression and voice in real time.

[0309] The emotion recognition engine of the terminal analyzes the acquired data and identifies the user's emotional state. Image processing technology and voice analysis technology are used for this analysis. The analysis result is sent to the server, and optimal content is selected using the generative AI model. The server operates on the cloud, enabling fast and secure data processing.

[0310] The selected content is sent to the device. The device uses augmented reality technology to display virtual animals and environments in the user's field of view. This display allows the user to experience the animals' behavior in real time and receive emotionally responsive audio guides and text information. For example, if the emotion recognition engine detects the user's tension, the device can display relaxed animals and play calming music.

[0311] An example of a prompt used in a generative AI model is the phrase, "Please specify animal content recommended for when the user is in a relaxed state." This prompt allows the server to generate and provide appropriate content to the user.

[0312] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0313] Step 1:

[0314] When a user puts on the visual device, the terminal first performs authentication. The inputs used are the user's facial image captured by the built-in camera and their voice recorded by the microphone. For data processing, a facial recognition algorithm and voice recognition software are used to log the user in. The output is the authentication success or failure status. Specifically, authentication is completed within a few seconds of the user putting on the device.

[0315] Step 2:

[0316] The device collects data using its built-in camera and microphone to analyze the user's emotional state. The input consists of image data showing the user's facial expressions and audio data showing their voice tone. For data processing, an emotion recognition engine analyzes this data to identify the user's current emotional state. The output is information about the analyzed emotional state, specifically states such as "relaxed," "excited," and "stressed."

[0317] Step 3:

[0318] The device sends the analyzed emotional state to the server. The emotional state information obtained in the previous step is used as input. A data transmission protocol is used to transmit the data to the cloud server in real time. The output is a confirmation notification that the server has received the emotional state information. Specifically, the data is transmitted encrypted, and confirmation is received within a few milliseconds.

[0319] Step 4:

[0320] The server runs a generative AI model based on the received emotional state to select the most suitable content. The inputs used are the user's emotional state information and the prompt "Please specify animal content recommended when the user is relaxed." During data calculation, the AI ​​algorithm compares the data against a historical database to select appropriate animal content. The output is the data of the selected content. In practice, the selection process is completed on the server within a few seconds.

[0321] Step 5:

[0322] The server sends the selected content to the terminal. Optimized content data is used as input. The data is delivered quickly and securely to the terminal via a transmission protocol. The output is the content data received by the terminal. Specifically, the content is compressed before transmission and decompressed after reception.

[0323] Step 6:

[0324] The device displays received content using augmented reality technology. It uses content data sent from a server as input. AR software processes the data, and a virtual animal appears in the user's field of view. The output is interactive content that the user can experience visually and aurally. Specifically, the virtual animal appears in the user's field of view, and an audio guide plays in accordance with its movements.

[0325] (Application Example 2)

[0326] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0327] Traditional zoo experiences have struggled to provide personalized content tailored to the individual emotions and interests of visitors. As a result, the information and visual enjoyment visitors experience are limited, creating a need for greater engagement. In particular, a system is needed that dynamically selects and delivers content based on the emotional state of each visitor.

[0328] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0329] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, information provision means for providing optimized animal content based on the analysis results, and visual display means for displaying virtual animals in the user's field of view using augmented reality technology. This makes it possible to provide a personalized zoo experience tailored to the emotional state of visitors.

[0330] An "emotion analysis device" is a technological device that analyzes a user's facial expressions and voice data to identify their emotional state in real time.

[0331] "Information provision means" refers to a technological device that provides animal-related content in audio or text format, tailored to the user's needs, based on the results of sentiment analysis.

[0332] A "visual display means" is a technological device that utilizes augmented reality technology to display the movements of animals in a virtual environment in real time within the user's field of view.

[0333] The embodiment of this invention consists mainly of a user-worn device, a cloud server for emotion analysis, and a data display system utilizing augmented reality (AR) technology. The system operates according to the following procedure.

[0334] When a user wears the device, it uses its camera and microphone to capture the user's facial expressions and voice in real time. This captured data is analyzed through a cloud-based emotion recognition API (e.g., Google Cloud Vision or Microsoft Azure Emotion API) as a means of emotion analysis. The results of the emotion analysis are sent to a server.

[0335] The server operates an information delivery system that selects animal content appropriate for the user based on the results of sentiment analysis. This content includes videos demonstrating animal behavior, related audio guides, and text information, and is optimized for the user's emotional state.

[0336] The selected content is displayed on the device via augmented reality technology. The visual display means shows virtual animals to the user's vision, enabling them to interactively experience the content.

[0337] As a concrete example, consider a scenario where a user is experiencing a virtual tour of a wildlife park. If the emotion analysis system detects that the user is relaxed, the server can select a video of an elephant leisurely walking across a vast grassland and present the user with calming music.

[0338] Examples of prompts include, "If the user is excited, select content showing active animal behaviors that are appropriate for them," and "For a relaxed user, present calming animal behaviors."

[0339] This allows users to enjoy a unique and immersive zoo experience based on their own emotional state.

[0340] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0341] Step 1:

[0342] The device captures the user's facial expressions and voice using its camera and microphone. The input is real-time video and audio, which is collected as digital data and prepared to be sent to an emotion analysis API.

[0343] Step 2:

[0344] Facial expression and audio data transmitted from the device are received by a server in the cloud. The server uses an emotion recognition API to analyze this data and determine the user's emotional state. The output is an indicator of the user's emotional state.

[0345] Step 3:

[0346] The server selects appropriate animal content based on the user's emotional state. This process searches a database to select the most suitable video, audio guide, and text information for that emotion. The input is an indicator of the emotional state, and the output is the selected content data.

[0347] Step 4:

[0348] The server transmits the selected animal content to the terminal. This includes appropriate augmented reality video data and associated audio data. The terminal receives this information and prepares it for use with its visual display device.

[0349] Step 5:

[0350] The device displays received animal content using augmented reality technology. Virtual animals are projected into the user's field of view, allowing for an interactive experience of the content. The input is content data from the server, and the output is an AR display for the user.

[0351] Step 6:

[0352] Users experience the displayed augmented reality animal content and gain further knowledge about the animals based on audio guides and text information. This process allows users to enjoy a personalized zoo experience tailored to their own emotions.

[0353] This process allows the entire system to work together to provide the user with the best possible experience.

[0354] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0355] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0356] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0357] [Third Embodiment]

[0358] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0359] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0360] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0361] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0362] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0363] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0364] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0365] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0366] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0367] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0368] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0369] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0370] One embodiment of this invention is to provide an augmented reality experience system for visitors to a zoo to enhance their enjoyment of animals. The system mainly consists of a visual device (terminal) worn by the user, sensing devices installed in the zoo, and a database and processing server in the cloud.

[0371] When a user arrives at the zoo, the device first authenticates their login information. This information is sent to the server, which also retrieves data from the user's past visits. This prepares the device to provide information optimized for the visitor.

[0372] When the user approaches the animal area, the device uses GPS and Bluetooth beacons to pinpoint the user's location and confirms the presence of animals through its camera. Based on this information, the server references the animals' past movements and behavioral patterns and uses generative AI to create virtual animal movement data. The device then uses AR technology to project the virtual animal into the user's field of view based on this data. At this time, the realistic movements and behaviors of the animals are visually reproduced.

[0373] Furthermore, the server sends descriptions and learning information about the animals to the device in audio and text formats. Users can continue observing the animals while listening to or reading this information, gaining a deeper understanding of interesting characteristics and the animals themselves.

[0374] As a concrete example, users in the lion area can virtually experience the lions moving around through ZOO Glasses, and learn about their ecology, characteristics, and conservation information. This system can provide visitors with a consistent educational and entertaining experience even when animals are not actually visible, enhancing the value of visiting the zoo.

[0375] The following describes the processing flow.

[0376] Step 1:

[0377] The user activates the ZOO glasses and enters their authentication information on the login screen. The device sends this information to the server to authenticate the user. The server confirms the user's authentication and sends data such as past visit history and personal settings to the device.

[0378] Step 2:

[0379] When a user approaches an animal area, the device uses its built-in GPS or Bluetooth beacon to determine the user's current location. The device activates its camera and performs image analysis on the captured video data to confirm the presence of animals.

[0380] Step 3:

[0381] The server receives the animal recognition results sent from the terminal and searches for past behavioral data related to that animal. The server uses generative AI to generate virtual animal data to reproduce the animal's natural movements and behaviors, and sends it to the terminal.

[0382] Step 4:

[0383] The device uses virtual animal data received from the server to create AR content, which is then displayed in the user's field of view in real time. Users can experience the virtualized animals' movements as if they were actually moving around right in front of them.

[0384] Step 5:

[0385] The server sends detailed information and entertainment information about animals to the device via voice and text. The device prioritizes and provides information of interest based on the user's actions. Users can request additional information via touch controls or voice commands.

[0386] Step 6:

[0387] When a user ends their visit, the device sends their browsing history and preferences to the server. The server stores this data and uses it to provide a personalized experience on their next visit. The user then logs out and exits the system.

[0388] (Example 1)

[0389] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0390] In recent years, zoos and museums have seen a growing demand for experiences that go beyond simple viewing, offering visitors deeper knowledge while simultaneously providing entertainment. However, conventional technology has struggled to effectively maintain visitors' interest during times when animals are inactive or when they are far away. Therefore, new methods are needed to provide visitors with consistent engagement.

[0391] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0392] In this invention, the server includes means for authenticating information entered by the user, means for determining the user's location using location acquisition technology, and means for sending prompts to a generating AI model to generate movement data of a virtual object. This allows visitors to experience the movements and behaviors of virtual animals through augmented reality technology, even when animals are not actually visible, while simultaneously gaining deep knowledge and entertainment about animals in audio and text formats.

[0393] "Means of authenticating user-entered information" refers to a function that verifies the legitimacy of a user based on the authentication information provided by the user.

[0394] "Means of determining the user's location using location acquisition technology" refers to a function that accurately measures the user's current location using technologies such as GPS and Bluetooth beacons.

[0395] "Means for recognizing target objects" refers to functions that use cameras and sensors to identify surrounding objects and acquire related information.

[0396] "A means of sending prompts to a generating AI model to generate movement data for a virtual target" refers to a function that gives instructions to an AI model and generates new movement data based on past data and current information.

[0397] "Means of visual display using augmented reality technology" refers to a technology that overlays virtual information created by computer graphics onto images of the real world.

[0398] "Means of providing information in audio and text format" refers to a function that transmits necessary information to users through audio or text data.

[0399] This invention provides an augmented reality system that allows visitors to virtually experience the movements and information of animals and exhibits that they cannot see in the real world, in zoos and museums. The system mainly consists of a visual display device worn by the user, a location acquisition device installed in the zoo, and a server in the cloud.

[0400] The device receives the user's authentication information upon their arrival at the zoo. The user's login information is sent to the server via the device, verifying their legitimacy. This prepares the system to provide a personalized experience based on past visit data.

[0401] When the user approaches the animal area, the device accurately measures their location using its built-in GPS and Bluetooth beacons. It also recognizes surrounding animals using its camera and sensors. This information is transmitted to a server via the device for data processing. The server then sends prompts to a generating AI model to generate movement data for virtual animals. An example of a specific prompt might be, "Generate the movements of a virtual animal based on the behavior patterns of a lion."

[0402] The device receives movement data of the generated virtual animal and projects the virtual animal into the user's field of view using augmented reality technology. This allows the user to experience the animal's movements in real time. Furthermore, the server transmits detailed information about the animal to the device in audio and text formats, providing it to the user. This allows the user to enjoy a highly entertaining experience while learning about the animal's ecology and characteristics.

[0403] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0404] Step 1:

[0405] When a user arrives at the zoo, the terminal receives login information from the user. This login information is sent to the server, which uses this information to authenticate the user. The user is entered with a user ID and password, and after authentication, the server retrieves the user's past visit data from the database. This process prepares the terminal to provide visitors with personalized experience information.

[0406] Step 2:

[0407] When a user approaches an animal area, the device uses GPS and Bluetooth beacons to determine the user's location. It then activates its camera to recognize animals within its field of view. The acquired location information and camera image data are used, and the device sends this data to a server. Based on this information, the server processes the virtual movement of animals, taking into account the user's current location and the surrounding environment.

[0408] Step 3:

[0409] The server uses the received location information and animal recognition data to send prompt messages to the generating AI model to generate movement data for a virtual animal. Specifically, prompts such as "Generate the movements of a virtual animal based on the behavior patterns of a lion" are used. The AI ​​model combines past movement data with real-time location information to output the movements of the virtual animal.

[0410] Step 4:

[0411] The device receives virtual animal movement data transmitted from the server. Based on this data, it uses augmented reality technology to project the virtual animal into the user's field of view. The input is virtual animal movement data, which the device converts into visual information and displays on the screen in real time.

[0412] Step 5:

[0413] The server transmits descriptive information about animals to the device in audio and text formats. The generated information includes details about the animals' ecology, characteristics, and conservation status. The device plays this information as an audio guide or displays it on the screen as text. Through this information, users can deepen their knowledge about animals and enjoy a sustained entertainment experience.

[0414] (Application Example 1)

[0415] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0416] The aim is to solve the problem that consumers find it difficult to have an efficient and personalized purchasing experience due to the time-consuming process of trying on clothes and searching for product information.

[0417] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0418] In this invention, the server includes a visual display means using augmented reality technology to virtually recreate a product display, an information distribution means to provide product information to the user, and a control means to switch content based on the user's location and product selection results. This enables consumers to quickly and individually optimize product selection and information acquisition.

[0419] "Product display" refers to a means of presenting products in a physical or virtual way, creating a situation where consumers can visually examine the products.

[0420] Augmented reality technology is a technique that overlays computer graphics and digital information onto the real world, and is used to enrich the user's visual experience.

[0421] "Visual display means" refers to a method or apparatus provided for a user to obtain information visually via a device, and may utilize augmented reality.

[0422] "Product information" refers to information about a product's features, specifications, price, usage instructions, and design, and is provided to assist users in making purchasing decisions.

[0423] "Information distribution means" refers to methods or devices used to transmit predetermined information to specific recipients, and these methods may be carried out using audio or text format.

[0424] "User location" refers to the geographical or spatial location where the user is currently located, as identified using technologies such as GPS and beacons.

[0425] "Product selection results" refer to information about the products selected by consumers and the decisions they made regarding those selections. This data is used by the system to control its next actions.

[0426] "Control means" refers to a device or method provided for a system to switch or adjust its operation based on predetermined standards or conditions.

[0427] A "virtual product" is a product that does not exist in reality but is reproduced on a computer screen using digital technology, and can be visually confirmed by consumers.

[0428] "Generative means" refers to processes and techniques used to create new data or information based on input data.

[0429] The invention will now be described in terms of embodiments. The system of this invention provides consumers with a real-time experience in which they can virtually try on products and obtain product information using augmented reality technology.

[0430] The main components of the system are a server, user terminals (e.g., smart glasses or smartphones), and beacon devices installed in stores. The server resides in the cloud and is responsible for generating and delivering personalized product information to users within the store using a generative AI model.

[0431] The server generates virtual products using a generative AI model based on the user's past purchase data. The generated data is sent to the user's terminal via a visual display device, allowing the user to experience virtual try-on using augmented reality technology.

[0432] When a user arrives at the store, a beacon device identifies the user's location and transmits this information to a server. Based on this information, the server controls the content displayed to the user based on their location and the products they have selected.

[0433] As a concrete example, when a user requests a virtual try-on of jeans at a store, the server sends a prompt to a generation AI model to generate data on jeans of the appropriate size and style. For instance, a prompt might say, "Generate jeans recommended for him / her based on their past purchase history." This allows the user to efficiently select the most suitable jeans.

[0434] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0435] Step 1:

[0436] The server detects when a user arrives at a store using signals from beacon devices. Based on this input, the server identifies the user's location and retrieves their past purchase history from the database. This completes the preparation of the user information.

[0437] Step 2:

[0438] The server analyzes the user's purchase history and creates prompts for the generating AI model. These prompts are sent to the generating AI model to generate product information suitable for the user. An example of a prompt is, "Based on the user's past purchase history, please generate jeans that are recommended for him / her."

[0439] Step 3:

[0440] The generative AI model performs data calculations based on prompt messages received from the server to generate virtual product data. The generated product data includes size, style, and color information. This data is then ready for use in subsequent processing.

[0441] Step 4:

[0442] The server transmits information about the generated products to the user's terminal via a visual display. The terminal uses AR technology to provide the user with a virtual try-on interface. The user can then check the appearance and fit of the virtually tried-on products.

[0443] Step 5:

[0444] As the user operates the terminal and selects or adjusts virtual products, the server updates the information and product suggestions to be provided next based on the user's selection. This control mechanism allows the content to dynamically change according to the user's needs.

[0445] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0446] This invention proposes a system that combines augmented reality technology and emotion recognition technology to make the zoo experience more interactive and personalized. The system consists of a visual device (terminal) worn by the user, various sensing devices, and a processing server in the cloud.

[0447] When a user visits the zoo, they first put on a device and log into the system after going through an authentication process. The device uses a built-in camera and microphone to understand the user's emotional state, activating an emotion recognition engine. To eliminate ambiguity, multiple indicators (such as facial expressions and tone of voice) are used to analyze emotions.

[0448] The results of the emotion analysis are sent to the server, which selects the most appropriate content based on the user's current emotional state. For example, if the user is relaxed, videos that recreate the leisurely behavior of animals will be selected. On the other hand, if the user is feeling happy or excited, content that emphasizes the active movements of animals will be provided.

[0449] The device uses augmented reality (AR) technology to display virtual animals in the user's field of view based on data received from the server. Users can experience the animals' behavior in real time, and receive emotionally responsive audio guides and text information, making their zoo experience even more engaging.

[0450] For example, if the emotion engine detects that a user in the lion area is feeling tense, the device can recreate the image of a lion lying calmly and play soothing background music. This allows the user to relax and enjoy their time at the zoo. This system, with its emotion recognition engine, can enhance educational value by providing diverse information about animals while also being attentive to the visitor's emotions.

[0451] The following describes the processing flow.

[0452] Step 1:

[0453] The user activates the ZOO glasses and enters their login information. The device sends the entered login information to the server for user authentication. If authentication is successful, the server sends past visit history and personal settings information to the device.

[0454] Step 2:

[0455] When a user approaches the animal area, the device uses its built-in GPS or Bluetooth beacon to pinpoint the user's location. Simultaneously, it activates its built-in camera to capture the user's facial expressions and uses its microphone to record their voice tone.

[0456] Step 3:

[0457] The device inputs captured facial expression and audio data into an emotion recognition engine. The emotion recognition engine analyzes this data to identify the user's emotional state. The analysis results are sent to the server in real time.

[0458] Step 4:

[0459] The server selects appropriate animal behavior content based on the user's emotional state. For example, if the server determines that the user is stressed, it will select calm animal movements and relaxing background music. The server then sends the selected content to the device.

[0460] Step 5:

[0461] The device uses content data received from the server to display virtual animals overlaid on real-world scenery using augmented reality (AR) technology. Users can experience animal behavior optimized in response to their emotions in real time.

[0462] Step 6:

[0463] Furthermore, the device uses audio and text information to provide users with interesting information about animals. This allows users to gain a deeper understanding of animals while simultaneously enjoying a visual experience.

[0464] Step 7:

[0465] When a user finishes their zoo visit, the device sends the user's activity history and emotional data to a server. The server stores this data and uses it to further customize the experience for their next visit. The user then logs out of the system and shuts down the device.

[0466] (Example 2)

[0467] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0468] Zoos and similar facilities face the challenge of providing visitors with interactive and personalized experiences. In particular, there is a need to offer experiences that foster deeper understanding and satisfaction by providing information tailored to the visitor's emotional state.

[0469] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0470] In this invention, the server includes means for analyzing the user's emotional state and selecting optimal content based on that state, means for displaying the selected content using augmented reality technology, and means for providing information in audio and text formats. This makes it possible to provide visitors with an interactive and personalized experience that responds to their emotions.

[0471] "Means for analyzing a user's emotional state" refers to a technological system that acquires data such as a user's facial expressions and tone of voice, analyzes it, and identifies the user's current emotional state.

[0472] "Means for selecting optimal content based on emotional state" refers to an algorithm or system that selects the most appropriate content to provide based on the analyzed emotional state of the user.

[0473] "Means of displaying using augmented reality technology" refers to a technology or device that uses AR technology to display selected content superimposed on the real world within the user's field of view.

[0474] "Means of providing information in audio and text format" refers to systems that provide users with detailed information about animals through audio guides or text messages.

[0475] A "processing device" is a computer-based device designed to process large amounts of data quickly, and cloud servers are an example of such devices.

[0476] A "generative AI model" refers to a set of technologies, including generative battle networks and natural language processing, that have algorithms to provide optimal content based on user sentiment data.

[0477] This invention is a technological system for personalizing and making the zoo experience interactive. Users wear a visual device, a terminal, when visiting the zoo. This terminal has a built-in camera and microphone that captures the user's facial expressions and voice in real time.

[0478] The device's emotion recognition engine analyzes acquired data to identify the user's emotional state. This analysis utilizes image processing and voice analysis technologies. The analysis results are sent to a server, where a generative AI model is used to select the most appropriate content. The server operates on the cloud, enabling fast and secure data processing.

[0479] The selected content is sent to the device. The device uses augmented reality technology to display virtual animals and environments in the user's field of view. This display allows the user to experience the animals' behavior in real time and receive emotionally responsive audio guides and text information. For example, if the emotion recognition engine detects the user's tension, the device can display relaxed animals and play calming music.

[0480] An example of a prompt used in a generative AI model is the phrase, "Please specify animal content recommended for when the user is in a relaxed state." This prompt allows the server to generate and provide appropriate content to the user.

[0481] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0482] Step 1:

[0483] When a user puts on the visual device, the terminal first performs authentication. The inputs used are the user's facial image captured by the built-in camera and their voice recorded by the microphone. For data processing, a facial recognition algorithm and voice recognition software are used to log the user in. The output is the authentication success or failure status. Specifically, authentication is completed within a few seconds of the user putting on the device.

[0484] Step 2:

[0485] The device collects data using its built-in camera and microphone to analyze the user's emotional state. The input consists of image data showing the user's facial expressions and audio data showing their voice tone. For data processing, an emotion recognition engine analyzes this data to identify the user's current emotional state. The output is information about the analyzed emotional state, specifically states such as "relaxed," "excited," and "stressed."

[0486] Step 3:

[0487] The device sends the analyzed emotional state to the server. The emotional state information obtained in the previous step is used as input. A data transmission protocol is used to transmit the data to the cloud server in real time. The output is a confirmation notification that the server has received the emotional state information. Specifically, the data is transmitted encrypted, and confirmation is received within a few milliseconds.

[0488] Step 4:

[0489] The server runs a generative AI model based on the received emotional state to select the most suitable content. The inputs used are the user's emotional state information and the prompt "Please specify animal content recommended when the user is relaxed." During data calculation, the AI ​​algorithm compares the data against a historical database to select appropriate animal content. The output is the data of the selected content. In practice, the selection process is completed on the server within a few seconds.

[0490] Step 5:

[0491] The server sends the selected content to the terminal. Optimized content data is used as input. The data is delivered quickly and securely to the terminal via a transmission protocol. The output is the content data received by the terminal. Specifically, the content is compressed before transmission and decompressed after reception.

[0492] Step 6:

[0493] The device displays received content using augmented reality technology. It uses content data sent from a server as input. AR software processes the data, and a virtual animal appears in the user's field of view. The output is interactive content that the user can experience visually and aurally. Specifically, the virtual animal appears in the user's field of view, and an audio guide plays in accordance with its movements.

[0494] (Application Example 2)

[0495] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0496] Traditional zoo experiences have struggled to provide personalized content tailored to the individual emotions and interests of visitors. As a result, the information and visual enjoyment visitors experience are limited, creating a need for greater engagement. In particular, a system is needed that dynamically selects and delivers content based on the emotional state of each visitor.

[0497] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0498] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, information provision means for providing optimized animal content based on the analysis results, and visual display means for displaying virtual animals in the user's field of view using augmented reality technology. This makes it possible to provide a personalized zoo experience tailored to the emotional state of visitors.

[0499] An "emotion analysis device" is a technological device that analyzes a user's facial expressions and voice data to identify their emotional state in real time.

[0500] "Information provision means" refers to a technological device that provides animal-related content in audio or text format, tailored to the user's needs, based on the results of sentiment analysis.

[0501] A "visual display means" is a technological device that utilizes augmented reality technology to display the movements of animals in a virtual environment in real time within the user's field of view.

[0502] The embodiment of this invention consists mainly of a user-worn device, a cloud server for emotion analysis, and a data display system utilizing augmented reality (AR) technology. The system operates according to the following procedure.

[0503] When a user wears the device, it uses its camera and microphone to capture the user's facial expressions and voice in real time. This captured data is analyzed through a cloud-based emotion recognition API (e.g., Google Cloud Vision or Microsoft Azure Emotion API) as a means of emotion analysis. The results of the emotion analysis are sent to a server.

[0504] The server operates an information delivery system that selects animal content appropriate for the user based on the results of sentiment analysis. This content includes videos demonstrating animal behavior, related audio guides, and text information, and is optimized for the user's emotional state.

[0505] The selected content is displayed on the device via augmented reality technology. The visual display means shows virtual animals to the user's vision, enabling them to interactively experience the content.

[0506] As a concrete example, consider a scenario where a user is experiencing a virtual tour of a wildlife park. If the emotion analysis system detects that the user is relaxed, the server can select a video of an elephant leisurely walking across a vast grassland and present the user with calming music.

[0507] Examples of prompts include, "If the user is excited, select content showing active animal behaviors that are appropriate for them," and "For a relaxed user, present calming animal behaviors."

[0508] This allows users to enjoy a unique and immersive zoo experience based on their own emotional state.

[0509] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0510] Step 1:

[0511] The device captures the user's facial expressions and voice using its camera and microphone. The input is real-time video and audio, which is collected as digital data and prepared to be sent to an emotion analysis API.

[0512] Step 2:

[0513] Facial expression and audio data transmitted from the device are received by a server in the cloud. The server uses an emotion recognition API to analyze this data and determine the user's emotional state. The output is an indicator of the user's emotional state.

[0514] Step 3:

[0515] The server selects appropriate animal content based on the user's emotional state. This process searches a database to select the most suitable video, audio guide, and text information for that emotion. The input is an indicator of the emotional state, and the output is the selected content data.

[0516] Step 4:

[0517] The server transmits the selected animal content to the terminal. This includes appropriate augmented reality video data and associated audio data. The terminal receives this information and prepares it for use with its visual display device.

[0518] Step 5:

[0519] The device displays received animal content using augmented reality technology. Virtual animals are projected into the user's field of view, allowing for an interactive experience of the content. The input is content data from the server, and the output is an AR display for the user.

[0520] Step 6:

[0521] Users experience the displayed augmented reality animal content and gain further knowledge about the animals based on audio guides and text information. This process allows users to enjoy a personalized zoo experience tailored to their own emotions.

[0522] This process allows the entire system to work together to provide the user with the best possible experience.

[0523] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0524] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0525] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0526] [Fourth Embodiment]

[0527] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0528] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0529] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0530] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0531] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0532] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0533] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0534] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0535] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0536] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0537] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0538] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0539] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0540] One embodiment of this invention is to provide an augmented reality experience system for visitors to a zoo to enhance their enjoyment of animals. The system mainly consists of a visual device (terminal) worn by the user, sensing devices installed in the zoo, and a database and processing server in the cloud.

[0541] When a user arrives at the zoo, the device first authenticates their login information. This information is sent to the server, which also retrieves data from the user's past visits. This prepares the device to provide information optimized for the visitor.

[0542] When the user approaches the animal area, the device uses GPS and Bluetooth beacons to pinpoint the user's location and confirms the presence of animals through its camera. Based on this information, the server references the animals' past movements and behavioral patterns and uses generative AI to create virtual animal movement data. The device then uses AR technology to project the virtual animal into the user's field of view based on this data. At this time, the realistic movements and behaviors of the animals are visually reproduced.

[0543] Furthermore, the server sends descriptions and learning information about the animals to the device in audio and text formats. Users can continue observing the animals while listening to or reading this information, gaining a deeper understanding of interesting characteristics and the animals themselves.

[0544] As a concrete example, users in the lion area can virtually experience the lions moving around through ZOO Glasses, and learn about their ecology, characteristics, and conservation information. This system can provide visitors with a consistent educational and entertaining experience even when animals are not actually visible, enhancing the value of visiting the zoo.

[0545] The following describes the processing flow.

[0546] Step 1:

[0547] The user activates the ZOO glasses and enters their authentication information on the login screen. The device sends this information to the server to authenticate the user. The server confirms the user's authentication and sends data such as past visit history and personal settings to the device.

[0548] Step 2:

[0549] When a user approaches an animal area, the device uses its built-in GPS or Bluetooth beacon to determine the user's current location. The device activates its camera and performs image analysis on the captured video data to confirm the presence of animals.

[0550] Step 3:

[0551] The server receives the animal recognition results sent from the terminal and searches for past behavioral data related to that animal. The server uses generative AI to generate virtual animal data to reproduce the animal's natural movements and behaviors, and sends it to the terminal.

[0552] Step 4:

[0553] The device uses virtual animal data received from the server to create AR content, which is then displayed in the user's field of view in real time. Users can experience the virtualized animals' movements as if they were actually moving around right in front of them.

[0554] Step 5:

[0555] The server sends detailed information and entertainment information about animals to the device via voice and text. The device prioritizes and provides information of interest based on the user's actions. Users can request additional information via touch controls or voice commands.

[0556] Step 6:

[0557] When a user ends their visit, the device sends their browsing history and preferences to the server. The server stores this data and uses it to provide a personalized experience on their next visit. The user then logs out and exits the system.

[0558] (Example 1)

[0559] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0560] In recent years, zoos and museums have seen a growing demand for experiences that go beyond simple viewing, offering visitors deeper knowledge while simultaneously providing entertainment. However, conventional technology has struggled to effectively maintain visitors' interest during times when animals are inactive or when they are far away. Therefore, new methods are needed to provide visitors with consistent engagement.

[0561] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0562] In this invention, the server includes means for authenticating information entered by the user, means for determining the user's location using location acquisition technology, and means for sending prompts to a generating AI model to generate movement data of a virtual object. This allows visitors to experience the movements and behaviors of virtual animals through augmented reality technology, even when animals are not actually visible, while simultaneously gaining deep knowledge and entertainment about animals in audio and text formats.

[0563] "Means of authenticating user-entered information" refers to a function that verifies the legitimacy of a user based on the authentication information provided by the user.

[0564] "Means of determining the user's location using location acquisition technology" refers to a function that accurately measures the user's current location using technologies such as GPS and Bluetooth beacons.

[0565] "Means for recognizing target objects" refers to functions that use cameras and sensors to identify surrounding objects and acquire related information.

[0566] "A means of sending prompts to a generating AI model to generate movement data for a virtual target" refers to a function that gives instructions to an AI model and generates new movement data based on past data and current information.

[0567] "Means of visual display using augmented reality technology" refers to a technology that overlays virtual information created by computer graphics onto images of the real world.

[0568] "Means of providing information in audio and text format" refers to a function that transmits necessary information to users through audio or text data.

[0569] This invention provides an augmented reality system that allows visitors to virtually experience the movements and information of animals and exhibits that they cannot see in the real world, in zoos and museums. The system mainly consists of a visual display device worn by the user, a location acquisition device installed in the zoo, and a server in the cloud.

[0570] The device receives the user's authentication information upon their arrival at the zoo. The user's login information is sent to the server via the device, verifying their legitimacy. This prepares the system to provide a personalized experience based on past visit data.

[0571] When the user approaches the animal area, the device accurately measures their location using its built-in GPS and Bluetooth beacons. It also recognizes surrounding animals using its camera and sensors. This information is transmitted to a server via the device for data processing. The server then sends prompts to a generating AI model to generate movement data for virtual animals. An example of a specific prompt might be, "Generate the movements of a virtual animal based on the behavior patterns of a lion."

[0572] The device receives movement data of the generated virtual animal and projects the virtual animal into the user's field of view using augmented reality technology. This allows the user to experience the animal's movements in real time. Furthermore, the server transmits detailed information about the animal to the device in audio and text formats, providing it to the user. This allows the user to enjoy a highly entertaining experience while learning about the animal's ecology and characteristics.

[0573] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0574] Step 1:

[0575] When a user arrives at the zoo, the terminal receives login information from the user. This login information is sent to the server, which uses this information to authenticate the user. The user is entered with a user ID and password, and after authentication, the server retrieves the user's past visit data from the database. This process prepares the terminal to provide visitors with personalized experience information.

[0576] Step 2:

[0577] When a user approaches an animal area, the device uses GPS and Bluetooth beacons to determine the user's location. It then activates its camera to recognize animals within its field of view. The acquired location information and camera image data are used, and the device sends this data to a server. Based on this information, the server processes the virtual movement of animals, taking into account the user's current location and the surrounding environment.

[0578] Step 3:

[0579] The server uses the received location information and animal recognition data to send prompt messages to the generating AI model to generate movement data for a virtual animal. Specifically, prompts such as "Generate the movements of a virtual animal based on the behavior patterns of a lion" are used. The AI ​​model combines past movement data with real-time location information to output the movements of the virtual animal.

[0580] Step 4:

[0581] The device receives virtual animal movement data transmitted from the server. Based on this data, it uses augmented reality technology to project the virtual animal into the user's field of view. The input is virtual animal movement data, which the device converts into visual information and displays on the screen in real time.

[0582] Step 5:

[0583] The server transmits descriptive information about animals to the device in audio and text formats. The generated information includes details about the animals' ecology, characteristics, and conservation status. The device plays this information as an audio guide or displays it on the screen as text. Through this information, users can deepen their knowledge about animals and enjoy a sustained entertainment experience.

[0584] (Application Example 1)

[0585] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0586] The aim is to solve the problem that consumers find it difficult to have an efficient and personalized purchasing experience due to the time-consuming process of trying on clothes and searching for product information.

[0587] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0588] In this invention, the server includes a visual display means using augmented reality technology to virtually recreate a product display, an information distribution means to provide product information to the user, and a control means to switch content based on the user's location and product selection results. This enables consumers to quickly and individually optimize product selection and information acquisition.

[0589] "Product display" refers to a means of presenting products in a physical or virtual way, creating a situation where consumers can visually examine the products.

[0590] Augmented reality technology is a technique that overlays computer graphics and digital information onto the real world, and is used to enrich the user's visual experience.

[0591] "Visual display means" refers to a method or apparatus provided for a user to obtain information visually via a device, and may utilize augmented reality.

[0592] "Product information" refers to information about a product's features, specifications, price, usage instructions, and design, and is provided to assist users in making purchasing decisions.

[0593] "Information distribution means" refers to methods or devices used to transmit predetermined information to specific recipients, and these methods may be carried out using audio or text format.

[0594] "User location" refers to the geographical or spatial location where the user is currently located, as identified using technologies such as GPS and beacons.

[0595] "Product selection results" refer to information about the products selected by consumers and the decisions they made regarding those selections. This data is used by the system to control its next actions.

[0596] "Control means" refers to a device or method provided for a system to switch or adjust its operation based on predetermined standards or conditions.

[0597] A "virtual product" is a product that does not exist in reality but is reproduced on a computer screen using digital technology, and can be visually confirmed by consumers.

[0598] "Generative means" refers to processes and techniques used to create new data or information based on input data.

[0599] The invention will now be described in terms of embodiments. The system of this invention provides consumers with a real-time experience in which they can virtually try on products and obtain product information using augmented reality technology.

[0600] The main components of the system are a server, user terminals (e.g., smart glasses or smartphones), and beacon devices installed in stores. The server resides in the cloud and is responsible for generating and delivering personalized product information to users within the store using a generative AI model.

[0601] The server generates virtual products using a generative AI model based on the user's past purchase data. The generated data is sent to the user's terminal via a visual display device, allowing the user to experience virtual try-on using augmented reality technology.

[0602] When a user arrives at the store, a beacon device identifies the user's location and transmits this information to a server. Based on this information, the server controls the content displayed to the user based on their location and the products they have selected.

[0603] As a concrete example, when a user requests a virtual try-on of jeans at a store, the server sends a prompt to a generation AI model to generate data on jeans of the appropriate size and style. For instance, a prompt might say, "Generate jeans recommended for him / her based on their past purchase history." This allows the user to efficiently select the most suitable jeans.

[0604] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0605] Step 1:

[0606] The server detects when a user arrives at a store using signals from beacon devices. Based on this input, the server identifies the user's location and retrieves their past purchase history from the database. This completes the preparation of the user information.

[0607] Step 2:

[0608] The server analyzes the user's purchase history and creates prompts for the generating AI model. These prompts are sent to the generating AI model to generate product information suitable for the user. An example of a prompt is, "Based on the user's past purchase history, please generate jeans that are recommended for him / her."

[0609] Step 3:

[0610] The generative AI model performs data calculations based on prompt messages received from the server to generate virtual product data. The generated product data includes size, style, and color information. This data is then ready for use in subsequent processing.

[0611] Step 4:

[0612] The server transmits information about the generated products to the user's terminal via a visual display. The terminal uses AR technology to provide the user with a virtual try-on interface. The user can then check the appearance and fit of the virtually tried-on products.

[0613] Step 5:

[0614] As the user operates the terminal and selects or adjusts virtual products, the server updates the information and product suggestions to be provided next based on the user's selection. This control mechanism allows the content to dynamically change according to the user's needs.

[0615] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0616] This invention proposes a system that combines augmented reality technology and emotion recognition technology to make the zoo experience more interactive and personalized. The system consists of a visual device (terminal) worn by the user, various sensing devices, and a processing server in the cloud.

[0617] When a user visits the zoo, they first put on a device and log into the system after going through an authentication process. The device uses a built-in camera and microphone to understand the user's emotional state, activating an emotion recognition engine. To eliminate ambiguity, multiple indicators (such as facial expressions and tone of voice) are used to analyze emotions.

[0618] The results of the emotion analysis are sent to the server, which selects the most appropriate content based on the user's current emotional state. For example, if the user is relaxed, videos that recreate the leisurely behavior of animals will be selected. On the other hand, if the user is feeling happy or excited, content that emphasizes the active movements of animals will be provided.

[0619] The device uses augmented reality (AR) technology to display virtual animals in the user's field of view based on data received from the server. Users can experience the animals' behavior in real time, and receive emotionally responsive audio guides and text information, making their zoo experience even more engaging.

[0620] For example, if the emotion engine detects that a user in the lion area is feeling tense, the device can recreate the image of a lion lying calmly and play soothing background music. This allows the user to relax and enjoy their time at the zoo. This system, with its emotion recognition engine, can enhance educational value by providing diverse information about animals while also being attentive to the visitor's emotions.

[0621] The following describes the processing flow.

[0622] Step 1:

[0623] The user activates the ZOO glasses and enters their login information. The device sends the entered login information to the server for user authentication. If authentication is successful, the server sends past visit history and personal settings information to the device.

[0624] Step 2:

[0625] When a user approaches the animal area, the device uses its built-in GPS or Bluetooth beacon to pinpoint the user's location. Simultaneously, it activates its built-in camera to capture the user's facial expressions and uses its microphone to record their voice tone.

[0626] Step 3:

[0627] The device inputs captured facial expression and audio data into an emotion recognition engine. The emotion recognition engine analyzes this data to identify the user's emotional state. The analysis results are sent to the server in real time.

[0628] Step 4:

[0629] The server selects appropriate animal behavior content based on the user's emotional state. For example, if the server determines that the user is stressed, it will select calm animal movements and relaxing background music. The server then sends the selected content to the device.

[0630] Step 5:

[0631] The device uses content data received from the server to display virtual animals overlaid on real-world scenery using augmented reality (AR) technology. Users can experience animal behavior optimized in response to their emotions in real time.

[0632] Step 6:

[0633] Furthermore, the device uses audio and text information to provide users with interesting information about animals. This allows users to gain a deeper understanding of animals while simultaneously enjoying a visual experience.

[0634] Step 7:

[0635] When a user finishes their zoo visit, the device sends the user's activity history and emotional data to a server. The server stores this data and uses it to further customize the experience for their next visit. The user then logs out of the system and shuts down the device.

[0636] (Example 2)

[0637] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0638] Zoos and similar facilities face the challenge of providing visitors with interactive and personalized experiences. In particular, there is a need to offer experiences that foster deeper understanding and satisfaction by providing information tailored to the visitor's emotional state.

[0639] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0640] In this invention, the server includes means for analyzing the user's emotional state and selecting optimal content based on that state, means for displaying the selected content using augmented reality technology, and means for providing information in audio and text formats. This makes it possible to provide visitors with an interactive and personalized experience that responds to their emotions.

[0641] "Means for analyzing a user's emotional state" refers to a technological system that acquires data such as a user's facial expressions and tone of voice, analyzes it, and identifies the user's current emotional state.

[0642] "Means for selecting optimal content based on emotional state" refers to an algorithm or system that selects the most appropriate content to provide based on the analyzed emotional state of the user.

[0643] "Means of displaying using augmented reality technology" refers to a technology or device that uses AR technology to display selected content superimposed on the real world within the user's field of view.

[0644] "Means of providing information in audio and text format" refers to systems that provide users with detailed information about animals through audio guides or text messages.

[0645] A "processing device" is a computer-based device designed to process large amounts of data quickly, and cloud servers are an example of such devices.

[0646] A "generative AI model" refers to a set of technologies, including generative battle networks and natural language processing, that have algorithms to provide optimal content based on user sentiment data.

[0647] This invention is a technological system for personalizing and making the zoo experience interactive. Users wear a visual device, a terminal, when visiting the zoo. This terminal has a built-in camera and microphone that captures the user's facial expressions and voice in real time.

[0648] The device's emotion recognition engine analyzes acquired data to identify the user's emotional state. This analysis utilizes image processing and voice analysis technologies. The analysis results are sent to a server, where a generative AI model is used to select the most appropriate content. The server operates on the cloud, enabling fast and secure data processing.

[0649] The selected content is sent to the device. The device uses augmented reality technology to display virtual animals and environments in the user's field of view. This display allows the user to experience the animals' behavior in real time and receive emotionally responsive audio guides and text information. For example, if the emotion recognition engine detects the user's tension, the device can display relaxed animals and play calming music.

[0650] An example of a prompt used in a generative AI model is the phrase, "Please specify animal content recommended for when the user is in a relaxed state." This prompt allows the server to generate and provide appropriate content to the user.

[0651] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0652] Step 1:

[0653] When a user puts on the visual device, the terminal first performs authentication. The inputs used are the user's facial image captured by the built-in camera and their voice recorded by the microphone. For data processing, a facial recognition algorithm and voice recognition software are used to log the user in. The output is the authentication success or failure status. Specifically, authentication is completed within a few seconds of the user putting on the device.

[0654] Step 2:

[0655] The device collects data using its built-in camera and microphone to analyze the user's emotional state. The input consists of image data showing the user's facial expressions and audio data showing their voice tone. For data processing, an emotion recognition engine analyzes this data to identify the user's current emotional state. The output is information about the analyzed emotional state, specifically states such as "relaxed," "excited," and "stressed."

[0656] Step 3:

[0657] The device sends the analyzed emotional state to the server. The emotional state information obtained in the previous step is used as input. A data transmission protocol is used to transmit the data to the cloud server in real time. The output is a confirmation notification that the server has received the emotional state information. Specifically, the data is transmitted encrypted, and confirmation is received within a few milliseconds.

[0658] Step 4:

[0659] The server runs a generative AI model based on the received emotional state to select the most suitable content. The inputs used are the user's emotional state information and the prompt "Please specify animal content recommended when the user is relaxed." During data calculation, the AI ​​algorithm compares the data against a historical database to select appropriate animal content. The output is the data of the selected content. In practice, the selection process is completed on the server within a few seconds.

[0660] Step 5:

[0661] The server sends the selected content to the terminal. Optimized content data is used as input. The data is delivered quickly and securely to the terminal via a transmission protocol. The output is the content data received by the terminal. Specifically, the content is compressed before transmission and decompressed after reception.

[0662] Step 6:

[0663] The device displays received content using augmented reality technology. It uses content data sent from a server as input. AR software processes the data, and a virtual animal appears in the user's field of view. The output is interactive content that the user can experience visually and aurally. Specifically, the virtual animal appears in the user's field of view, and an audio guide plays in accordance with its movements.

[0664] (Application Example 2)

[0665] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0666] Traditional zoo experiences have struggled to provide personalized content tailored to the individual emotions and interests of visitors. As a result, the information and visual enjoyment visitors experience are limited, creating a need for greater engagement. In particular, a system is needed that dynamically selects and delivers content based on the emotional state of each visitor.

[0667] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0668] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, information provision means for providing optimized animal content based on the analysis results, and visual display means for displaying virtual animals in the user's field of view using augmented reality technology. This makes it possible to provide a personalized zoo experience tailored to the emotional state of visitors.

[0669] An "emotion analysis device" is a technological device that analyzes a user's facial expressions and voice data to identify their emotional state in real time.

[0670] "Information provision means" refers to a technological device that provides animal-related content in audio or text format, tailored to the user's needs, based on the results of sentiment analysis.

[0671] A "visual display means" is a technological device that utilizes augmented reality technology to display the movements of animals in a virtual environment in real time within the user's field of view.

[0672] The embodiment of this invention consists mainly of a user-worn device, a cloud server for emotion analysis, and a data display system utilizing augmented reality (AR) technology. The system operates according to the following procedure.

[0673] When a user wears the device, it uses its camera and microphone to capture the user's facial expressions and voice in real time. This captured data is analyzed through a cloud-based emotion recognition API (e.g., Google Cloud Vision or Microsoft Azure Emotion API) as a means of emotion analysis. The results of the emotion analysis are sent to a server.

[0674] The server operates an information delivery system that selects animal content appropriate for the user based on the results of sentiment analysis. This content includes videos demonstrating animal behavior, related audio guides, and text information, and is optimized for the user's emotional state.

[0675] The selected content is displayed on the device via augmented reality technology. The visual display means shows virtual animals to the user's vision, enabling them to interactively experience the content.

[0676] As a concrete example, consider a scenario where a user is experiencing a virtual tour of a wildlife park. If the emotion analysis system detects that the user is relaxed, the server can select a video of an elephant leisurely walking across a vast grassland and present the user with calming music.

[0677] Examples of prompts include, "If the user is excited, select content showing active animal behaviors that are appropriate for them," and "For a relaxed user, present calming animal behaviors."

[0678] This allows users to enjoy a unique and immersive zoo experience based on their own emotional state.

[0679] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0680] Step 1:

[0681] The device captures the user's facial expressions and voice using its camera and microphone. The input is real-time video and audio, which is collected as digital data and prepared to be sent to an emotion analysis API.

[0682] Step 2:

[0683] Facial expression and audio data transmitted from the device are received by a server in the cloud. The server uses an emotion recognition API to analyze this data and determine the user's emotional state. The output is an indicator of the user's emotional state.

[0684] Step 3:

[0685] The server selects appropriate animal content based on the user's emotional state. This process searches a database to select the most suitable video, audio guide, and text information for that emotion. The input is an indicator of the emotional state, and the output is the selected content data.

[0686] Step 4:

[0687] The server transmits the selected animal content to the terminal. This includes appropriate augmented reality video data and associated audio data. The terminal receives this information and prepares it for use with its visual display device.

[0688] Step 5:

[0689] The device displays received animal content using augmented reality technology. Virtual animals are projected into the user's field of view, allowing for an interactive experience of the content. The input is content data from the server, and the output is an AR display for the user.

[0690] Step 6:

[0691] Users experience the displayed augmented reality animal content and gain further knowledge about the animals based on audio guides and text information. This process allows users to enjoy a personalized zoo experience tailored to their own emotions.

[0692] This process allows the entire system to work together to provide the user with the best possible experience.

[0693] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0694] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0695] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0696] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0697] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0698] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0699] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0700] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0701] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0702] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0703] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0704] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0705] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0706] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0707] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0708] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0709] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0710] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0711] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0712] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0713] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0714] The following is further disclosed regarding the embodiments described above.

[0715] (Claim 1)

[0716] A visual display means using augmented reality technology to virtually reproduce animal behavior,

[0717] An information distribution method that provides users with characteristic information about animals,

[0718] A control means for switching content based on the user's location and the results of animal recognition,

[0719] A system that includes this.

[0720] (Claim 2)

[0721] The system according to claim 1, wherein the visual display means includes a generation means for generating virtual animals based on past animal behavior data.

[0722] (Claim 3)

[0723] The system according to claim 1, wherein the information distribution means provides information about animals in audio and text format.

[0724] "Example 1"

[0725] (Claim 1)

[0726] A means of authenticating the information entered by the user,

[0727] A means of determining the user's location using location acquisition technology,

[0728] Means for recognizing an object,

[0729] A means of sending prompts to a generating AI model to generate motion data of a virtual object,

[0730] A means of performing visual displays using augmented reality technology,

[0731] Means for providing information in audio and text formats,

[0732] A system that includes this.

[0733] (Claim 2)

[0734] The system according to claim 1, wherein the generation means includes means for generating a virtual target based on past operation data.

[0735] (Claim 3)

[0736] The system according to claim 1, wherein the information provision means provides users with an educational and entertaining experience.

[0737] "Application Example 1"

[0738] (Claim 1)

[0739] A visual display means using augmented reality technology to virtually recreate product displays,

[0740] A means of providing product information to users,

[0741] A control means for switching content based on the user's location and product selection results,

[0742] A system that includes this.

[0743] (Claim 2)

[0744] The system according to claim 1, wherein the visual display means includes a generation means for generating virtual products based on past purchase data.

[0745] (Claim 3)

[0746] The system according to claim 1, wherein the information distribution means provides information regarding the characteristics of the product in audio and text format.

[0747] "Example 2 of combining an emotion engine"

[0748] (Claim 1)

[0749] A means of analyzing the user's emotional state and selecting the optimal content based on that emotional state,

[0750] A means of displaying selected content using augmented reality technology,

[0751] Means for providing information in audio and text formats,

[0752] A system that includes this.

[0753] (Claim 2)

[0754] The system according to claim 1, wherein the display means has a generation AI model that selects optimal content using a processing device that operates on the cloud.

[0755] (Claim 3)

[0756] The system according to claim 1, wherein the information providing means provides audio guides and text information corresponding to emotional states.

[0757] "Application example 2 when combining with an emotional engine"

[0758] (Claim 1)

[0759] A means of analyzing the emotional state of a user,

[0760] An information provision method that provides optimized animal content based on analysis results,

[0761] A visual display means that uses augmented reality technology to display virtual animals in the user's field of view,

[0762] A system that includes this.

[0763] (Claim 2)

[0764] The system according to claim 1, wherein the emotion analysis means analyzes emotions using facial expression and voice data.

[0765] (Claim 3)

[0766] The system according to claim 1, wherein the information providing means generates voice guidance and text based on emotional state. [Explanation of Symbols]

[0767] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A visual display means using augmented reality technology to virtually reproduce animal behavior, An information distribution method that provides users with characteristic information about animals, A control means for switching content based on the user's location and the results of animal recognition, A system that includes this.

2. The system according to claim 1, wherein the visual display means includes a generation means for generating a virtual animal based on past animal behavior data.

3. The system according to claim 1, wherein the information distribution means provides information about animals in audio and text format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A