system

The system addresses the lack of real-time feedback and natural animations in virtual animal interactions by using a generative AI model to provide immediate and intuitive responses, enhancing user engagement and educational value.

JP2026068489APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing virtual animal interaction systems lack real-time user operation feedback and natural animations, leading to reduced educational effectiveness and a less immersive experience.

Method used

A system that includes means for displaying virtual animals, sensing user input, analyzing operation data, and generating natural animations using a generative AI model to provide immediate and intuitive responses.

Benefits of technology

Enables real-time, natural, and educational interactions with virtual animals, enhancing user engagement and educational value through seamless and realistic movements and expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068489000001_ABST
    Figure 2026068489000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of displaying virtual animals, Operation data acquisition means that senses user operations and generates data, An analysis means that analyzes operational data and generates the response of a virtual animal, An animation generation means that generates an animation of a virtual animal based on the generated reaction, A display control means for sending and displaying animations on a user terminal, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Safe interaction with virtual animals is highly educational, but the prior art lacks real-time user operation and its feedback, and there are limitations in the smoothness of animations and the generation of natural movements. As a result, there is a problem that it is difficult for users to actually feel as if they are communicating with animals, and the educational effect is reduced.

Means for Solving the Problems

[0005] This invention enables immediate data analysis and response generation based on user input by using means for displaying virtual animals and means for sensing user input and acquiring operation data. Furthermore, by including animation generation means that seamlessly generate animal movements and facial expressions using an image generation engine, it enables real-time and natural feedback to the user, effectively solving the problem.

[0006] A "virtual animal" is an animal that is recreated on a computer using digital technology, and it is something that users can learn from and enjoy through interaction.

[0007] "Display means" refers to a device or software that has the function of displaying a virtual animal or its actions on an output device so that the user can visually recognize it.

[0008] "Operation data acquisition means" refers to a device or software function that senses information about touches and actions entered by the user and acquires it as data.

[0009] "Analysis means" refers to the function of a device or software that performs processing to derive the response of a virtual animal based on acquired operational data.

[0010] "Animation generation means" refers to a device or software function that generates continuous changes in the movements and facial expressions of a virtual animal as image data and expresses them as smooth motion.

[0011] "Display control means" refers to a device or software function that appropriately transmits generated animation data to a user terminal and controls it so that the user can display it.

[0012] "Real-time" refers to a processing speed that produces a nearly instantaneous response to user actions, and is a characteristic that provides a seamless user experience. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention relates to a system for safely and educationally facilitating interaction with virtual animals. This system allows users to interact with virtual animals in real time by operating a terminal. Specific embodiments of the invention are described below.

[0035] The user first launches the application on their device and accesses the virtual safari park screen. Here, a list of animals the user can choose from is displayed, and if they select, for example, a lion, a virtual image of that lion will be displayed on their device.

[0036] Users interact with the lion by touching or swiping on the device's touchscreen. These user actions are detected by the device and collected as operation data. This operation data includes information such as the location of the touch, the speed of the movement, and the direction.

[0037] The terminal sends the acquired operation data to the server. The server analyzes this data and generates the appropriate responses for the animal in real time. This process utilizes a generative AI model, which assigns natural and intuitive movements and expressions to the virtual animal in response to user input.

[0038] Based on the analysis results, the server generates animations that represent the animals' reactions. Image generation technology is used to smooth the movements, providing a more realistic experience. The system also generates ambient sounds and sound effects of the animals' movements, creating a more immersive experience for the user.

[0039] The generated animation data is immediately transmitted to the terminal via the network. The terminal receives this data and displays the movements of the virtual animal in real time in response to the user's actions. As a result, the user can enjoy an interactive learning experience powered by the latest technology.

[0040] For example, when a user performs an action such as stroking the lion's head, the server generates an action where the lion narrows its eyes in apparent pleasure and lightly wags its tail. This result is immediately displayed on the device, allowing the user to enjoy interacting with the lion.

[0041] In this way, the system of the present invention can provide an interactive experience with virtual animals that has educational value and is highly safe.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user launches the application on their device and accesses the virtual safari park environment. A list of animals is displayed on the screen, and the user selects an animal they are interested in.

[0045] Step 2:

[0046] The device receives the user's touch input and displays an image of the selected animal (for example, a lion) on the screen. At this point, the user can begin interacting with the animal, such as touching or stroking it.

[0047] Step 3:

[0048] User actions are detected via the device's touchscreen and captured as operation data. This data includes the touch location, amount of movement, direction, and speed of movement.

[0049] Step 4:

[0050] The terminal transmits detected operation data to the server in real time. The data includes the user's ID, the animal's ID, and detailed information about the operation.

[0051] Step 5:

[0052] The server analyzes the received operation data. Using a generative AI model, it determines the animal's response (movement and facial expression) based on the acquired information. For example, if the user performs an action of stroking the lion's head, it will generate a response in which the lion narrows its eyes in apparent pleasure.

[0053] Step 6:

[0054] The server creates animations corresponding to the reactions of the generated animals. This uses image generation technology to produce smooth movements and realistic facial expressions.

[0055] Step 7:

[0056] The generated animation data is compressed and sent to the device in the optimal format. This enables real-time interaction with minimal latency.

[0057] Step 8:

[0058] The device unpacks the received animation data and displays it on the user's screen in real time. The user can then see the animal's reaction in response to their actions, allowing for a more interactive experience.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] There is a growing need to provide more natural, safe, and educational interactions with virtual beings. However, conventional technologies suffer from a lack of real-time capabilities, resulting in delayed responses to user input or unnatural movements and expressions. It is necessary to overcome these challenges and provide users with a realistic experience.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for acquiring operation information that senses user operations and generates information, means for analyzing operation information and generating responses of a virtual organism, and means for generating natural movements and facial expressions using a generative model. This makes it possible to generate immediate and natural responses to user operations and provide the user with a real-time, immersive interaction.

[0064] A "virtual organism" refers to a virtual animal or living being generated on a computer system, which a user can interact with through an interface.

[0065] "Display means" refers to a device or mechanism for outputting images or animations of virtual creatures so that users can visually recognize them.

[0066] "Operation information acquisition means" refers to a device or software that detects operations performed by a user and acquires that information as digital data.

[0067] "Analysis means" refers to a device or program that analyzes acquired operational information and performs processing to determine an appropriate response from a virtual organism based on that analysis.

[0068] "Motion generation means" refers to a mechanism or software that generates the movements and behaviors of a virtual organism based on the reactions determined by the analysis means.

[0069] "Display control means" refers to a device or software for playing back the movements and animations of a generated virtual creature and presenting them effectively to the user.

[0070] A "generative model" refers to a mathematical model or algorithm that uses machine learning or artificial intelligence techniques to generate natural movements and facial expressions for virtual beings.

[0071] "Information and communication network" refers to a network structure used to send and receive digital data, and includes the internet and dedicated lines.

[0072] This invention relates to a system for realizing interaction with virtual creatures. The user operates a terminal to generate real-time responses from the virtual animal. Specific embodiments of this system are described below.

[0073] The user first launches an application installed on their device. This displays a virtual safari park on the screen, allowing the user to select their favorite virtual animal. For example, if the user selects a lion, a virtual image of that lion will be displayed on the device's screen.

[0074] When a user touches or swipes an animal using the device's touchscreen, these actions are detected, and the device acquires action data such as the location, speed, and direction of the touch. This action data is transmitted to the server in real time.

[0075] The server is responsible for analyzing the received operation data. This analysis uses a generative AI model to generate natural movements of a virtual animal in response to user actions. Specifically, the generative AI model uses the prompt "Generate the lion's reaction when it is petted" to generate the lion's movements and facial expressions.

[0076] Based on the analysis results, the server animates the virtual animal's reactions and sends the generated data to the terminal. The terminal receives this data and displays the virtual lion's movements and expressions in real time in conjunction with the user's actions. This process allows the user to become familiar with the virtual lion and enjoy an interactive experience.

[0077] For example, if a user gently strokes the lion's ears, the server generates a motion where the lion happily wags its tail, and displays the result on the terminal. In this way, the virtual creature responds immediately to the user's actions, aiming to provide a more realistic and educational experience.

[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0079] Step 1:

[0080] The user launches an application on their device. This displays a virtual safari park on the screen, and a screen for selecting an animal appears. When the user selects a lion, its virtual image is displayed on the device's screen. The input is the user's selection information, and the output is the virtual image of the selected animal.

[0081] Step 2:

[0082] The user touches or swipes on the selected animal on the device's touchscreen. The device detects these actions and generates action data, which includes information such as the position, speed, and direction of the action. The input is the user's touch actions, and the output is the action data.

[0083] Step 3:

[0084] The terminal sends the generated operation data to the server. The transmitted data includes details such as the location, speed, and direction of the touch. The input is the operation data, and the output is the transmission of data to the server.

[0085] Step 4:

[0086] The server analyzes the received operation data. This analysis process uses a generative AI model to determine the animal's response using prompt statements. For example, based on the prompt statement "Generate the lion's response when petted," it generates natural lion behavior. The input is operation data and prompt statements, and the output is the lion's response data.

[0087] Step 5:

[0088] The server generates animations of virtual animals based on the analysis results. The generated motion data is created by the server along with ambient sounds and sound effects. For example, it might add the actions of a lion squinting and wagging its tail, along with a purring sound. The input is reaction data, and the output is animation data and sound data.

[0089] Step 6:

[0090] The server sends the generated animation and audio data to the terminal. The terminal receives this in real time and displays the movements of the virtual lion in response to user input. The input is animation and audio data, and the output is a real-time display of the virtual animal.

[0091] (Application Example 1)

[0092] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0093] Conventional virtual experience systems have limited interaction with virtual creatures, making it difficult to provide users with sufficient educational value or effective guidance. Furthermore, the insufficient real-time motion generation resulted in unnatural interactions.

[0094] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0095] In this invention, the server includes operation data acquisition means that senses user operations and generates data, analysis means that analyzes the operation data and generates responses from the virtual creature, and guide generation means that provides guidance from the virtual creature based on information related to the user's environment. This enables the user to receive real-time guidance from the virtual creature through intuitive and flexible interaction with educational value.

[0096] A "virtual organism" is an electronic living entity generated by a computer system, which is displayed as an object that the user interacts with.

[0097] "Display means" refers to a function or device for displaying images and actions of virtual creatures on a user's terminal.

[0098] "User operation" refers to input operations and instructions performed by the user via a terminal, including those detected using touchscreens or sensors.

[0099] "Operation data acquisition means" refers to a function or device for sensing user operations and collecting and generating data based on those operations.

[0100] "Analysis means" refers to a data analysis function or device for generating and determining the responses and actions of a virtual organism in real time based on acquired operational data.

[0101] "Motion generation means" refers to a function or device that generates movements based on the analyzed results in order to give continuity and naturalness to the movements and facial expressions of a virtual organism.

[0102] "Display control means" refers to a control function or device that transmits generated actions or responses to a user terminal and displays them.

[0103] A "guide generation means" is a function or device that uses the user's environmental information to generate guide information for a virtual creature to provide appropriate guidance and explanations.

[0104] The program for the system to implement this invention aims to realize real-time interaction between a virtual organism and a user. This system mainly consists of three elements: a server, a terminal, and a user.

[0105] First, it is assumed that users will use smart glasses or head-mounted displays as smart devices. Users will wear these devices and interact with virtual creatures in a virtual space. The terminals have built-in display means and are equipped with means for acquiring operation data that receives user input. Through this operation, users can touch the virtual creatures and give them instructions.

[0106] The server analyzes user operation data received via operation data acquisition means using a generative AI model. This analysis process determines the virtual creature's response to specific user actions. Based on the analysis results, the server creates corresponding animations and motion data using motion generation means that generate the virtual creature's movements and expressions. This process utilizes a media generation engine to enable natural and smooth movements.

[0107] The generated motion data is sent to the terminal by the display control means and presented to the user in real time. This allows the user to experience intuitive and seamless interaction with the virtual creature. Furthermore, the server includes a guide generation means that analyzes the user's environment information and provides guidance to the user based on that information. This function allows the virtual creature to provide appropriate advice and explanations in real time according to the user's surroundings.

[0108] As a concrete example, if a user interacts with a virtual lion to learn about a specific product, the server will generate actions for the lion to provide a concise and easy-to-understand explanation of that product. The lion will guide the user in a friendly tone, saying, "This product is made from environmentally friendly materials and is a sustainable choice." An example of a prompt would be, "The user has shown interest in a new eco-bag. As the lion guide, what will you tell the user about the product?"

[0109] This allows users to enjoy a rich virtual experience that goes beyond mere visual enjoyment, enabling them to actually learn.

[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0111] Step 1:

[0112] Users access a virtual space using smart devices. Users wear smart glasses or head-mounted displays and interact with virtual beings. The terminal receives user input through an operation data acquisition mechanism, obtaining user actions and location information as input. This data is transmitted to the server.

[0113] Step 2:

[0114] The server analyzes the user's input data using a generated AI model. This analysis process involves data calculations to determine the virtual creature's response to specific user actions. For example, the touch location and direction of movement are input to the AI ​​model as prompts, and the corresponding response from the virtual creature is generated as output.

[0115] Step 3:

[0116] Based on the analysis results, the server uses motion generation tools to generate the movements and facial expressions of the virtual creature. A media generation engine is then used to create smooth and realistic animations. The generated animation data is provided as output from the motion generation tools.

[0117] Step 4:

[0118] The generated motion data and animations are sent to the terminal and presented to the user in real time by a display control system. Based on the received motion data, the terminal allows the user to observe the movements of the virtual creature through the display.

[0119] Step 5:

[0120] The server further analyzes the user's environmental information and uses a guide generation mechanism to enable the virtual creature to provide appropriate guidance to the user. For example, if a user shows interest in a particular product, that information is input as environmental data, and a guide prompt is generated. This prompt is input into a generating AI model, and the virtual creature's guide information is generated as output.

[0121] Step 6:

[0122] Ultimately, the generated guide information and actions are displayed on the terminal, allowing the user to receive real-time guidance and explanations from the virtual creature. Through this, the user can choose to continue the operation or make a purchase decision.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention relates to a system that, in addition to interaction with a virtual animal, recognizes the user's emotions and provides a personalized experience. This system can provide a more interactive and personalized learning experience by having the virtual animal respond in a way that takes the user's emotional state into account.

[0125] By launching an application on their device, users can access a virtual safari park and select one of the displayed animals. For example, if a lion is selected, a lion will appear on the screen. At this point, the user can begin interacting with the animal through touch controls.

[0126] The device uses cameras and microphones to recognize not only user actions but also the user's facial expressions and tone of voice in real time. This data is analyzed via an emotion engine to estimate the user's emotional state. This information is sent to the server along with the operation data.

[0127] The server generates virtual animal responses that correspond to the user's emotions, based on acquired operation data and estimated emotional information. For example, if the user is smiling, the lion can perform friendly actions, and if the user is feeling down, it can perform encouraging actions.

[0128] Next, the server generates an animation based on the appropriate animal response. Image generation technology is used for this generation, providing continuous and realistic motion. The animation data is compressed and sent to the terminal.

[0129] The device decompresses the transmitted animation data and displays it on the user's screen in real time. This allows users to interact with animals that react differently to their emotions, resulting in a more personalized experience.

[0130] For example, if the emotion engine determines that the user is happy as they touch the lion, the server will generate an animation in which the lion moves in a more friendly manner and behaves in a welcoming manner towards the user. Through this process, the user can enjoy a unique interaction experience.

[0131] This invention makes it possible to provide users with a more empathetic and intuitive learning environment by incorporating emotion recognition technology into their interactions with virtual animals.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] The user launches an application on their device and accesses the virtual safari park environment. A list of animals is displayed on the device, and the user selects a virtual animal they are interested in.

[0135] Step 2:

[0136] The device displays the selected virtual animal on the screen. The user begins interacting with the virtual animal on the touchscreen and performs actions. At the same time, the device also acquires data on the user's facial expressions and voice.

[0137] Step 3:

[0138] The device collects user operation data and simultaneously incorporates facial expression and voice data acquired through the camera and microphone into the emotion engine. This allows the user's emotional state to be analyzed.

[0139] Step 4:

[0140] The terminal combines operational data and analyzed emotional data, and sends this to the server. The server analyzes the received information and determines the virtual animal's response based on the user's emotions.

[0141] Step 5:

[0142] The server generates responses that adjust the actions and facial expressions of the virtual animal based on the user's emotional state. For example, if the user seems happy, the lion will perform playful actions, while if the user seems sad, it will generate encouraging actions.

[0143] Step 6:

[0144] Based on the generated animal reactions, the server creates a continuous animation. High-quality, realistic animations are produced using image generation technology.

[0145] Step 7:

[0146] The server sends the generated animation data to the terminal. Since the data is transmitted in real time, the user can receive immediate feedback.

[0147] Step 8:

[0148] The device decompresses the received animation data and displays it immediately on the screen. Users can directly see the virtual animal's reactions in response to their own actions and emotions, resulting in a more empathetic experience.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0151] Current virtual animal interaction systems have limitations in their response to user input and are unable to provide personalized experiences that incorporate emotional recognition. Therefore, there is a need to provide more personalized interactions that take user emotions into consideration.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for displaying a virtual animal, means for acquiring data by sensing user actions and emotions and generating data, and means for analyzing the acquired action data and emotion data and generating responses for the virtual animal. This allows the virtual animal to show individual responses in accordance with the user's emotions, enabling an interactive and personalized experience.

[0154] A "virtual animal" is a computer-simulated animal created by a computer program that can interact with the user.

[0155] "Display means" refers to technologies and devices for visually displaying information on a user's terminal, including displays and screens.

[0156] "Data acquisition methods" refer to technologies and techniques for detecting user actions and emotional states and collecting them as data.

[0157] "Analysis means" refers to a method or apparatus for determining the response of a virtual animal using a computer algorithm based on acquired data.

[0158] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generatively creates new content and responses.

[0159] "Animation generation means" refers to technologies and methods for creating images or videos to continuously represent the movements of a virtual animal.

[0160] "Display control means" refers to technologies and devices that transmit generated animations to the user's terminal and control them to execute them appropriately.

[0161] This invention is a system that provides a personalized experience that takes into account the user's emotions through interaction with virtual animals. Users can launch a specific application on their device, access a virtual safari, select an animal, and begin interacting with it.

[0162] Users can directly control virtual animals using a touch panel. The device uses its built-in camera and microphone to sense the user's facial expressions and voice tone, collecting emotional data in real time. This data is temporarily stored on the device and sent to the emotion engine. The emotion engine analyzes this data using machine learning models to estimate the user's emotional state. The machine learning models used include deep learning techniques.

[0163] The estimated emotion information is sent to the server along with the user's interaction data. The server uses a generative AI model to generate appropriate virtual animal responses based on this information. For example, if the user smiles, the lion is programmed to perform friendly actions. Generative AI models may utilize technologies such as natural language processing or image generation.

[0164] The server generates an animation based on the determined animal's response. The generated animation is then sent from the server to the terminal after its data size is reduced using compression technology. The terminal decompresses the received animation and displays it on the user's screen in real time.

[0165] For example, when a user touches a lion, if the emotion engine detects the user's smile, the server generates an animation of the lion welcoming the user. Through such interactions, users can enjoy a personalized experience that responds to their emotions in real time.

[0166] An example of a prompt might be, "Create an animation showing how a lion behaves in a friendly manner when a user touches it with a smile." Through this prompt, the generative AI model can generate natural reactions that meet the user's expectations.

[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0168] Step 1:

[0169] The user launches the application on their device and accesses a virtual safari. From the animals displayed, they select, for example, a lion. The selection is made by tapping on the touch panel. The input here is the user's selection, and the selected animal is displayed on the screen as output.

[0170] Step 2:

[0171] The device activates its built-in camera and microphone to sense the user's facial expressions and voice tone in real time. This allows for the collection of emotional data. The user's facial expressions and voice data are used as input, and emotional data is generated as output for transmission to the emotion engine.

[0172] Step 3:

[0173] The device sends the collected emotional data to the emotion engine, which analyzes the data using a machine learning model. The analysis estimates the user's emotional state. The input is emotional data, and the output is estimated emotional information. This estimated information, along with the user's interaction data, is used in the next step.

[0174] Step 4:

[0175] The server receives operation data and emotion information sent from the terminal. Using a generative AI model, it generates a virtual animal response that corresponds to the user's emotions. Specifically, if the user is smiling, it generates an action where the lion approaches in a friendly manner. The input is operation data and emotion information, and the output is virtual animal response data.

[0176] Step 5:

[0177] The server creates animations based on the generated reaction data. Image generation technology is used to achieve realistic movement. The input is the reaction data of a virtual animal, and the output is compressed animation data.

[0178] Step 6:

[0179] The device receives and decompresses compressed animation data sent from the server. It then displays the animation on the user's screen in real time. The user can enjoy interacting with an animal that responds to emotions. The input is compressed animation data, and the output is the animal's actions displayed on the user's screen.

[0180] (Application Example 2)

[0181] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0182] In modern production environments, there is a need for efficient work support that takes into account the emotions and physical condition of workers. However, conventional systems have difficulty responding appropriately to changes in workers' emotions, limiting improvements in work efficiency and safety. Therefore, there is a need for a system that can recognize workers' emotional states in real time and provide individualized and appropriate support accordingly.

[0183] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0184] In this invention, the server includes means for displaying a virtual organism, means for acquiring operation information that senses operations and generates information, and means for recognizing the user's emotional state. This makes it possible to provide appropriate work support immediately according to the user's emotional state.

[0185] A "virtual organism" refers to a living being simulated on a computer, which exhibits various reactions and actions through interaction with the user.

[0186] "Display means" refers to functions that visualize virtual creatures on a computer screen or device, and includes technologies for presenting information to the user in real time.

[0187] "Operation information acquisition means" refers to a function that detects the user's physical operations and inputs and collects and processes information generated based on them.

[0188] "Analysis means" refers to a function that includes a process for analyzing acquired operational information and emotional state information to generate appropriate responses and actions for the virtual organism.

[0189] "Visual representation generation means" refers to technology for visualizing the actions and reactions of virtual organisms and presenting them to the user as seamless and realistic images.

[0190] "Display control means" refers to a technology that has the function of appropriately transmitting the generated visual representation to the user terminal and displaying it on the screen.

[0191] "Emotion recognition means" refers to a function that detects and analyzes the user's emotional state from their facial expressions, tone of voice, etc.

[0192] "Work support means" refers to a process that includes functions to provide specific work support and interaction tailored to the user's emotional state.

[0193] This invention is a system that recognizes the emotional state of a user and provides individualized and appropriate support tailored to that user. This system utilizes virtual organisms and is designed to support users in working efficiently and safely in their work environment.

[0194] The server uses emotion recognition means to analyze data collected from the camera and microphone of the user's smart device. This analysis uses voice analysis software and facial recognition technology as examples. It detects changes in the user's facial expressions and tone of voice in real time and determines their emotional state. This information is processed via operation information acquisition means and treated as basic data for the virtual organism to show appropriate responses as needed.

[0195] The terminal receives visual representation data transmitted from the server and displays it on the user's screen. The visual representation generation means provides seamless and realistic movement of the virtual creature, which the user can visually confirm. For example, if the user shows signs of fatigue, the virtual creature displays encouraging messages or suggestions for improving work methods. This interaction allows the user to relax and increase efficiency while working.

[0196] As a concrete example, if the emotion recognition system detects signs of stress while the user is working, the server generates a message such as "We recommend taking a break to relax" and presents it to the user as a visual representation. An example of a prompt message to the generating AI model would be: "We have emotion analysis data for the user. He appears to be tired at the moment. Please generate support or guidance to prompt the system."

[0197] This makes it possible to provide users with a better work environment and personalized support.

[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0199] Step 1:

[0200] The user puts on a smart device and begins working. The device activates its camera and microphone to capture the user's facial expressions and voice data. The input consists of video and audio data, and the output is prepared for analysis of this data. The device collects this data from sensors in real time.

[0201] Step 2:

[0202] The terminal sends the collected data to the server. The server processes the received video and audio data using analysis software. The input is the raw video and audio data sent from the terminal, and the output is extracted feature data. Specifically, the server uses facial recognition technology to extract the user's facial features and voice analysis technology to analyze their tone.

[0203] Step 3:

[0204] Based on the analysis results, the server uses a generative AI model to estimate the user's emotional state. This model combines facial and vocal features to generate emotion labels. The input is feature data, and the output is an emotional state (e.g., stress, exhilaration, fatigue). The server accurately estimates the emotional state and stores it in a database for the next step.

[0205] Step 4:

[0206] The server generates interactions with the virtual creature based on its estimated emotional state. Specific actions and messages are determined by a generating AI model. Inputs are emotional states and past interaction data, while output is the virtual creature's response (e.g., soothing actions or relaxing messages). The server analyzes the user's past emotions and selects the optimal response.

[0207] Step 5:

[0208] The server visualizes the reactions of the generated virtual organism through a visual representation generation system and transmits that data to the terminal. The input is the reactions of the virtual organism, and the output is effective visual representation data. The server performs the necessary processing for visualization, compresses the data, and transmits it.

[0209] Step 6:

[0210] The terminal decompresses the received visual representation data and displays it on the user's device. The input is compressed visual representation data, and the output is the animation and interaction that the user actually sees. The terminal displays the virtual creature's actions on the home screen and prepares to take further action information to check the user's response.

[0211] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0212] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0213] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0214] [Second Embodiment]

[0215] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0216] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0217] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0218] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0219] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0221] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0222] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0223] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0224] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0225] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0226] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0227] This invention relates to a system for safely and educationally facilitating interaction with virtual animals. This system allows users to interact with virtual animals in real time by operating a terminal. Specific embodiments of the invention are described below.

[0228] The user first launches the application on their device and accesses the virtual safari park screen. Here, a list of animals the user can choose from is displayed, and if they select, for example, a lion, a virtual image of that lion will be displayed on their device.

[0229] Users interact with the lion by touching or swiping on the device's touchscreen. These user actions are detected by the device and collected as operation data. This operation data includes information such as the location of the touch, the speed of the movement, and the direction.

[0230] The terminal sends the acquired operation data to the server. The server analyzes this data and generates the appropriate responses for the animal in real time. This process utilizes a generative AI model, which assigns natural and intuitive movements and expressions to the virtual animal in response to user input.

[0231] Based on the analysis results, the server generates animations that represent the animals' reactions. Image generation technology is used to smooth the movements, providing a more realistic experience. The system also generates ambient sounds and sound effects of the animals' movements, creating a more immersive experience for the user.

[0232] The generated animation data is immediately transmitted to the terminal via the network. The terminal receives this data and displays the movements of the virtual animal in real time in response to the user's actions. As a result, the user can enjoy an interactive learning experience powered by the latest technology.

[0233] For example, when a user performs an action such as stroking the lion's head, the server generates an action where the lion narrows its eyes in apparent pleasure and lightly wags its tail. This result is immediately displayed on the device, allowing the user to enjoy interacting with the lion.

[0234] In this way, the system of the present invention can provide an interactive experience with virtual animals that has educational value and is highly safe.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] The user launches the application on their device and accesses the virtual safari park environment. A list of animals is displayed on the screen, and the user selects an animal they are interested in.

[0238] Step 2:

[0239] The device receives the user's touch input and displays an image of the selected animal (for example, a lion) on the screen. At this point, the user can begin interacting with the animal, such as touching or stroking it.

[0240] Step 3:

[0241] User actions are detected via the device's touchscreen and captured as operation data. This data includes the touch location, amount of movement, direction, and speed of movement.

[0242] Step 4:

[0243] The terminal transmits detected operation data to the server in real time. The data includes the user's ID, the animal's ID, and detailed information about the operation.

[0244] Step 5:

[0245] The server analyzes the received operation data. Using a generative AI model, it determines the animal's response (movement and facial expression) based on the acquired information. For example, if the user performs an action of stroking the lion's head, it will generate a response in which the lion narrows its eyes in apparent pleasure.

[0246] Step 6:

[0247] The server creates animations corresponding to the reactions of the generated animals. This uses image generation technology to produce smooth movements and realistic facial expressions.

[0248] Step 7:

[0249] The generated animation data is compressed and sent to the device in the optimal format. This enables real-time interaction with minimal latency.

[0250] Step 8:

[0251] The device unpacks the received animation data and displays it on the user's screen in real time. The user can then see the animal's reaction in response to their actions, allowing for a more interactive experience.

[0252] (Example 1)

[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0254] There is a growing need to provide more natural, safe, and educational interactions with virtual beings. However, conventional technologies suffer from a lack of real-time capabilities, resulting in delayed responses to user input or unnatural movements and expressions. It is necessary to overcome these challenges and provide users with a realistic experience.

[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0256] In this invention, the server includes means for acquiring operation information that senses user operations and generates information, means for analyzing operation information and generating responses of a virtual organism, and means for generating natural movements and facial expressions using a generative model. This makes it possible to generate immediate and natural responses to user operations and provide the user with a real-time, immersive interaction.

[0257] A "virtual organism" refers to a virtual animal or living being generated on a computer system, which a user can interact with through an interface.

[0258] "Display means" refers to a device or mechanism for outputting images or animations of virtual creatures so that users can visually recognize them.

[0259] "Operation information acquisition means" refers to a device or software that detects operations performed by a user and acquires that information as digital data.

[0260] "Analysis means" refers to a device or program that analyzes acquired operational information and performs processing to determine an appropriate response from a virtual organism based on that analysis.

[0261] "Motion generation means" refers to a mechanism or software that generates the movements and behaviors of a virtual organism based on the reactions determined by the analysis means.

[0262] "Display control means" refers to a device or software for playing back the movements and animations of a generated virtual creature and presenting them effectively to the user.

[0263] A "generative model" refers to a mathematical model or algorithm that uses machine learning or artificial intelligence techniques to generate natural movements and facial expressions for virtual beings.

[0264] "Information and communication network" refers to a network structure used to send and receive digital data, and includes the internet and dedicated lines.

[0265] This invention relates to a system for realizing interaction with virtual creatures. The user operates a terminal to generate real-time responses from the virtual animal. Specific embodiments of this system are described below.

[0266] The user first launches an application installed on their device. This displays a virtual safari park on the screen, allowing the user to select their favorite virtual animal. For example, if the user selects a lion, a virtual image of that lion will be displayed on the device's screen.

[0267] When a user touches or swipes an animal using the device's touchscreen, these actions are detected, and the device acquires action data such as the location, speed, and direction of the touch. This action data is transmitted to the server in real time.

[0268] The server is responsible for analyzing the received operation data. This analysis uses a generative AI model to generate natural movements of a virtual animal in response to user actions. Specifically, the generative AI model uses the prompt "Generate the lion's reaction when it is petted" to generate the lion's movements and facial expressions.

[0269] Based on the analysis results, the server animates the virtual animal's reactions and sends the generated data to the terminal. The terminal receives this data and displays the virtual lion's movements and expressions in real time in conjunction with the user's actions. This process allows the user to become familiar with the virtual lion and enjoy an interactive experience.

[0270] For example, if a user gently strokes the lion's ears, the server generates a motion where the lion happily wags its tail, and displays the result on the terminal. In this way, the virtual creature responds immediately to the user's actions, aiming to provide a more realistic and educational experience.

[0271] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0272] Step 1:

[0273] The user launches an application on their device. This displays a virtual safari park on the screen, and a screen for selecting an animal appears. When the user selects a lion, its virtual image is displayed on the device's screen. The input is the user's selection information, and the output is the virtual image of the selected animal.

[0274] Step 2:

[0275] The user touches or swipes on the selected animal on the device's touchscreen. The device detects these actions and generates action data, which includes information such as the position, speed, and direction of the action. The input is the user's touch actions, and the output is the action data.

[0276] Step 3:

[0277] The terminal sends the generated operation data to the server. The transmitted data includes details such as the location, speed, and direction of the touch. The input is the operation data, and the output is the transmission of data to the server.

[0278] Step 4:

[0279] The server analyzes the received operation data. This analysis process uses a generative AI model to determine the animal's response using prompt statements. For example, based on the prompt statement "Generate the lion's response when petted," it generates natural lion behavior. The input is operation data and prompt statements, and the output is the lion's response data.

[0280] Step 5:

[0281] The server generates animations of virtual animals based on the analysis results. The generated motion data is created by the server along with ambient sounds and sound effects. For example, it might add the actions of a lion squinting and wagging its tail, along with a purring sound. The input is reaction data, and the output is animation data and sound data.

[0282] Step 6:

[0283] The server transmits the generated animation data and voice data to the terminal. The terminal receives this in real time and displays the movement of the virtual lion according to the user's operation. The input is the animation data and voice data, and the output is the real-time display of the virtual animal.

[0284] (Application Example 1)

[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0286] In the conventional virtual experience system, the interaction with virtual organisms is limited, and it is difficult to provide sufficient educational value and effective guidance to users. In addition, due to insufficient real-time motion generation, there is a problem of lack of naturalness in interaction.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0288] In this invention, the server includes an operation data acquisition means for sensing user operations and generating data, an analysis means for analyzing the operation data and generating responses of virtual organisms, and a guidance generation means for providing guidance for virtual organisms based on information related to the user's environment. As a result, the user can receive real-time guidance from virtual organisms through an intuitive and flexible interaction with educational value.

[0289] A "virtual organism" is an electronic organism generated by a computer system and is displayed as an object for the user to interact with.

[0290] A "display means" is a function or device for displaying an image or movement of a virtual organism on a user terminal.

[0291] "User operation" refers to input operations and instructions performed by the user via a terminal, including those detected using touchscreens or sensors.

[0292] "Operation data acquisition means" refers to a function or device for sensing user operations and collecting and generating data based on those operations.

[0293] "Analysis means" refers to a data analysis function or device for generating and determining the responses and actions of a virtual organism in real time based on acquired operational data.

[0294] "Motion generation means" refers to a function or device that generates movements based on the analyzed results in order to give continuity and naturalness to the movements and facial expressions of a virtual organism.

[0295] "Display control means" refers to a control function or device that transmits generated actions or responses to a user terminal and displays them.

[0296] A "guide generation means" is a function or device that uses the user's environmental information to generate guide information for a virtual creature to provide appropriate guidance and explanations.

[0297] The program for the system to implement this invention aims to realize real-time interaction between a virtual organism and a user. This system mainly consists of three elements: a server, a terminal, and a user.

[0298] First, it is assumed that users will use smart glasses or head-mounted displays as smart devices. Users will wear these devices and interact with virtual creatures in a virtual space. The terminals have built-in display means and are equipped with means for acquiring operation data that receives user input. Through this operation, users can touch the virtual creatures and give them instructions.

[0299] The server analyzes user operation data received via operation data acquisition means using a generative AI model. This analysis process determines the virtual creature's response to specific user actions. Based on the analysis results, the server creates corresponding animations and motion data using motion generation means that generate the virtual creature's movements and expressions. This process utilizes a media generation engine to enable natural and smooth movements.

[0300] The generated motion data is sent to the terminal by the display control means and presented to the user in real time. This allows the user to experience intuitive and seamless interaction with the virtual creature. Furthermore, the server includes a guide generation means that analyzes the user's environment information and provides guidance to the user based on that information. This function allows the virtual creature to provide appropriate advice and explanations in real time according to the user's surroundings.

[0301] As a concrete example, if a user interacts with a virtual lion to learn about a specific product, the server will generate actions for the lion to provide a concise and easy-to-understand explanation of that product. The lion will guide the user in a friendly tone, saying, "This product is made from environmentally friendly materials and is a sustainable choice." An example of a prompt would be, "The user has shown interest in a new eco-bag. As the lion guide, what will you tell the user about the product?"

[0302] This allows users to enjoy a rich virtual experience that goes beyond mere visual enjoyment, enabling them to actually learn.

[0303] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0304] Step 1:

[0305] The user accesses the virtual space using a smart device. The user wears smart glasses or a head-mounted display and interacts with virtual creatures. The terminal receives the user's operation input through the operation data acquisition means and obtains the user's motion and position information as input. These data are transmitted to the server.

[0306] Step 2:

[0307] The server analyzes the received operation data of the user using the generated AI model. In this analysis process, data calculations are performed to determine the responses of virtual creatures to specific operations of the user. For example, the touch position, movement direction, etc. are input into the AI model as prompts, and the corresponding responses of virtual creatures are generated as outputs.

[0308] Step 3:

[0309] Based on the analysis results, the server uses the action generation means to generate the actions and expressions of virtual creatures. The process of generating smooth and realistic animations using the media generation engine is carried out here. The animation data generated is prepared as the output from the action generation means.

[0310] Step 4:

[0311] The generated action data and animations are transmitted to the terminal and presented to the user in real time by the display control means. The terminal enables the user to observe the actions of virtual creatures through the display based on the received action data.

[0312] Step 5:

[0313] The server further analyzes the user's environmental information and uses a guide generation mechanism to enable the virtual creature to provide appropriate guidance to the user. For example, if a user shows interest in a particular product, that information is input as environmental data, and a guide prompt is generated. This prompt is input into a generating AI model, and the virtual creature's guide information is generated as output.

[0314] Step 6:

[0315] Ultimately, the generated guide information and actions are displayed on the terminal, allowing the user to receive real-time guidance and explanations from the virtual creature. Through this, the user can choose to continue the operation or make a purchase decision.

[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0317] This invention relates to a system that, in addition to interaction with a virtual animal, recognizes the user's emotions and provides a personalized experience. This system can provide a more interactive and personalized learning experience by having the virtual animal respond in a way that takes the user's emotional state into account.

[0318] By launching an application on their device, users can access a virtual safari park and select one of the displayed animals. For example, if a lion is selected, a lion will appear on the screen. At this point, the user can begin interacting with the animal through touch controls.

[0319] The device uses cameras and microphones to recognize not only user actions but also the user's facial expressions and tone of voice in real time. This data is analyzed via an emotion engine to estimate the user's emotional state. This information is sent to the server along with the operation data.

[0320] The server generates virtual animal responses that correspond to the user's emotions, based on acquired operation data and estimated emotional information. For example, if the user is smiling, the lion can perform friendly actions, and if the user is feeling down, it can perform encouraging actions.

[0321] Next, the server generates an animation based on the appropriate animal response. Image generation technology is used for this generation, providing continuous and realistic motion. The animation data is compressed and sent to the terminal.

[0322] The device decompresses the transmitted animation data and displays it on the user's screen in real time. This allows users to interact with animals that react differently to their emotions, resulting in a more personalized experience.

[0323] For example, if the emotion engine determines that the user is happy as they touch the lion, the server will generate an animation in which the lion moves in a more friendly manner and behaves in a welcoming manner towards the user. Through this process, the user can enjoy a unique interaction experience.

[0324] This invention makes it possible to provide users with a more empathetic and intuitive learning environment by incorporating emotion recognition technology into interactions with virtual animals.

[0325] The following describes the processing flow.

[0326] Step 1:

[0327] The user launches an application on their device and accesses the virtual safari park environment. A list of animals is displayed on the device, and the user selects a virtual animal they are interested in.

[0328] Step 2:

[0329] The device displays the selected virtual animal on the screen. The user begins interacting with the virtual animal on the touchscreen and performs actions. At the same time, the device also acquires data on the user's facial expressions and voice.

[0330] Step 3:

[0331] The device collects user operation data and simultaneously incorporates facial expression and voice data acquired through the camera and microphone into the emotion engine. This allows the user's emotional state to be analyzed.

[0332] Step 4:

[0333] The terminal combines operational data and analyzed emotional data and sends it to the server. The server analyzes the received information and determines the virtual animal's response based on the user's emotions.

[0334] Step 5:

[0335] The server generates responses that adjust the actions and facial expressions of the virtual animal based on the user's emotional state. For example, if the user seems happy, the lion will perform playful actions, while if the user seems sad, it will generate encouraging actions.

[0336] Step 6:

[0337] Based on the generated animal reactions, the server creates a continuous animation. High-quality, realistic animations are produced using image generation technology.

[0338] Step 7:

[0339] The server sends the generated animation data to the terminal. Since the data is transmitted in real time, the user can receive immediate feedback.

[0340] Step 8:

[0341] The device decompresses the received animation data and displays it immediately on the screen. Users can directly see the virtual animal's reactions in response to their own actions and emotions, resulting in a more empathetic experience.

[0342] (Example 2)

[0343] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0344] Current virtual animal interaction systems have limitations in their response to user input and are unable to provide personalized experiences that incorporate emotional recognition. Therefore, there is a need to provide more personalized interactions that take user emotions into consideration.

[0345] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0346] In this invention, the server includes means for displaying a virtual animal, means for acquiring data by sensing user actions and emotions and generating data, and means for analyzing the acquired action data and emotion data and generating responses for the virtual animal. This allows the virtual animal to show individual responses in accordance with the user's emotions, enabling an interactive and personalized experience.

[0347] A "virtual animal" is a computer-simulated animal created by a computer program that can interact with the user.

[0348] "Display means" refers to technologies and devices for visually displaying information on a user's terminal, including displays and screens.

[0349] "Data acquisition methods" refer to technologies and techniques for detecting user actions and emotional states and collecting them as data.

[0350] "Analysis means" refers to a method or apparatus for determining the response of a virtual animal using a computer algorithm based on acquired data.

[0351] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generatively creates new content and responses.

[0352] "Animation generation means" refers to technologies and methods for creating images or videos to continuously represent the movements of a virtual animal.

[0353] "Display control means" refers to technologies and devices that transmit generated animations to the user's terminal and control them to execute them appropriately.

[0354] This invention is a system that provides a personalized experience that takes into account the user's emotions through interaction with virtual animals. Users can launch a specific application on their device, access a virtual safari, select an animal, and begin interacting with it.

[0355] Users can directly control virtual animals using a touch panel. The device uses its built-in camera and microphone to sense the user's facial expressions and voice tone, collecting emotional data in real time. This data is temporarily stored on the device and sent to the emotion engine. The emotion engine analyzes this data using machine learning models to estimate the user's emotional state. The machine learning models used include deep learning techniques.

[0356] The estimated emotion information is sent to the server along with the user's interaction data. The server uses a generative AI model to generate appropriate virtual animal responses based on this information. For example, if the user smiles, the lion is programmed to perform friendly actions. Generative AI models may utilize technologies such as natural language processing or image generation.

[0357] The server generates an animation based on the determined animal's response. The generated animation is then sent from the server to the terminal after its data size is reduced using compression technology. The terminal decompresses the received animation and displays it on the user's screen in real time.

[0358] For example, when a user touches a lion, if the emotion engine detects the user's smile, the server generates an animation of the lion welcoming the user. Through such interactions, users can enjoy a personalized experience that responds to their emotions in real time.

[0359] An example of a prompt might be, "Create an animation showing how a lion behaves in a friendly manner when a user touches it with a smile." Through this prompt, the generative AI model can generate natural reactions that meet the user's expectations.

[0360] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0361] Step 1:

[0362] The user launches the application on their device and accesses a virtual safari. From the animals displayed, they select, for example, a lion. The selection is made by tapping on the touch panel. The input here is the user's selection, and the selected animal is displayed on the screen as output.

[0363] Step 2:

[0364] The device activates its built-in camera and microphone to sense the user's facial expressions and voice tone in real time. This allows for the collection of emotional data. The user's facial expressions and voice data are used as input, and emotional data is generated as output for transmission to the emotion engine.

[0365] Step 3:

[0366] The device sends the collected emotional data to the emotion engine, which analyzes the data using a machine learning model. The analysis then estimates the user's emotional state. The input is emotional data, and the output is estimated emotional information. This estimated information, along with the user's interaction data, is used in the next step.

[0367] Step 4:

[0368] The server receives operation data and emotion information sent from the terminal. Using a generative AI model, it generates a virtual animal response that corresponds to the user's emotions. Specifically, if the user is smiling, it generates an action where the lion approaches in a friendly manner. The input is operation data and emotion information, and the output is virtual animal response data.

[0369] Step 5:

[0370] The server creates animations based on the generated reaction data. Image generation technology is used to achieve realistic movement. The input is the reaction data of a virtual animal, and the output is compressed animation data.

[0371] Step 6:

[0372] The device receives and decompresses compressed animation data sent from the server. It then displays the animation on the user's screen in real time. The user can enjoy interacting with an animal that responds to emotions. The input is compressed animation data, and the output is the animal's actions displayed on the user's screen.

[0373] (Application Example 2)

[0374] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0375] In modern production environments, there is a need for efficient work support that takes into account the emotions and physical condition of workers. However, conventional systems have difficulty responding appropriately to changes in workers' emotions, limiting improvements in work efficiency and safety. Therefore, there is a need for a system that can recognize workers' emotional states in real time and provide individualized and appropriate support accordingly.

[0376] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0377] In this invention, the server includes means for displaying a virtual organism, means for acquiring operation information that senses operations and generates information, and means for recognizing the user's emotional state. This makes it possible to provide appropriate work support immediately according to the user's emotional state.

[0378] A "virtual organism" refers to a living being simulated on a computer, which exhibits various reactions and actions through interaction with the user.

[0379] "Display means" refers to functions that visualize virtual creatures on a computer screen or device, and includes technologies for presenting information to the user in real time.

[0380] "Operation information acquisition means" refers to a function that detects the user's physical operations and inputs and collects and processes information generated based on them.

[0381] "Analysis means" refers to a function that includes a process for analyzing acquired operational information and emotional state information to generate appropriate responses and actions for the virtual organism.

[0382] "Visual representation generation means" refers to technology for visualizing the actions and reactions of virtual organisms and presenting them to the user as seamless and realistic images.

[0383] "Display control means" refers to a technology that has the function of appropriately transmitting the generated visual representation to the user terminal and displaying it on the screen.

[0384] "Emotion recognition means" refers to a function that detects and analyzes the user's emotional state from their facial expressions, tone of voice, etc.

[0385] "Work support means" refers to a process that includes functions to provide specific work support and interaction tailored to the user's emotional state.

[0386] This invention is a system that recognizes the emotional state of a user and provides individualized and appropriate support tailored to that user. This system utilizes virtual organisms and is designed to support users in working efficiently and safely in their work environment.

[0387] The server uses emotion recognition means to analyze data collected from the camera and microphone of the user's smart device. This analysis uses voice analysis software and facial recognition technology as examples. It detects changes in the user's facial expressions and tone of voice in real time and determines their emotional state. This information is processed via operation information acquisition means and treated as basic data for the virtual organism to show appropriate responses as needed.

[0388] The terminal receives visual representation data transmitted from the server and displays it on the user's screen. The visual representation generation means provides seamless and realistic movement of the virtual creature, which the user can visually confirm. For example, if the user shows signs of fatigue, the virtual creature displays encouraging messages or suggestions for improving work methods. This interaction allows the user to relax and increase efficiency while working.

[0389] As a concrete example, if the emotion recognition system detects signs of stress while the user is working, the server generates a message such as "We recommend taking a break to relax" and presents it to the user as a visual representation. An example of a prompt message to the generating AI model would be: "We have emotion analysis data for the user. He appears to be tired at the moment. Please generate support or guidance to prompt the system."

[0390] This makes it possible to provide users with a better work environment and personalized support.

[0391] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0392] Step 1:

[0393] The user puts on a smart device and begins working. The device activates its camera and microphone to capture the user's facial expressions and voice data. The input consists of video and audio data, and the output is prepared for analysis of this data. The device collects this data from sensors in real time.

[0394] Step 2:

[0395] The terminal sends the collected data to the server. The server processes the received video and audio data using analysis software. The input is the raw video and audio data sent from the terminal, and the output is extracted feature data. Specifically, the server uses facial recognition technology to extract the user's facial features and voice analysis technology to analyze their tone.

[0396] Step 3:

[0397] Based on the analysis results, the server uses a generative AI model to estimate the user's emotional state. This model combines facial and vocal features to generate emotion labels. The input is feature data, and the output is an emotional state (e.g., stress, exhilaration, fatigue). The server accurately estimates the emotional state and stores it in a database for the next step.

[0398] Step 4:

[0399] The server generates interactions with the virtual creature based on its estimated emotional state. Specific actions and messages are determined by a generating AI model. Inputs are emotional states and past interaction data, while output is the virtual creature's response (e.g., soothing actions or relaxing messages). The server analyzes the user's past emotions and selects the optimal response.

[0400] Step 5:

[0401] The server visualizes the reactions of the generated virtual organism through a visual representation generation system and transmits that data to the terminal. The input is the reactions of the virtual organism, and the output is effective visual representation data. The server performs the necessary processing for visualization, compresses the data, and transmits it.

[0402] Step 6:

[0403] The terminal decompresses the received visual representation data and displays it on the user's device. The input is compressed visual representation data, and the output is the animation and interaction that the user actually sees. The terminal displays the virtual creature's actions on the home screen and prepares to take further action information to check the user's response.

[0404] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0405] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0406] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0407] [Third Embodiment]

[0408] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0409] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0410] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0411] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0412] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0413] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0414] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0415] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0416] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0417] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0418] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0419] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0420] This invention relates to a system for safely and educationally facilitating interaction with virtual animals. This system allows users to interact with virtual animals in real time by operating a terminal. Specific embodiments of the invention are described below.

[0421] The user first launches the application on their device and accesses the virtual safari park screen. Here, a list of animals the user can choose from is displayed, and if they select, for example, a lion, a virtual image of that lion will be displayed on their device.

[0422] Users interact with the lion by touching or swiping on the device's touchscreen. These user actions are detected by the device and collected as operation data. This operation data includes information such as the location of the touch, the speed of the movement, and the direction.

[0423] The terminal sends the acquired operation data to the server. The server analyzes this data and generates the appropriate responses for the animal in real time. This process utilizes a generative AI model, which assigns natural and intuitive movements and expressions to the virtual animal in response to user input.

[0424] Based on the analysis results, the server generates animations that represent the animals' reactions. Image generation technology is used to smooth the movements, providing a more realistic experience. The system also generates ambient sounds and sound effects of the animals' movements, creating a more immersive experience for the user.

[0425] The generated animation data is immediately transmitted to the terminal via the network. The terminal receives this data and displays the movements of the virtual animal in real time in response to the user's actions. As a result, the user can enjoy an interactive learning experience powered by the latest technology.

[0426] For example, when a user performs an action such as stroking the lion's head, the server generates an action where the lion narrows its eyes in apparent pleasure and lightly wags its tail. This result is immediately displayed on the device, allowing the user to enjoy interacting with the lion.

[0427] In this way, the system of the present invention can provide an interactive experience with virtual animals that has educational value and is highly safe.

[0428] The following describes the processing flow.

[0429] Step 1:

[0430] The user launches the application on their device and accesses the virtual safari park environment. A list of animals is displayed on the screen, and the user selects an animal they are interested in.

[0431] Step 2:

[0432] The device receives the user's touch input and displays an image of the selected animal (for example, a lion) on the screen. At this point, the user can begin interacting with the animal, such as touching or stroking it.

[0433] Step 3:

[0434] User actions are detected via the device's touchscreen and captured as operation data. This data includes the touch location, amount of movement, direction, and speed of movement.

[0435] Step 4:

[0436] The terminal transmits detected operation data to the server in real time. The data includes the user's ID, the animal's ID, and detailed information about the operation.

[0437] Step 5:

[0438] The server analyzes the received operation data. Using a generative AI model, it determines the animal's response (movement and facial expression) based on the acquired information. For example, if the user performs an action of stroking the lion's head, it will generate a response in which the lion narrows its eyes in apparent pleasure.

[0439] Step 6:

[0440] The server creates animations corresponding to the reactions of the generated animals. This uses image generation technology to produce smooth movements and realistic facial expressions.

[0441] Step 7:

[0442] The generated animation data is compressed and sent to the device in the optimal format. This enables real-time interaction with minimal latency.

[0443] Step 8:

[0444] The device unpacks the received animation data and displays it on the user's screen in real time. The user can then see the animal's reaction in response to their actions, allowing for a more interactive experience.

[0445] (Example 1)

[0446] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0447] There is a growing need to provide more natural, safe, and educational interactions with virtual beings. However, conventional technologies suffer from a lack of real-time capabilities, resulting in delayed responses to user input or unnatural movements and expressions. It is necessary to overcome these challenges and provide users with a realistic experience.

[0448] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0449] In this invention, the server includes means for acquiring operation information that senses user operations and generates information, means for analyzing operation information and generating responses of a virtual organism, and means for generating natural movements and facial expressions using a generative model. This makes it possible to generate immediate and natural responses to user operations and provide the user with a real-time, immersive interaction.

[0450] A "virtual organism" refers to a virtual animal or living being generated on a computer system, which a user can interact with through an interface.

[0451] "Display means" refers to a device or mechanism for outputting images or animations of virtual creatures so that users can visually recognize them.

[0452] "Operation information acquisition means" refers to a device or software that detects operations performed by a user and acquires that information as digital data.

[0453] "Analysis means" refers to a device or program that analyzes acquired operational information and performs processing to determine an appropriate response from a virtual organism based on that analysis.

[0454] "Motion generation means" refers to a mechanism or software that generates the movements and behaviors of a virtual organism based on the reactions determined by the analysis means.

[0455] "Display control means" refers to a device or software for playing back the movements and animations of a generated virtual creature and presenting them effectively to the user.

[0456] A "generative model" refers to a mathematical model or algorithm that uses machine learning or artificial intelligence techniques to generate natural movements and facial expressions for virtual beings.

[0457] "Information and communication network" refers to a network structure used to send and receive digital data, and includes the internet and dedicated lines.

[0458] This invention relates to a system for realizing interaction with virtual creatures. The user operates a terminal to generate real-time responses from the virtual animal. Specific embodiments of this system are described below.

[0459] The user first launches an application installed on their device. This displays a virtual safari park on the screen, allowing the user to select their favorite virtual animal. For example, if the user selects a lion, a virtual image of that lion will be displayed on the device's screen.

[0460] When a user touches or swipes an animal using the device's touchscreen, these actions are detected, and the device acquires action data such as the location, speed, and direction of the touch. This action data is transmitted to the server in real time.

[0461] The server is responsible for analyzing the received operation data. This analysis uses a generative AI model to generate natural movements of a virtual animal in response to user actions. Specifically, the generative AI model uses the prompt "Generate the lion's reaction when it is petted" to generate the lion's movements and facial expressions.

[0462] Based on the analysis results, the server animates the virtual animal's reactions and sends the generated data to the terminal. The terminal receives this data and displays the virtual lion's movements and expressions in real time in conjunction with the user's actions. This process allows the user to become familiar with the virtual lion and enjoy an interactive experience.

[0463] For example, if a user gently strokes the lion's ears, the server generates a motion where the lion happily wags its tail, and displays the result on the terminal. In this way, the virtual creature responds immediately to the user's actions, aiming to provide a more realistic and educational experience.

[0464] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0465] Step 1:

[0466] The user launches an application on their device. This displays a virtual safari park on the screen, and a screen for selecting an animal appears. When the user selects a lion, its virtual image is displayed on the device's screen. The input is the user's selection information, and the output is the virtual image of the selected animal.

[0467] Step 2:

[0468] The user touches or swipes on the selected animal on the device's touchscreen. The device detects these actions and generates action data, which includes information such as the position, speed, and direction of the action. The input is the user's touch actions, and the output is the action data.

[0469] Step 3:

[0470] The terminal sends the generated operation data to the server. The transmitted data includes details such as the location, speed, and direction of the touch. The input is the operation data, and the output is the transmission of data to the server.

[0471] Step 4:

[0472] The server analyzes the received operation data. This analysis process uses a generative AI model to determine the animal's response using prompt statements. For example, based on the prompt statement "Generate the lion's response when petted," it generates natural lion behavior. The input is operation data and prompt statements, and the output is the lion's response data.

[0473] Step 5:

[0474] The server generates animations of virtual animals based on the analysis results. The generated motion data is created by the server along with ambient sounds and sound effects. For example, it might add the actions of a lion squinting and wagging its tail, along with a purring sound. The input is reaction data, and the output is animation data and sound data.

[0475] Step 6:

[0476] The server sends the generated animation and audio data to the terminal. The terminal receives this in real time and displays the movements of the virtual lion in response to user input. The input is animation and audio data, and the output is a real-time display of the virtual animal.

[0477] (Application Example 1)

[0478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0479] Conventional virtual experience systems have limited interaction with virtual creatures, making it difficult to provide users with sufficient educational value or effective guidance. Furthermore, the insufficient real-time motion generation resulted in unnatural interactions.

[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0481] In this invention, the server includes operation data acquisition means that senses user operations and generates data, analysis means that analyzes the operation data and generates responses from the virtual creature, and guide generation means that provides guidance from the virtual creature based on information related to the user's environment. This enables the user to receive real-time guidance from the virtual creature through an intuitive and flexible interaction with educational value.

[0482] A "virtual organism" is an electronic living entity generated by a computer system, which is displayed as an object that the user interacts with.

[0483] "Display means" refers to a function or device for displaying images and actions of virtual creatures on a user's terminal.

[0484] "User operation" refers to input operations and instructions performed by the user via a terminal, including those detected using touchscreens or sensors.

[0485] "Operation data acquisition means" refers to a function or device for sensing user operations and collecting and generating data based on those operations.

[0486] "Analysis means" refers to a data analysis function or device for generating and determining the responses and actions of a virtual organism in real time based on acquired operational data.

[0487] "Motion generation means" refers to a function or device that generates movements based on the analyzed results in order to give continuity and naturalness to the movements and facial expressions of a virtual organism.

[0488] "Display control means" refers to a control function or device that transmits generated actions or responses to a user terminal and displays them.

[0489] A "guide generation means" is a function or device that uses the user's environmental information to generate guide information for a virtual creature to provide appropriate guidance and explanations.

[0490] The program for the system to implement this invention aims to realize real-time interaction between a virtual organism and a user. This system mainly consists of three elements: a server, a terminal, and a user.

[0491] First, it is assumed that users will use smart glasses or head-mounted displays as smart devices. Users will wear these devices and interact with virtual creatures in a virtual space. The terminals have built-in display means and are equipped with means for acquiring operation data that receives user input. Through this operation, users can touch the virtual creatures and give them instructions.

[0492] The server analyzes user operation data received via operation data acquisition means using a generative AI model. This analysis process determines the virtual creature's response to specific user actions. Based on the analysis results, the server creates corresponding animations and motion data using motion generation means that generate the virtual creature's movements and expressions. This process utilizes a media generation engine to enable natural and smooth movements.

[0493] The generated motion data is sent to the terminal by the display control means and presented to the user in real time. This allows the user to experience intuitive and seamless interaction with the virtual creature. Furthermore, the server includes a guide generation means that analyzes the user's environment information and provides guidance to the user based on that information. This function allows the virtual creature to provide appropriate advice and explanations in real time according to the user's surroundings.

[0494] As a concrete example, if a user interacts with a virtual lion to learn about a specific product, the server will generate actions for the lion to provide a concise and easy-to-understand explanation of that product. The lion will guide the user in a friendly tone, saying, "This product is made from environmentally friendly materials and is a sustainable choice." An example of a prompt would be, "The user has shown interest in a new eco-bag. As the lion guide, what will you tell the user about the product?"

[0495] This allows users to enjoy a rich virtual experience that goes beyond mere visual enjoyment, enabling them to actually learn.

[0496] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0497] Step 1:

[0498] Users access a virtual space using smart devices. Users wear smart glasses or head-mounted displays and interact with virtual beings. The terminal receives user input through an operation data acquisition mechanism, obtaining user actions and location information as input. This data is transmitted to the server.

[0499] Step 2:

[0500] The server analyzes the user's input data using a generated AI model. This analysis process involves data calculations to determine the virtual creature's response to specific user actions. For example, the touch location and direction of movement are input to the AI ​​model as prompts, and the corresponding response from the virtual creature is generated as output.

[0501] Step 3:

[0502] Based on the analysis results, the server uses motion generation tools to generate the movements and facial expressions of the virtual creature. A media generation engine is then used to create smooth and realistic animations. The generated animation data is provided as output from the motion generation tools.

[0503] Step 4:

[0504] The generated motion data and animations are sent to the terminal and presented to the user in real time by a display control system. Based on the received motion data, the terminal allows the user to observe the movements of the virtual creature through the display.

[0505] Step 5:

[0506] The server further analyzes the user's environmental information and uses a guide generation mechanism to enable the virtual creature to provide appropriate guidance to the user. For example, if a user shows interest in a particular product, that information is input as environmental data, and a guide prompt is generated. This prompt is input into a generating AI model, and the virtual creature's guide information is generated as output.

[0507] Step 6:

[0508] Ultimately, the generated guide information and actions are displayed on the terminal, allowing the user to receive real-time guidance and explanations from the virtual creature. Through this, the user can choose to continue the operation or make a purchase decision.

[0509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0510] This invention relates to a system that, in addition to interaction with a virtual animal, recognizes the user's emotions and provides a personalized experience. This system can provide a more interactive and personalized learning experience by having the virtual animal respond in a way that takes the user's emotional state into account.

[0511] By launching an application on their device, users can access a virtual safari park and select one of the displayed animals. For example, if a lion is selected, a lion will appear on the screen. At this point, the user can begin interacting with the animal through touch controls.

[0512] The device uses cameras and microphones to recognize not only user actions but also the user's facial expressions and tone of voice in real time. This data is analyzed via an emotion engine to estimate the user's emotional state. This information is sent to the server along with the operation data.

[0513] The server generates virtual animal responses that correspond to the user's emotions, based on acquired operation data and estimated emotional information. For example, if the user is smiling, the lion can perform friendly actions, and if the user is feeling down, it can perform encouraging actions.

[0514] Next, the server generates an animation based on the appropriate animal response. Image generation technology is used for this generation, providing continuous and realistic motion. The animation data is compressed and sent to the terminal.

[0515] The device decompresses the transmitted animation data and displays it on the user's screen in real time. This allows users to interact with animals that react differently to their emotions, resulting in a more personalized experience.

[0516] For example, if the emotion engine determines that the user is happy as they touch the lion, the server will generate an animation in which the lion moves in a more friendly manner and behaves in a welcoming manner towards the user. Through this process, the user can enjoy a unique interaction experience.

[0517] This invention makes it possible to provide users with a more empathetic and intuitive learning environment by incorporating emotion recognition technology into interactions with virtual animals.

[0518] The following describes the processing flow.

[0519] Step 1:

[0520] The user launches an application on their device and accesses the virtual safari park environment. A list of animals is displayed on the device, and the user selects a virtual animal they are interested in.

[0521] Step 2:

[0522] The device displays the selected virtual animal on the screen. The user begins interacting with the virtual animal on the touchscreen and performs actions. At the same time, the device also acquires data on the user's facial expressions and voice.

[0523] Step 3:

[0524] The device collects user operation data and simultaneously incorporates facial expression and voice data acquired through the camera and microphone into the emotion engine. This allows the user's emotional state to be analyzed.

[0525] Step 4:

[0526] The terminal combines operational data and analyzed emotional data and sends it to the server. The server analyzes the received information and determines the virtual animal's response based on the user's emotions.

[0527] Step 5:

[0528] The server generates responses that adjust the actions and facial expressions of the virtual animal based on the user's emotional state. For example, if the user seems happy, the lion will perform playful actions, while if the user seems sad, it will generate encouraging actions.

[0529] Step 6:

[0530] Based on the generated animal reactions, the server creates a continuous animation. High-quality, realistic animations are produced using image generation technology.

[0531] Step 7:

[0532] The server sends the generated animation data to the terminal. Since the data is transmitted in real time, the user can receive immediate feedback.

[0533] Step 8:

[0534] The device decompresses the received animation data and displays it immediately on the screen. Users can directly see the virtual animal's reactions in response to their own actions and emotions, resulting in a more empathetic experience.

[0535] (Example 2)

[0536] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0537] Current virtual animal interaction systems have limitations in their response to user input and are unable to provide personalized experiences that incorporate emotional recognition. Therefore, there is a need to provide more personalized interactions that take user emotions into consideration.

[0538] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0539] In this invention, the server includes means for displaying a virtual animal, means for acquiring data by sensing user actions and emotions and generating data, and means for analyzing the acquired action data and emotion data and generating responses for the virtual animal. This allows the virtual animal to show individual responses in accordance with the user's emotions, enabling an interactive and personalized experience.

[0540] A "virtual animal" is a computer-simulated animal created by a computer program that can interact with the user.

[0541] "Display means" refers to technologies and devices for visually displaying information on a user's terminal, including displays and screens.

[0542] "Data acquisition methods" refer to technologies and techniques for detecting user actions and emotional states and collecting them as data.

[0543] "Analysis means" refers to a method or apparatus for determining the response of a virtual animal using a computer algorithm based on acquired data.

[0544] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generatively creates new content and responses.

[0545] "Animation generation means" refers to technologies and methods for creating images or videos to continuously represent the movements of a virtual animal.

[0546] "Display control means" refers to technologies and devices that transmit generated animations to the user's terminal and control them to execute them appropriately.

[0547] This invention is a system that provides a personalized experience that takes into account the user's emotions through interaction with virtual animals. Users can launch a specific application on their device, access a virtual safari, select an animal, and begin interacting with it.

[0548] Users can directly control virtual animals using a touch panel. The device uses its built-in camera and microphone to sense the user's facial expressions and voice tone, collecting emotional data in real time. This data is temporarily stored on the device and sent to the emotion engine. The emotion engine analyzes this data using machine learning models to estimate the user's emotional state. The machine learning models used include deep learning techniques.

[0549] The estimated emotion information is sent to the server along with the user's interaction data. The server uses a generative AI model to generate appropriate virtual animal responses based on this information. For example, if the user smiles, the lion is programmed to perform friendly actions. Generative AI models may utilize technologies such as natural language processing or image generation.

[0550] The server generates an animation based on the determined animal's response. The generated animation is then sent from the server to the terminal after its data size is reduced using compression technology. The terminal decompresses the received animation and displays it on the user's screen in real time.

[0551] For example, when a user touches a lion, if the emotion engine detects the user's smile, the server generates an animation of the lion welcoming the user. Through such interactions, users can enjoy a personalized experience that responds to their emotions in real time.

[0552] An example of a prompt might be, "Create an animation showing how a lion behaves in a friendly manner when a user touches it with a smile." Through this prompt, the generative AI model can generate natural reactions that meet the user's expectations.

[0553] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0554] Step 1:

[0555] The user launches the application on their device and accesses a virtual safari. From the animals displayed, they select, for example, a lion. The selection is made by tapping on the touch panel. The input here is the user's selection, and the selected animal is displayed on the screen as output.

[0556] Step 2:

[0557] The device activates its built-in camera and microphone to sense the user's facial expressions and voice tone in real time. This allows for the collection of emotional data. The user's facial expressions and voice data are used as input, and emotional data is generated as output for transmission to the emotion engine.

[0558] Step 3:

[0559] The device sends the collected emotional data to the emotion engine, which analyzes the data using a machine learning model. The analysis then estimates the user's emotional state. The input is emotional data, and the output is estimated emotional information. This estimated information, along with the user's interaction data, is used in the next step.

[0560] Step 4:

[0561] The server receives operation data and emotion information sent from the terminal. Using a generative AI model, it generates a virtual animal response that corresponds to the user's emotions. Specifically, if the user is smiling, it generates an action where the lion approaches in a friendly manner. The input is operation data and emotion information, and the output is virtual animal response data.

[0562] Step 5:

[0563] The server creates animations based on the generated reaction data. Image generation technology is used to achieve realistic movement. The input is the reaction data of a virtual animal, and the output is compressed animation data.

[0564] Step 6:

[0565] The device receives and decompresses compressed animation data sent from the server. It then displays the animation on the user's screen in real time. The user can enjoy interacting with an animal that responds to emotions. The input is compressed animation data, and the output is the animal's actions displayed on the user's screen.

[0566] (Application Example 2)

[0567] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0568] In modern production environments, there is a need for efficient work support that takes into account the emotions and physical condition of workers. However, conventional systems have difficulty responding appropriately to changes in workers' emotions, limiting improvements in work efficiency and safety. Therefore, there is a need for a system that can recognize workers' emotional states in real time and provide individualized and appropriate support accordingly.

[0569] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0570] In this invention, the server includes means for displaying a virtual organism, means for acquiring operation information that senses operations and generates information, and means for recognizing the user's emotional state. This makes it possible to provide appropriate work support immediately according to the user's emotional state.

[0571] A "virtual organism" refers to a living being simulated on a computer, which exhibits various reactions and actions through interaction with the user.

[0572] "Display means" refers to functions that visualize virtual creatures on a computer screen or device, and includes technologies for presenting information to the user in real time.

[0573] "Operation information acquisition means" refers to a function that detects the user's physical operations and inputs and collects and processes information generated based on them.

[0574] "Analysis means" refers to a function that includes a process for analyzing acquired operational information and emotional state information to generate appropriate responses and actions for the virtual organism.

[0575] "Visual representation generation means" refers to technology for visualizing the actions and reactions of virtual organisms and presenting them to the user as seamless and realistic images.

[0576] "Display control means" refers to a technology that has the function of appropriately transmitting the generated visual representation to the user terminal and displaying it on the screen.

[0577] "Emotion recognition means" refers to a function that detects and analyzes the user's emotional state from their facial expressions, tone of voice, etc.

[0578] "Work support means" refers to a process that includes functions to provide specific work support and interaction tailored to the user's emotional state.

[0579] This invention is a system that recognizes the emotional state of a user and provides individualized and appropriate support tailored to that user. This system utilizes virtual organisms and is designed to support users in working efficiently and safely in their work environment.

[0580] The server uses emotion recognition means to analyze data collected from the camera and microphone of the user's smart device. This analysis uses voice analysis software and facial recognition technology as examples. It detects changes in the user's facial expressions and tone of voice in real time and determines their emotional state. This information is processed via operation information acquisition means and treated as basic data for the virtual organism to show appropriate responses as needed.

[0581] The terminal receives visual representation data transmitted from the server and displays it on the user's screen. The visual representation generation means provides seamless and realistic movement of the virtual creature, which the user can visually confirm. For example, if the user shows signs of fatigue, the virtual creature displays encouraging messages or suggestions for improving work methods. This interaction allows the user to relax and increase efficiency while working.

[0582] As a concrete example, if the emotion recognition system detects signs of stress while the user is working, the server generates a message such as "We recommend taking a break to relax" and presents it to the user as a visual representation. An example of a prompt message to the generating AI model would be: "We have emotion analysis data for the user. He appears to be tired at the moment. Please generate support or guidance to prompt the system."

[0583] This makes it possible to provide users with a better work environment and personalized support.

[0584] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0585] Step 1:

[0586] The user puts on a smart device and begins working. The device activates its camera and microphone to capture the user's facial expressions and voice data. The input consists of video and audio data, and the output is prepared for analysis of this data. The device collects this data from sensors in real time.

[0587] Step 2:

[0588] The terminal sends the collected data to the server. The server processes the received video and audio data using analysis software. The input is the raw video and audio data sent from the terminal, and the output is extracted feature data. Specifically, the server uses facial recognition technology to extract the user's facial features and voice analysis technology to analyze their tone.

[0589] Step 3:

[0590] Based on the analysis results, the server uses a generative AI model to estimate the user's emotional state. This model combines facial and vocal features to generate emotion labels. The input is feature data, and the output is an emotional state (e.g., stress, exhilaration, fatigue). The server accurately estimates the emotional state and stores it in a database for the next step.

[0591] Step 4:

[0592] The server generates interactions with the virtual creature based on its estimated emotional state. Specific actions and messages are determined by a generating AI model. Inputs are emotional states and past interaction data, while output is the virtual creature's response (e.g., soothing actions or relaxing messages). The server analyzes the user's past emotions and selects the optimal response.

[0593] Step 5:

[0594] The server visualizes the reactions of the generated virtual organism through a visual representation generation system and transmits that data to the terminal. The input is the reactions of the virtual organism, and the output is effective visual representation data. The server performs the necessary processing for visualization, compresses the data, and transmits it.

[0595] Step 6:

[0596] The terminal decompresses the received visual representation data and displays it on the user's device. The input is compressed visual representation data, and the output is the animation and interaction that the user actually sees. The terminal displays the virtual creature's actions on the home screen and prepares to take further action information to check the user's response.

[0597] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0598] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0599] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0600] [Fourth Embodiment]

[0601] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0602] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0603] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0604] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0605] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0606] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0607] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0608] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0609] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0610] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0611] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0612] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0613] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0614] This invention relates to a system for safely and educationally facilitating interaction with virtual animals. This system allows users to interact with virtual animals in real time by operating a terminal. Specific embodiments of the invention are described below.

[0615] The user first launches the application on their device and accesses the virtual safari park screen. Here, a list of animals the user can choose from is displayed, and if they select, for example, a lion, a virtual image of that lion will be displayed on their device.

[0616] Users interact with the lion by touching or swiping on the device's touchscreen. These user actions are detected by the device and collected as operation data. This operation data includes information such as the location of the touch, the speed of the movement, and the direction.

[0617] The terminal sends the acquired operation data to the server. The server analyzes this data and generates the appropriate responses for the animal in real time. This process utilizes a generative AI model, which assigns natural and intuitive movements and expressions to the virtual animal in response to user input.

[0618] Based on the analysis results, the server generates animations that represent the animals' reactions. Image generation technology is used to smooth the movements, providing a more realistic experience. The system also generates ambient sounds and sound effects of the animals' movements, creating a more immersive experience for the user.

[0619] The generated animation data is immediately transmitted to the terminal via the network. The terminal receives this data and displays the movements of the virtual animal in real time in response to the user's actions. As a result, the user can enjoy an interactive learning experience powered by the latest technology.

[0620] For example, when a user performs an action such as stroking the lion's head, the server generates an action where the lion narrows its eyes in apparent pleasure and lightly wags its tail. This result is immediately displayed on the device, allowing the user to enjoy interacting with the lion.

[0621] In this way, the system of the present invention can provide an interactive experience with virtual animals that has educational value and is highly safe.

[0622] The following describes the processing flow.

[0623] Step 1:

[0624] The user launches the application on their device and accesses the virtual safari park environment. A list of animals is displayed on the screen, and the user selects an animal they are interested in.

[0625] Step 2:

[0626] The device receives the user's touch input and displays an image of the selected animal (for example, a lion) on the screen. At this point, the user can begin interacting with the animal, such as touching or stroking it.

[0627] Step 3:

[0628] User actions are detected via the device's touchscreen and captured as operation data. This data includes the touch location, amount of movement, direction, and speed of movement.

[0629] Step 4:

[0630] The terminal transmits detected operation data to the server in real time. The data includes the user's ID, the animal's ID, and detailed information about the operation.

[0631] Step 5:

[0632] The server analyzes the received operation data. Using a generative AI model, it determines the animal's response (movement and facial expression) based on the acquired information. For example, if the user performs an action of stroking the lion's head, it will generate a response in which the lion narrows its eyes in apparent pleasure.

[0633] Step 6:

[0634] The server creates animations corresponding to the reactions of the generated animals. This uses image generation technology to produce smooth movements and realistic facial expressions.

[0635] Step 7:

[0636] The generated animation data is compressed and sent to the device in the optimal format. This enables real-time interaction with minimal latency.

[0637] Step 8:

[0638] The device unpacks the received animation data and displays it on the user's screen in real time. The user can then see the animal's reaction in response to their actions, allowing for a more interactive experience.

[0639] (Example 1)

[0640] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0641] There is a growing need to provide more natural, safe, and educational interactions with virtual beings. However, conventional technologies suffer from a lack of real-time capabilities, resulting in delayed responses to user input or unnatural movements and expressions. It is necessary to overcome these challenges and provide users with a realistic experience.

[0642] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0643] In this invention, the server includes means for acquiring operation information that senses user operations and generates information, means for analyzing operation information and generating responses of a virtual organism, and means for generating natural movements and facial expressions using a generative model. This makes it possible to generate immediate and natural responses to user operations and provide the user with a real-time, immersive interaction.

[0644] A "virtual organism" refers to a virtual animal or living being generated on a computer system, which a user can interact with through an interface.

[0645] "Display means" refers to a device or mechanism for outputting images or animations of virtual creatures so that users can visually recognize them.

[0646] "Operation information acquisition means" refers to a device or software that detects operations performed by a user and acquires that information as digital data.

[0647] "Analysis means" refers to a device or program that analyzes acquired operational information and performs processing to determine an appropriate response from a virtual organism based on that analysis.

[0648] "Motion generation means" refers to a mechanism or software that generates the movements and behaviors of a virtual organism based on the reactions determined by the analysis means.

[0649] "Display control means" refers to a device or software for playing back the movements and animations of a generated virtual creature and presenting them effectively to the user.

[0650] A "generative model" refers to a mathematical model or algorithm that uses machine learning or artificial intelligence techniques to generate natural movements and facial expressions for virtual beings.

[0651] "Information and communication network" refers to a network structure used to send and receive digital data, and includes the internet and dedicated lines.

[0652] This invention relates to a system for realizing interaction with virtual creatures. The user operates a terminal to generate real-time responses from the virtual animal. Specific embodiments of this system are described below.

[0653] The user first launches an application installed on their device. This displays a virtual safari park on the screen, allowing the user to select their favorite virtual animal. For example, if the user selects a lion, a virtual image of that lion will be displayed on the device's screen.

[0654] When a user touches or swipes an animal using the device's touchscreen, these actions are detected, and the device acquires action data such as the location, speed, and direction of the touch. This action data is transmitted to the server in real time.

[0655] The server is responsible for analyzing the received operation data. This analysis uses a generative AI model to generate natural movements of a virtual animal in response to user actions. Specifically, the generative AI model uses the prompt "Generate the lion's reaction when it is petted" to generate the lion's movements and facial expressions.

[0656] Based on the analysis results, the server animates the virtual animal's reactions and sends the generated data to the terminal. The terminal receives this data and displays the virtual lion's movements and expressions in real time in conjunction with the user's actions. This process allows the user to become familiar with the virtual lion and enjoy an interactive experience.

[0657] For example, if a user gently strokes the lion's ears, the server generates a motion where the lion happily wags its tail, and displays the result on the terminal. In this way, the virtual creature responds immediately to the user's actions, aiming to provide a more realistic and educational experience.

[0658] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0659] Step 1:

[0660] The user launches an application on their device. This displays a virtual safari park on the screen, and a screen for selecting an animal appears. When the user selects a lion, its virtual image is displayed on the device's screen. The input is the user's selection information, and the output is the virtual image of the selected animal.

[0661] Step 2:

[0662] The user touches or swipes on the selected animal on the device's touchscreen. The device detects these actions and generates action data, which includes information such as the position, speed, and direction of the action. The input is the user's touch actions, and the output is the action data.

[0663] Step 3:

[0664] The terminal sends the generated operation data to the server. The transmitted data includes details such as the location, speed, and direction of the touch. The input is the operation data, and the output is the transmission of data to the server.

[0665] Step 4:

[0666] The server analyzes the received operation data. This analysis process uses a generative AI model to determine the animal's response using prompt statements. For example, based on the prompt statement "Generate the lion's response when petted," it generates natural lion behavior. The input is operation data and prompt statements, and the output is the lion's response data.

[0667] Step 5:

[0668] The server generates animations of virtual animals based on the analysis results. The generated motion data is created by the server along with ambient sounds and sound effects. For example, it might add the actions of a lion squinting and wagging its tail, along with a purring sound. The input is reaction data, and the output is animation data and sound data.

[0669] Step 6:

[0670] The server sends the generated animation and audio data to the terminal. The terminal receives this in real time and displays the movements of the virtual lion in response to user input. The input is animation and audio data, and the output is a real-time display of the virtual animal.

[0671] (Application Example 1)

[0672] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0673] Conventional virtual experience systems have limited interaction with virtual creatures, making it difficult to provide users with sufficient educational value or effective guidance. Furthermore, the insufficient real-time motion generation resulted in unnatural interactions.

[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0675] In this invention, the server includes operation data acquisition means that senses user operations and generates data, analysis means that analyzes the operation data and generates responses from the virtual creature, and guide generation means that provides guidance from the virtual creature based on information related to the user's environment. This enables the user to receive real-time guidance from the virtual creature through an intuitive and flexible interaction with educational value.

[0676] A "virtual organism" is an electronic living entity generated by a computer system, which is displayed as an object that the user interacts with.

[0677] "Display means" refers to a function or device for displaying images and actions of virtual creatures on a user's terminal.

[0678] "User operation" refers to input operations and instructions performed by the user via a terminal, including those detected using touchscreens or sensors.

[0679] "Operation data acquisition means" refers to a function or device for sensing user operations and collecting and generating data based on those operations.

[0680] "Analysis means" refers to a data analysis function or device for generating and determining the responses and actions of a virtual organism in real time based on acquired operational data.

[0681] "Motion generation means" refers to a function or device that generates movements based on the analyzed results in order to give continuity and naturalness to the movements and facial expressions of a virtual organism.

[0682] "Display control means" refers to a control function or device that transmits generated actions or responses to a user terminal and displays them.

[0683] A "guide generation means" is a function or device that uses the user's environmental information to generate guide information for a virtual creature to provide appropriate guidance and explanations.

[0684] The program for the system to implement this invention aims to realize real-time interaction between a virtual organism and a user. This system mainly consists of three elements: a server, a terminal, and a user.

[0685] First, it is assumed that users will use smart glasses or head-mounted displays as smart devices. Users will wear these devices and interact with virtual creatures in a virtual space. The terminals have built-in display means and are equipped with means for acquiring operation data that receives user input. Through this operation, users can touch the virtual creatures and give them instructions.

[0686] The server analyzes user operation data received via operation data acquisition means using a generative AI model. This analysis process determines the virtual creature's response to specific user actions. Based on the analysis results, the server creates corresponding animations and motion data using motion generation means that generate the virtual creature's movements and expressions. This process utilizes a media generation engine to enable natural and smooth movements.

[0687] The generated motion data is sent to the terminal by the display control means and presented to the user in real time. This allows the user to experience intuitive and seamless interaction with the virtual creature. Furthermore, the server includes a guide generation means that analyzes the user's environment information and provides guidance to the user based on that information. This function allows the virtual creature to provide appropriate advice and explanations in real time according to the user's surroundings.

[0688] As a concrete example, if a user interacts with a virtual lion to learn about a specific product, the server will generate actions for the lion to provide a concise and easy-to-understand explanation of that product. The lion will guide the user in a friendly tone, saying, "This product is made from environmentally friendly materials and is a sustainable choice." An example of a prompt would be, "The user has shown interest in a new eco-bag. As the lion guide, what will you tell the user about the product?"

[0689] This allows users to enjoy a rich virtual experience that goes beyond mere visual enjoyment, enabling them to actually learn.

[0690] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0691] Step 1:

[0692] Users access a virtual space using smart devices. Users wear smart glasses or head-mounted displays and interact with virtual beings. The terminal receives user input through an operation data acquisition mechanism, obtaining user actions and location information as input. This data is transmitted to the server.

[0693] Step 2:

[0694] The server analyzes the user's input data using a generated AI model. This analysis process involves data calculations to determine the virtual creature's response to specific user actions. For example, the touch location and direction of movement are input to the AI ​​model as prompts, and the corresponding response from the virtual creature is generated as output.

[0695] Step 3:

[0696] Based on the analysis results, the server uses motion generation tools to generate the movements and facial expressions of the virtual creature. A media generation engine is then used to create smooth and realistic animations. The generated animation data is provided as output from the motion generation tools.

[0697] Step 4:

[0698] The generated motion data and animations are sent to the terminal and presented to the user in real time by a display control system. Based on the received motion data, the terminal allows the user to observe the movements of the virtual creature through the display.

[0699] Step 5:

[0700] The server further analyzes the user's environmental information and uses a guide generation mechanism to enable the virtual creature to provide appropriate guidance to the user. For example, if a user shows interest in a particular product, that information is input as environmental data, and a guide prompt is generated. This prompt is input into a generating AI model, and the virtual creature's guide information is generated as output.

[0701] Step 6:

[0702] Ultimately, the generated guide information and actions are displayed on the terminal, allowing the user to receive real-time guidance and explanations from the virtual creature. Through this, the user can choose to continue the operation or make a purchase decision.

[0703] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0704] This invention relates to a system that, in addition to interaction with a virtual animal, recognizes the user's emotions and provides a personalized experience. This system can provide a more interactive and personalized learning experience by having the virtual animal respond in a way that takes the user's emotional state into account.

[0705] By launching an application on their device, users can access a virtual safari park and select one of the displayed animals. For example, if a lion is selected, a lion will appear on the screen. At this point, the user can begin interacting with the animal through touch controls.

[0706] The device uses cameras and microphones to recognize not only user actions but also the user's facial expressions and tone of voice in real time. This data is analyzed via an emotion engine to estimate the user's emotional state. This information is sent to the server along with the operation data.

[0707] The server generates virtual animal responses that correspond to the user's emotions, based on acquired operation data and estimated emotional information. For example, if the user is smiling, the lion can perform friendly actions, and if the user is feeling down, it can perform encouraging actions.

[0708] Next, the server generates an animation based on the appropriate animal response. Image generation technology is used for this generation, providing continuous and realistic motion. The animation data is compressed and sent to the terminal.

[0709] The device decompresses the transmitted animation data and displays it on the user's screen in real time. This allows users to interact with animals that react differently to their emotions, resulting in a more personalized experience.

[0710] For example, if the emotion engine determines that the user is happy as they touch the lion, the server will generate an animation in which the lion moves in a more friendly manner and behaves in a welcoming manner towards the user. Through this process, the user can enjoy a unique interaction experience.

[0711] This invention makes it possible to provide users with a more empathetic and intuitive learning environment by incorporating emotion recognition technology into interactions with virtual animals.

[0712] The following describes the processing flow.

[0713] Step 1:

[0714] The user launches an application on their device and accesses the virtual safari park environment. A list of animals is displayed on the device, and the user selects a virtual animal they are interested in.

[0715] Step 2:

[0716] The device displays the selected virtual animal on the screen. The user begins interacting with the virtual animal on the touchscreen and performs actions. At the same time, the device also acquires data on the user's facial expressions and voice.

[0717] Step 3:

[0718] The device collects user operation data and simultaneously incorporates facial expression and voice data acquired through the camera and microphone into the emotion engine. This allows the user's emotional state to be analyzed.

[0719] Step 4:

[0720] The terminal combines operational data and analyzed emotional data and sends it to the server. The server analyzes the received information and determines the virtual animal's response based on the user's emotions.

[0721] Step 5:

[0722] The server generates responses that adjust the actions and facial expressions of the virtual animal based on the user's emotional state. For example, if the user seems happy, the lion will perform playful actions, while if the user seems sad, it will generate encouraging actions.

[0723] Step 6:

[0724] Based on the generated animal reactions, the server creates a continuous animation. High-quality, realistic animations are produced using image generation technology.

[0725] Step 7:

[0726] The server sends the generated animation data to the terminal. Since the data is transmitted in real time, the user can receive immediate feedback.

[0727] Step 8:

[0728] The device decompresses the received animation data and displays it immediately on the screen. Users can directly see the virtual animal's reactions in response to their own actions and emotions, resulting in a more empathetic experience.

[0729] (Example 2)

[0730] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0731] Current virtual animal interaction systems have limitations in their response to user input and are unable to provide personalized experiences that incorporate emotional recognition. Therefore, there is a need to provide more personalized interactions that take user emotions into consideration.

[0732] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0733] In this invention, the server includes means for displaying a virtual animal, means for acquiring data by sensing user actions and emotions and generating data, and means for analyzing the acquired action data and emotion data and generating responses for the virtual animal. This allows the virtual animal to show individual responses in accordance with the user's emotions, enabling an interactive and personalized experience.

[0734] A "virtual animal" is a computer-simulated animal created by a computer program that can interact with the user.

[0735] "Display means" refers to technologies and devices for visually displaying information on a user's terminal, including displays and screens.

[0736] "Data acquisition methods" refer to technologies and techniques for detecting user actions and emotional states and collecting them as data.

[0737] "Analysis means" refers to a method or apparatus for determining the response of a virtual animal using a computer algorithm based on acquired data.

[0738] A "generative AI model" is a type of artificial intelligence that learns from large amounts of data and generatively creates new content and responses.

[0739] "Animation generation means" refers to technologies and methods for creating images or videos to continuously represent the movements of a virtual animal.

[0740] "Display control means" refers to technologies and devices that transmit generated animations to the user's terminal and control them to execute them appropriately.

[0741] This invention is a system that provides a personalized experience that takes into account the user's emotions through interaction with virtual animals. Users can launch a specific application on their device, access a virtual safari, select an animal, and begin interacting with it.

[0742] Users can directly control virtual animals using a touch panel. The device uses its built-in camera and microphone to sense the user's facial expressions and voice tone, collecting emotional data in real time. This data is temporarily stored on the device and sent to the emotion engine. The emotion engine analyzes this data using machine learning models to estimate the user's emotional state. The machine learning models used include deep learning techniques.

[0743] The estimated emotion information is sent to the server along with the user's interaction data. The server uses a generative AI model to generate appropriate virtual animal responses based on this information. For example, if the user smiles, the lion is programmed to perform friendly actions. Generative AI models may utilize technologies such as natural language processing or image generation.

[0744] The server generates an animation based on the determined animal's response. The generated animation is then sent from the server to the terminal after its data size is reduced using compression technology. The terminal decompresses the received animation and displays it on the user's screen in real time.

[0745] For example, when a user touches a lion, if the emotion engine detects the user's smile, the server generates an animation of the lion welcoming the user. Through such interactions, users can enjoy a personalized experience that responds to their emotions in real time.

[0746] An example of a prompt might be, "Create an animation showing how a lion behaves in a friendly manner when a user touches it with a smile." Through this prompt, the generative AI model can generate natural reactions that meet the user's expectations.

[0747] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0748] Step 1:

[0749] The user launches the application on their device and accesses a virtual safari. From the animals displayed, they select, for example, a lion. The selection is made by tapping on the touch panel. The input here is the user's selection, and the selected animal is displayed on the screen as output.

[0750] Step 2:

[0751] The device activates its built-in camera and microphone to sense the user's facial expressions and voice tone in real time. This allows for the collection of emotional data. The user's facial expressions and voice data are used as input, and emotional data is generated as output for transmission to the emotion engine.

[0752] Step 3:

[0753] The device sends the collected emotional data to the emotion engine, which analyzes the data using a machine learning model. The analysis then estimates the user's emotional state. The input is emotional data, and the output is estimated emotional information. This estimated information, along with the user's interaction data, is used in the next step.

[0754] Step 4:

[0755] The server receives operation data and emotion information sent from the terminal. Using a generative AI model, it generates a virtual animal response that corresponds to the user's emotions. Specifically, if the user is smiling, it generates an action where the lion approaches in a friendly manner. The input is operation data and emotion information, and the output is virtual animal response data.

[0756] Step 5:

[0757] The server creates animations based on the generated reaction data. Image generation technology is used to achieve realistic movement. The input is the reaction data of a virtual animal, and the output is compressed animation data.

[0758] Step 6:

[0759] The device receives and decompresses compressed animation data sent from the server. It then displays the animation on the user's screen in real time. The user can enjoy interacting with an animal that responds to emotions. The input is compressed animation data, and the output is the animal's actions displayed on the user's screen.

[0760] (Application Example 2)

[0761] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0762] In modern production environments, there is a need for efficient work support that takes into account the emotions and physical condition of workers. However, conventional systems have difficulty responding appropriately to changes in workers' emotions, limiting improvements in work efficiency and safety. Therefore, there is a need for a system that can recognize workers' emotional states in real time and provide individualized and appropriate support accordingly.

[0763] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0764] In this invention, the server includes means for displaying a virtual organism, means for acquiring operation information that senses operations and generates information, and means for recognizing the user's emotional state. This makes it possible to provide appropriate work support immediately according to the user's emotional state.

[0765] A "virtual organism" refers to a living being simulated on a computer, which exhibits various reactions and actions through interaction with the user.

[0766] "Display means" refers to functions that visualize virtual creatures on a computer screen or device, and includes technologies for presenting information to the user in real time.

[0767] "Operation information acquisition means" refers to a function that detects the user's physical operations and inputs and collects and processes information generated based on them.

[0768] "Analysis means" refers to a function that includes a process for analyzing acquired operational information and emotional state information to generate appropriate responses and actions for the virtual organism.

[0769] "Visual representation generation means" refers to technology for visualizing the actions and reactions of virtual organisms and presenting them to the user as seamless and realistic images.

[0770] "Display control means" refers to a technology that has the function of appropriately transmitting the generated visual representation to the user terminal and displaying it on the screen.

[0771] "Emotion recognition means" refers to a function that detects and analyzes the user's emotional state from their facial expressions, tone of voice, etc.

[0772] "Work support means" refers to a process that includes functions to provide specific work support and interaction tailored to the user's emotional state.

[0773] This invention is a system that recognizes the emotional state of a user and provides individualized and appropriate support tailored to that user. This system utilizes virtual organisms and is designed to support users in working efficiently and safely in their work environment.

[0774] The server uses emotion recognition means to analyze data collected from the camera and microphone of the user's smart device. This analysis uses voice analysis software and facial recognition technology as examples. It detects changes in the user's facial expressions and tone of voice in real time and determines their emotional state. This information is processed via operation information acquisition means and treated as basic data for the virtual organism to show appropriate responses as needed.

[0775] The terminal receives visual representation data transmitted from the server and displays it on the user's screen. The visual representation generation means provides seamless and realistic movement of the virtual creature, which the user can visually confirm. For example, if the user shows signs of fatigue, the virtual creature displays encouraging messages or suggestions for improving work methods. This interaction allows the user to relax and increase efficiency while working.

[0776] As a concrete example, if the emotion recognition system detects signs of stress while the user is working, the server generates a message such as "We recommend taking a break to relax" and presents it to the user as a visual representation. An example of a prompt message to the generating AI model would be: "We have emotion analysis data for the user. He appears to be tired at the moment. Please generate support or guidance to prompt the system."

[0777] This makes it possible to provide users with a better work environment and personalized support.

[0778] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0779] Step 1:

[0780] The user puts on a smart device and begins working. The device activates its camera and microphone to capture the user's facial expressions and voice data. The input consists of video and audio data, and the output is prepared for analysis of this data. The device collects this data from sensors in real time.

[0781] Step 2:

[0782] The terminal sends the collected data to the server. The server processes the received video and audio data using analysis software. The input is the raw video and audio data sent from the terminal, and the output is extracted feature data. Specifically, the server uses facial recognition technology to extract the user's facial features and voice analysis technology to analyze their tone.

[0783] Step 3:

[0784] Based on the analysis results, the server uses a generative AI model to estimate the user's emotional state. This model combines facial and vocal features to generate emotion labels. The input is feature data, and the output is an emotional state (e.g., stress, exhilaration, fatigue). The server accurately estimates the emotional state and stores it in a database for the next step.

[0785] Step 4:

[0786] The server generates interactions with the virtual creature based on its estimated emotional state. Specific actions and messages are determined by a generating AI model. Inputs are emotional states and past interaction data, while output is the virtual creature's response (e.g., soothing actions or relaxing messages). The server analyzes the user's past emotions and selects the optimal response.

[0787] Step 5:

[0788] The server visualizes the reactions of the generated virtual organism through a visual representation generation system and transmits that data to the terminal. The input is the reactions of the virtual organism, and the output is effective visual representation data. The server performs the necessary processing for visualization, compresses the data, and transmits it.

[0789] Step 6:

[0790] The terminal decompresses the received visual representation data and displays it on the user's device. The input is compressed visual representation data, and the output is the animation and interaction that the user actually sees. The terminal displays the virtual creature's actions on the home screen and prepares to take further action information to check the user's response.

[0791] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0792] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0793] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0794] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0795] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0796] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0797] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0798] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0799] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0800] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0801] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0802] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0803] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0804] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0805] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0806] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0807] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0808] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0809] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0810] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0811] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0812] The following is further disclosed regarding the embodiments described above.

[0813] (Claim 1)

[0814] A means of displaying virtual animals,

[0815] Operation data acquisition means that senses user operations and generates data,

[0816] An analysis means that analyzes operational data and generates the response of a virtual animal,

[0817] An animation generation means that generates an animation of a virtual animal based on the generated reaction,

[0818] A display control means for sending and displaying animations on a user terminal,

[0819] A system that includes this.

[0820] (Claim 2)

[0821] The system according to claim 1, comprising an analysis means for analyzing data generated based on user operations in real time and instantly generating the response of a virtual animal.

[0822] (Claim 3)

[0823] The system according to claim 1, comprising an animation generation means that provides seamless animation using an image generation engine to continuously generate the movements and facial expressions of a virtual animal.

[0824] "Example 1"

[0825] (Claim 1)

[0826] A means of displaying a virtual creature,

[0827] Operation information acquisition means that senses the user's actions and generates information,

[0828] An analytical means that analyzes operational information and generates the response of a virtual organism,

[0829] A motion generation means that generates the movement of a virtual organism based on the generated reaction,

[0830] A display control means that transmits and displays the movement to the user device,

[0831] A means of generating natural movements and facial expressions using a generative model,

[0832] A means of transmitting motion information via an information and communication network,

[0833] A system that includes this.

[0834] (Claim 2)

[0835] The system according to claim 1, comprising analytical means for analyzing information generated based on user operations in real time and instantly generating responses of a virtual organism.

[0836] (Claim 3)

[0837] The system according to claim 1, comprising motion generation means that provides seamless movement using image generation technology in order to continuously generate the actions and facial expressions of a virtual organism.

[0838] "Application Example 1"

[0839] (Claim 1)

[0840] Means of displaying virtual creatures,

[0841] Operation data acquisition means that senses user operations and generates data,

[0842] An analytical means for analyzing operational data and generating responses from a virtual organism,

[0843] A motion generation means that generates the actions of a virtual organism based on the generated response,

[0844] A display control means that transmits and displays the operation on the user terminal,

[0845] A guide generation means that provides a guide for a virtual creature based on information related to the user's environment,

[0846] A system that includes this.

[0847] (Claim 2)

[0848] The system according to claim 1, comprising an analysis means for instantly analyzing data generated based on user operations and instantly generating a response from a virtual organism.

[0849] (Claim 3)

[0850] The system according to claim 1, comprising motion generation means that provides smooth motion using a media generation engine in order to continuously generate the movements and facial expressions of a virtual organism.

[0851] "Example 2 of combining an emotion engine"

[0852] (Claim 1)

[0853] A means of displaying virtual animals,

[0854] A data acquisition means that senses user actions and emotions and generates data,

[0855] An analysis means that analyzes acquired operation data and emotional data to generate responses of a virtual animal,

[0856] An animation generation means that generates animations of virtual animals based on generated responses using a generative AI model,

[0857] A display control means that compresses the animation, transmits it to the user terminal, and controls its display,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, comprising an analysis means for analyzing data generated based on user operations and emotions in real time and instantly generating movements that respond to the emotions of a virtual animal.

[0861] (Claim 3)

[0862] The system according to claim 1, comprising an animation generation means that provides seamless animation using a generation AI model and image generation technology to continuously and realistically generate the movements and facial expressions of a virtual animal.

[0863] "Application example 2 when combining with an emotional engine"

[0864] (Claim 1)

[0865] Means of displaying virtual creatures,

[0866] Operation information acquisition means that senses operations and generates information,

[0867] An analytical means that analyzes operational information and generates the response of a virtual organism,

[0868] A visual representation generation means that generates a visual representation of a virtual organism based on the generated reaction,

[0869] A display control means that transmits and displays a visual representation on the user's terminal,

[0870] An emotion recognition means for recognizing the emotional state of the user,

[0871] Work support tools for providing support according to emotional state,

[0872] A system that includes this.

[0873] (Claim 2)

[0874] The system according to claim 1, comprising an analysis means for analyzing information generated in real time based on the user's actions and emotional state, and for instantly generating responses and work support for a virtual organism.

[0875] (Claim 3)

[0876] The system according to claim 1, comprising a visual representation generation means that provides seamless visual representation using visual generation technology in order to continuously generate the movements and facial expressions of a virtual organism. [Explanation of Symbols]

[0877] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of displaying virtual animals, Operation data acquisition means that senses user operations and generates data, An analysis means that analyzes operational data and generates the response of a virtual animal, An animation generation means that generates an animation of a virtual animal based on the generated reaction, A display control means for sending and displaying animations on a user terminal, A system that includes this.

2. The system according to claim 1, comprising an analysis means for analyzing data generated based on user operations in real time and instantly generating the response of a virtual animal.

3. The system according to claim 1, comprising an animation generation means that provides seamless animation using an image generation engine to continuously generate the movements and facial expressions of a virtual animal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A