System
A system that analyzes pet data to generate an AI model for virtual interaction addresses pet loss distress, providing a realistic and emotional support experience.
Patent Information
- Application Number
- JP2024121460
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
The increasing number of pet losses is causing significant psychological distress, necessitating methods to ease this burden and help individuals cope with the loss.
A system that uploads pet image and video data to a server for analysis, generating an AI model that recreates the pet's appearance and behavior in a virtual space, allowing users to interact with their pet through VR or AR devices.
The system provides a realistic and moving experience, helping users cope with pet loss by recreating memories and interactions with their pets, offering psychological support.
Smart Images

Figure 2026019712000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, the number of people keeping pets such as dogs and cats at home has increased, but the problem of pet loss, which can cause serious psychological shock when parted with a pet, is also on the rise. Pet loss can be so psychologically taxing that it requires psychological treatment, so there is a need for methods to ease this burden and help people gradually come to terms with parting with their pets. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for uploading image and video data of pet animals from a terminal to a server, a means for the server to analyze the uploaded data and generate an AI model that captures the animal's characteristics, and a means for recreating the generated AI model in a virtual space to simulate the animal's life. Furthermore, the server includes a means for facial recognition, body shape identification, and extraction of movement patterns of the animal from the image and video data. In the virtual space, the AI model operates and responds to user operations, allowing the user to enjoy interacting with their pet. This is expected to not only ease the shock of pet loss but also gradually make it easier to accept the loss.
[0006] A "server" is a computer system that processes and stores data over a network.
[0007] A "terminal" is a device used by a user to input and receive information, including a smartphone, PC, etc.
[0008] "Image and video data" refers to information in the form of photographs and movies of pets, and is basic data for capturing the characteristics of pets visually and aurally.
[0009] "Upload" is the act of a user sending data from a terminal to a server.
[0010] "Analysis" is the process by which the server extracts and deciphers the characteristics of animals based on the uploaded image and video data.
[0011] An "AI model" is a collection of digital data and programs that uses artificial intelligence technology to reproduce the appearance and movements of an animal.
[0012] A "virtual space" is a virtual environment created using computer technology in which users can interact.
[0013] "Simulation" means reproducing the behavior and characteristics of real animals in a virtual space.
[0014] "Facial recognition" is a technology that uses image analysis technology to identify an animal's face and capture its characteristics.
[0015] "Body shape identification" is a technique that uses image analysis technology to identify the shape and size of an animal's body and capture its characteristics.
[0016] "Movement patterns" are data that indicate the physical movements and specific behaviors of animals.
[0017] "User" refers to a person who uses the system to upload images and videos of their pet and interact with the AI model recreated in a virtual space. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses image and video data of pets and analyzes them to generate an AI model that reproduces the appearance and behavior of the pet, and then recreates it in a virtual space.
[0040] Explanation of system program processing
[0041] 1. Data collection and upload
[0042] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[0043] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0044] 2. Data processing and AI model generation
[0045] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[0046] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[0047] The server generates an AI model that replicates the pet's appearance and behavior based on the analysis results, and the generated AI model is tuned to replicate the pet's characteristics as faithfully as possible.
[0048] Example: The server creates a 3D model based on identified face, body shape, and fur color data, reproduces animal sounds using audio data, and generates animations based on movement patterns.
[0049] 3. Providing AI models and verifying their operation
[0050] The server sends the generated AI model to the user's device, which must support a VR headset or AR app.
[0051] Example: A server sends the generated AI model, including the 3D model and audio files, in a format compatible with the user's VR headset.
[0052] Users can then load the AI model using VR or AR devices, reunite with their pet in a virtual space, and enjoy interacting with it.
[0053] Example: A user puts on a VR headset, launches a system-specific app, loads a 3D model of their pet dog, and plays with it in a virtual garden.
[0054] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of their pet, providing users with a highly realistic and moving experience.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[0058] Step 2:
[0059] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[0060] Step 3:
[0061] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[0062] Step 4:
[0063] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[0064] Step 5:
[0065] The server begins analyzing the stored image and video data, which includes:
[0066] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[0067] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[0068] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[0069] Step 6:
[0070] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[0071] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0072] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0073] Movement animation: Generates animation based on the extracted movement patterns.
[0074] Step 7:
[0075] The server sends the generated AI model to the user's device. The data sent must be in a format compatible with VR or AR.
[0076] Step 8:
[0077] The user launches a VR headset or AR app and loads the sent AI model, recreating the experience of interacting with a pet in a virtual space.
[0078] Step 9:
[0079] The device operates the AI model in response to user input within the virtual space, allowing interaction with the pet. This recreates the natural presence and vitality of dinosaurs. This scenario is translated into English.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] Pet loss is a serious psychological issue for many people. The present invention aims to alleviate pet loss by using pet image and video data to create an AI model that faithfully reproduces a pet's appearance and behavior. However, conventional technologies have difficulty accurately capturing a pet's characteristics and lack the means to allow users to fully enjoy interacting with a pet recreated in a virtual space. The purpose of the present invention is to solve these problems.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes means for uploading image and video data of the pet from the terminal to the server, means for analyzing the uploaded data and generating an AI model that captures the characteristics of the pet, means for reproducing the generated AI model in a virtual space and simulating the life of the pet, and means for the user to display the generated AI model in the virtual space using a VR or AR device and interact with it. This allows the generation of a highly accurate AI model that faithfully reproduces the characteristics of the pet, allowing the user to enjoy interacting with the pet in the virtual space.
[0085] A "server" is a computing device that receives and analyzes data and generates and serves AI models.
[0086] A "terminal" is a computer device that is operated by a user and that uploads and receives data.
[0087] A "user" is an individual who uses a terminal to provide image and video data of their pet to the server and interacts with the generated AI model.
[0088] "Image and video data" refers to still image and video files that record the appearance and behavior of your pet.
[0089] "Upload" is the act of sending data from a user's terminal to a server.
[0090] "Analysis" is the process by which the server processes the received data to extract and identify the pet's characteristics.
[0091] An "AI model" is an artificial intelligence model that is generated based on analyzed data and is used to reproduce the appearance and behavior of a pet.
[0092] "Virtual space" is a virtual three-dimensional space generated by computer simulation.
[0093] "Simulation" is the process of recreating the life and behavior of a real pet in a virtual space.
[0094] "VR equipment" refers to equipment for experiencing virtual reality, including devices such as headsets.
[0095] An "AR device" is a device for experiencing augmented reality, and includes devices such as smartphones and AR glasses.
[0096] "Interaction" refers to two-way communication and manipulation between the user and a pet recreated in a virtual space.
[0097] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses images and video data of the user's pet and analyzes them to generate an AI model that reproduces the pet's appearance and behavior, then recreates it in a virtual space.
[0098] First, users upload image and video data of their pets to the server from their own devices. Devices can be smartphones, tablets, or PCs, and data can be easily selected and sent via a dedicated upload interface. For example, a user can select 10 photos of dogs and three videos from their smartphone gallery and click the "Upload" button.
[0099] The server then receives the uploaded data and analyzes it. It uses a TensorFlow-based image recognition model to recognize the animal's face, body shape, and coat color from the image. For videos, it uses a time-series data analysis tool to analyze movement patterns, and an audio analysis tool to analyze the animal's barks from the audio data. For example, the server uses an image analysis engine to identify the dog's face and body shape, and a video analysis engine to extract movement patterns. It also uses an audio analysis engine to extract bark characteristics.
[0100] Based on this analytical data, the server generates a highly accurate AI model that reproduces the pet's appearance and behavior. The server uses 3D modeling tools to create a three-dimensional model of the pet and applies movement animations to it. It also adds audio data to reproduce its cries. For example, the server creates a 3D model based on the identified face, body shape, and fur color data, implements movement animations based on movement patterns, and adds cries data.
[0101] The generated AI model is then sent back to the device. The user's device must support a VR headset or AR application. The server checks file compatibility, converts the format if necessary, and then sends the AI model to the device. The user can then load the AI model using a VR or AR device, reunite with their pet in a virtual space, and enjoy interacting with it. For example, a user can put on a VR headset, launch a dedicated app, load a 3D model of their pet, and play with it in a virtual garden.
[0102] An example of a prompt is as follows:
[0103] "I've uploaded photos and videos of my dog. Please generate a realistic 3D model based on them."
[0104] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of a pet, providing users with a highly realistic and moving experience.
[0105] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0106] Step 1:
[0107] The user selects and uploads pet image and video data from the device. Specifically, the user opens the gallery or file manager on the device and selects multiple images and videos. For example, the user selects 10 dog photos and 3 videos from the smartphone gallery and clicks the "Upload" button. The image and video data selected by the user are used as input, and this data is sent to the server as output.
[0108] Step 2:
[0109] The server receives the image and video data sent by the user. Specifically, the server checks the integrity of the data and verifies that all files have been received correctly. The input is the image and video data from the user, and the output is the storage and verification of the data.
[0110] Step 3:
[0111] The server analyzes the received image data. It uses a TensorFlow-based image recognition model to perform facial recognition, body shape identification, and coat color analysis. Specifically, the image analysis engine extracts the facial position, body shape outline, and coat color characteristics for each image. For example, it identifies the facial feature points (eyes, nose, and mouth) of a dog and stores that information in a database. Using the received image data as input, facial recognition and body shape identification data are generated as output.
[0112] Step 4:
[0113] The server analyzes the received video data. It uses a time-series data analysis tool to analyze movement patterns and extract the pet's barks from the audio data. To determine specific behaviors, the video analysis engine analyzes each frame and extracts behavioral patterns while detecting the continuity of movement. The audio analysis engine also extracts sound wave characteristics and detects bark patterns and frequencies. For example, a dog's behaviors such as running, eating, and sleeping are recorded along a time axis. The received video data is used as input, and movement patterns and bark data are generated as output.
[0114] Step 5:
[0115] The server generates an AI model based on the analysis data. 3D modeling tools are used to create a three-dimensional model of the pet, and movement animations are applied to that model. Analyzed audio data is then added to recreate the animal's barks. Specific movements are then converted by the server into a 3D model of the face, body shape, and coat color, and the movement patterns obtained from video analysis are incorporated into the animation using motion capture technology. For example, a dog's running motion and barking sounds are realistically recreated. The analysis results are used as input, and a complete AI model is generated as output.
[0116] Step 6:
[0117] The server sends the generated AI model to the user's device. The transmission format is compatible with VR or AR. Specifically, the server checks file compatibility and converts the format if necessary. The data is then seamlessly transferred to the user's device. The generated AI model is used as input, and the AI model is sent to the user's device as output.
[0118] Step 7:
[0119] The user loads the sent AI model using a VR or AR device. Using a dedicated application, the user interacts with their pet in a virtual space. Specifically, the user puts on a VR headset, launches a dedicated application, and loads the AI model. The user can then play with their pet dog in a virtual garden. The AI model sent from the server is used as input, and interaction in a virtual space is realized as output.
[0120] (Application example 1)
[0121] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0122] To alleviate the psychological shock and sadness caused by pet loss, there is a need for a system that can accurately recreate memories of past pets and allow users to re-interact with their pets in a virtual space. Furthermore, there is a need for a method to enable such a system to provide a moving experience without users having to physically visit a store.
[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0124] In this invention, the server includes means for uploading image and video data of the animal from the terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for reproducing the generated AI model in a virtual reality space and simulating the life of the animal, and means for the user to access the virtual space on the terminal and enjoy interacting with the animal's AI model. This allows the user to ease the loss of their pet and have a moving experience through interacting with the pet in the virtual space.
[0125] 1. "Domestic animals" refers to all animals kept as pets in homes.
[0126] 2. "Image and video data" refers to digital data consisting of photographs and videos of pets.
[0127] 3. "Device" refers to electronic devices that can connect to the Internet, such as smartphones, tablets, and PCs.
[0128] 4. "Server" refers to a central processing unit used for data analysis and AI model generation.
[0129] 5. "Upload" refers to the act of sending digital data from a device to a server.
[0130] 6. “Analysis” refers to the act of extracting information based on image or video data.
[0131] 7. "Feature-capturing AI model" refers to an artificial intelligence model that faithfully reproduces the appearance and behavior of a pet.
[0132] 8. "Virtual reality space" refers to a virtual environment created using VR technology.
[0133] 9. "Simulate" refers to the act of recreating real-world actions or environments.
[0134] 10. "Accessing a virtual space" refers to the act of a user connecting to a virtual reality environment using a VR headset or device.
[0135] 11. “Interaction” refers to the act of a user interacting with an AI model.
[0136] This invention is a system aimed at alleviating the loss of a pet. It analyzes image and video data of pets to generate an AI model of the pet, which is then recreated in a virtual space, allowing users to relive memories with their pet and providing a moving experience.
[0137] This system consists of the following main elements:
[0138] 1. Terminal
[0139] Users use devices such as smartphones or tablets to collect images and video data of their pets and upload them to a server. A dedicated application is installed on the device, allowing users to easily select and send data.
[0140] 2. Server
[0141] The server receives the uploaded image and video data and analyzes it, which includes the following processes:
[0142] Image analysis uses image recognition software such as TensorFlow to recognize pet faces, identify body shapes, and analyze fur color.
[0143] Video analysis uses video analysis software such as OpenCV to extract pet movements and behavior patterns.
[0144] Based on the analysis results, an AI model is generated using a 3D modeling tool such as Unity. This AI model is generated by integrating the analysis data from TensorFlow and OpenCV to faithfully reproduce the appearance and behavior of the pet.
[0145] 3. Virtual Reality Space
[0146] The AI model generated by the server is reproduced in a virtual reality space, which can be accessed through a VR headset or an AR app. The generated AI model is configured to operate in real time in the virtual reality space using VR development tools such as Unity.
[0147] By wearing a VR headset, users can access a virtual reality space and enjoy interacting with their pets.
[0148] 4. User Interaction
[0149] Users can interact with the AI model in real time through a VR headset. For example, if a user calls a pet's name in the virtual space, the AI model will respond by turning toward the user or performing a specific action, providing an experience similar to interacting with a real pet.
[0150] A specific example of this system is shown below.
[0151] Specific examples
[0152] Users upload using their smartphones
[0153] Users select 10 photos of their dog and three videos from their smartphone gallery and upload them to the server using a dedicated app.
[0154] Server analysis process
[0155] The server uses TensorFlow to analyze images and OpenCV to analyze videos, which allows for facial recognition, body shape identification, sound analysis, and movement pattern extraction of pets.
[0156] Generating AI models
[0157] Based on the analysis results, the server uses Unity to generate a 3D model and create an AI model.
[0158] Reproduction in virtual reality space
[0159] Users put on a VR headset and access the virtual reality space. They then launch the system's dedicated app, load a 3D model of their pet dog, and play with it in a virtual garden.
[0160] Prompt Sentence Examples
[0161] AI model generation prompts
[0162] Please generate a 3D model that is as faithful as possible to my pet's appearance based on past images and videos of my pet.
[0163] Interaction prompts
[0164] Place the generated 3D model in your virtual garden and simulate its movements as you play with your pet.
[0165] In this way, users can ease the loss of their pet and relive their memories with their pet through an emotional experience in a virtual reality space.
[0166] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0167] Step 1:
[0168] The user uploads pet image and video data using a device. Specifically, the user launches a dedicated application on the device, selects pet photos and videos from the gallery, and clicks the "Upload" button. This sends the selected data to the server. The input is the pet image and video data, and the output is the data transfer to the server.
[0169] Step 2:
[0170] The server receives the uploaded image and video data. The received data is temporarily stored in the server's storage. The input is the image and video data sent by the user, and the output is the data saved in the server's storage.
[0171] Step 3:
[0172] The server uses TensorFlow to analyze the received image data. First, the server preprocesses the image data and performs facial recognition, body shape identification, and fur color analysis. Specifically, the server converts the image data into RGB format and inputs each data into a model to extract features. The input is the saved image data, and the output is feature data related to the pet's facial recognition, body shape, and fur color.
[0173] Step 4:
[0174] The server uses OpenCV to analyze the received video data. The server extracts motion patterns from the video data frame by frame and generates time-series data. Specifically, it extracts frames from the video data and detects motion changes using methods such as optical flow. The input is the saved video data, and the output is time-series data on motion patterns.
[0175] Step 5:
[0176] The server integrates the analyzed image and video data to generate an AI model that captures the pet's characteristics. To do this, a 3D model is built using Unity based on the analysis results of TensorFlow and OpenCV, and movement patterns are added. The input is feature data and time-series data, and the output is the pet's AI model.
[0177] Step 6:
[0178] The server sends the generated AI model to the user's device. Specifically, the generated AI model is sent as data in a format compatible with VR headsets and AR apps. The input is the generated AI model, and the output is data transferred to the user's device (VR headset or smartphone).
[0179] Step 7:
[0180] The user puts on a VR headset and launches a dedicated app to load the generated AI model. Specifically, the user selects the sent AI model from the VR headset app and accesses the virtual space. To interact with the pet in the VR space, the user interacts with the AI model through voice recognition and gesture control. The input is the AI model and the VR headset, and the output is the interaction with the pet in the virtual space.
[0181] Through these steps, the system can ease the loss of a pet and provide a moving experience for users.
[0182] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0183] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[0184] Explanation of system program processing
[0185] 1. Data collection and upload
[0186] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[0187] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0188] 2. Data processing and AI model generation
[0189] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[0190] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[0191] The server uses the analysis results to generate an AI model that replicates the pet's appearance and behavior. This model includes the following elements:
[0192] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0193] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0194] Movement animation: Generates animation based on the extracted movement patterns.
[0195] 3. Sentiment analysis and AI model tuning
[0196] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[0197] Example: The server uses a facial recognition algorithm to detect a user's smiling or sad facial expressions, and a speech analysis tool to analyze changes in the tone of their voice.
[0198] 4. Providing AI models and verifying their operation
[0199] The server generates an AI model and sends the emotion analysis results to the user's device. The data is in a format compatible with VR or AR.
[0200] Example: A server sends data containing the generated 3D model, audio files, and emotion analysis results to a user's VR headset.
[0201] The user launches a VR headset or AR app and loads the AI model, recreating the experience of interacting with a pet in a virtual space. Furthermore, the AI model's behavior changes depending on the user's emotional state.
[0202] Example: A user puts on a VR headset, launches a dedicated app for the system, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[0203] This system not only helps users ease the shock of losing a pet and recreate memories with their pet, but also allows them to enjoy interactions that are in tune with their emotions. The introduction of an emotion engine allows users to have a more realistic and moving experience, and provides psychological support to help them move on to a new life.
[0204] The processing flow will be explained below.
[0205] Step 1:
[0206] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[0207] Step 2:
[0208] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[0209] Step 3:
[0210] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[0211] Step 4:
[0212] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[0213] Step 5:
[0214] The server begins analyzing the stored image and video data, which includes:
[0215] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[0216] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[0217] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[0218] Step 6:
[0219] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[0220] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0221] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0222] Movement animation: Generates animation based on the extracted movement patterns.
[0223] Step 7:
[0224] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[0225] Example: The server uses a facial recognition algorithm to detect the user's facial expressions (e.g., smiling, sad), and a voice analysis tool to analyze changes in voice tone.
[0226] Step 8:
[0227] The server then sends the generated AI model and emotion analysis results to the user's device in a format compatible with VR or AR.
[0228] Step 9:
[0229] The user launches the VR headset or AR app and loads the AI model, which then begins interacting with the pet in the virtual space.
[0230] Example: A user puts on a VR headset, launches a dedicated app, and loads a 3D model of their pet dog. If the user is smiling, the AI model of the pet will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of the pet will quietly cuddle with the user.
[0231] This series of processes not only helps users ease the shock of losing a pet, but also provides an emotionally responsive, interactive experience.
[0232] Example 2
[0233] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0234] Users who are grieving the loss of their pet need a system that allows them to not only reminisce about their memories with their deceased pet, but also enjoy emotionally rich interactions. However, existing systems lack the functionality to empathize with the user's emotions and are insufficient in providing psychological support. Therefore, a system that dynamically analyzes the user's emotions and adjusts the behavior of the pet's AI model to provide more emotionally rich interactions is desired.
[0235] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0236] In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for recreating the generated AI model in a virtual space and simulating the life of the animal, means for activating an emotion engine that detects a user's emotions and analyzing emotions from the user's facial expressions and voice using a camera and microphone, and means for adjusting the operation of the AI model based on the emotion analysis results, thereby enabling users to ease the shock of losing their pet and enjoy emotionally sensitive interactions.
[0237] The "server" is a central processing unit that receives and analyzes data uploaded from the user's device, and generates and adjusts the AI model.
[0238] A "terminal" is a device used by a user, such as a computer or smartphone, through which data is uploaded or downloaded.
[0239] "Upload" means a function or interface for transmitting image or video data from a terminal to a server.
[0240] "Analysis" refers to the process of processing the data received by the server and extracting the animal's characteristics and behavioral patterns.
[0241] An "AI model" is an artificial intelligence model that is generated based on analyzed data and reproduces the appearance and behavior of animals.
[0242] A "virtual space" is an imaginary environment generated using computer graphics and related technologies.
[0243] The "simulation" means virtually recreating the generated AI model in a virtual space and imitating the animal's life.
[0244] An "emotion engine" is a collection of software and algorithms for analyzing a user's emotions.
[0245] A "camera" is a photographing device for capturing the user's facial expression.
[0246] A "microphone" is a recording device for capturing the user's voice.
[0247] "Emotion analysis" is the process of determining a user's current emotional state based on data captured using a camera or microphone.
[0248] "Behavior adjustment" is the process of changing the behavior and reactions of an AI model based on the results of emotion analysis.
[0249] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[0250] The system of the present invention mainly uses the following hardware and software:
[0251] Hardware: Servers, devices (smartphones, PCs, VR headsets), cameras, microphones
[0252] Software: TensorFlow-based image recognition models, voice analysis tools, time-series data analysis tools, 3D modeling software, voice synthesis technology, emotion recognition algorithms
[0253] Specific embodiments will be described below.
[0254] Data collection and upload
[0255] Users select images and videos of their pets from their devices and send them to the server using the system's upload interface. For example, a user selects 10 dog photos and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0256] Data processing and AI model generation
[0257] The server receives the uploaded images and videos and stores them in the initial data storage. It then uses a TensorFlow-based image recognition model to analyze the images and perform facial recognition, body shape identification, and fur color analysis of the pet. It also analyzes the video to extract the pet's movements and behavior patterns. Based on these analysis results, the server generates an AI model with the following elements:
[0258] 1. 3D model: A 3D model is created using the analyzed face, body shape, and fur color data.
[0259] 2. Audio file: The extracted call data is reproduced using voice synthesis technology.
[0260] 3. Motion animation: Generate animation based on the extracted motion patterns.
[0261] Sentiment analysis and AI model tuning
[0262] The server activates an emotion engine to monitor the user's emotions, using a camera and microphone to capture the user's facial expressions and voice in real time. For example, the camera captures the user's face and uses a facial recognition algorithm to detect smiling or sad expressions. The audio recorded by the microphone is then analyzed for changes in tone using a voice analysis tool. As a result, the server adjusts the behavior of the AI model according to the user's emotional state. For example, if the user is smiling, the AI pet will move around energetically. Conversely, if the user is sad, the AI pet will quietly cuddle.
[0263] Providing AI models and verifying their operation
[0264] The server sends the generated AI model and emotion analysis results to the user's device (e.g., a VR headset). The user then launches the VR headset or AR app and loads the sent AI model, recreating interaction with a pet in a virtual space. The AI pet's behavior changes depending on the user's emotional state. For example, a user puts on a VR headset, launches a system-specific app, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[0265] Prompt Sentence Examples
[0266] "I want to be reunited with my pet. Please use these photos and videos to create an AI model."
[0267] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0268] Step 1:
[0269] User selects data and uploads it on the device
[0270] The user selects images and video data of their pet from their device (smartphone or PC).
[0271] What happens: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0272] Input: Image and video files stored in your phone's gallery.
[0273] Output: Image and video data sent to the server.
[0274] Step 2:
[0275] The server receives and stores the data
[0276] The server stores the image and video data received from the terminal in an initial data storage.
[0277] Specific operation: The server saves the uploaded 10 images and 3 videos to the file system or database.
[0278] Input: Image and video data submitted by the user.
[0279] Output: Image and video files saved to storage.
[0280] Step 3:
[0281] The server analyzes the data
[0282] The server analyzes the stored data using a TensorFlow-based image recognition model.
[0283] How it works: The server uses image recognition models to identify pet faces, body shapes, and coat colors, and uses time-series data analysis tools to extract movement and behavior patterns from video.
[0284] Input: Stored image and video data.
[0285] Output: Face recognition results, body shape identification results, coat color data, movement pattern data.
[0286] Step 4:
[0287] The server generates the AI model
[0288] Based on the analysis results, the server generates an AI model that includes a 3D model, audio files, and motion animations.
[0289] Specific movements: The server uses 3D modeling software to create a 3D model based on the pet's face, body shape, and fur color data. It then uses voice synthesis technology to reproduce the pet's barking sound and generates movement animations based on the movement pattern data.
[0290] Input: Face recognition results, body shape identification results, coat color data, movement pattern data.
[0291] Output: 3D model, reconstructed audio files, movement animations.
[0292] Step 5:
[0293] The server starts the emotion engine
[0294] The server detects the user's emotions by capturing the user's facial expressions and voice in real time using the device's camera and microphone.
[0295] How it works: The server takes a picture of the user's face with the device's camera, uses a facial recognition algorithm to detect smiling or sad expressions, and uses a voice analysis tool to analyze changes in the tone of the voice recorded by the microphone.
[0296] Input: Video and audio recording of the user's face.
[0297] Output: User sentiment analysis results.
[0298] Step 6:
[0299] The server adjusts the AI model
[0300] The server adjusts the behavior of the generated AI model based on the results of emotion analysis.
[0301] Specific behavior: Based on the results of emotion analysis, the AI model's animations and reactions are changed in real time. For example, if the user is smiling, the AI pet will move around energetically and wag its tail. If the user is sad, the AI pet will quietly cuddle.
[0302] Input: User sentiment analysis results, generated AI model.
[0303] Output: The movement animations and reactions of the tuned AI model.
[0304] Step 7:
[0305] The server sends the AI model to the device.
[0306] The server sends the generated AI model and emotion analysis results to the user's device.
[0307] Specific operation: The server sends data including the generated 3D model, audio files, movement animations, and emotion analysis results to the user's VR headset.
[0308] Input: Tuned AI model, sentiment analysis results.
[0309] Output: Data sent to the user's device.
[0310] Step 8:
[0311] The user launches the AI model and checks its operation
[0312] The user launches a VR headset or AR app and displays the sent AI model in a virtual space.
[0313] Specific operation: The user puts on the VR headset, launches the system's dedicated app, loads the 3D model of the pet, enjoys interacting with the AI pet model, and sees its behavior change in real time.
[0314] Input: User's emotional state, transmitted data.
[0315] Output: Interaction with a pet in a virtual space.
[0316] This system allows users to relive memories with their pets while enjoying emotionally rich interactions.
[0317] (Application example 2)
[0318] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0319] There is a need for a system that can vividly recreate memories of pets, empathize with the user's emotional state, and provide real-time interaction while easing the psychological shock of pet loss. However, conventional systems lack the technology to flexibly respond to changes in the user's emotions, making it difficult to achieve more emotionally rich interactions.
[0320] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for the server to use an emotion engine that analyzes the user's emotional state and dynamically adjust the behavior of the generated AI model, and means for reproducing the generated AI model in a virtual space, simulating the life of the animal, and interacting with the animal in accordance with the user's emotional state. This makes it possible to provide more emotionally rich interactions in real time that are in line with the user's emotional state while easing the loss of a pet.
[0321] "Image" is data containing static visual information of an animal.
[0322] "Video data" refers to data that includes dynamic visual and audio information such as animal movements and sounds.
[0323] "Terminal" refers to a device used by a user, such as a smartphone, tablet, or PC.
[0324] A "server" is a computer system that receives, analyzes, stores, and transmits data over a network.
[0325] "Analysis" is the process of extracting necessary information from image and video data and identifying its features.
[0326] An "AI model" is an artificial intelligence software model created to reproduce an animal's shape, movement patterns, sounds, etc.
[0327] An "emotion engine" is a technology for analyzing a user's emotional state from their facial expressions and voice.
[0328] "Dynamic adjustment" refers to changing the behavior of an AI model in response to the user's emotional state in real time.
[0329] A "virtual space" is a virtual environment created using computer technology.
[0330] "Simulation" means virtually reproducing the life and behavior of real animals.
[0331] "Interaction" refers to the interaction between a user and an AI model.
[0332] "Real-time" refers to processing occurring immediately, without delay.
[0333] The embodiment of this invention is a system that allows a user to upload images and video data of animals kept by the user from a terminal to a server, and the server analyzes this data and generates an AI model that captures the characteristics of the animal.Furthermore, the system uses an emotion engine to analyze the user's emotional state and dynamically adjusts the behavior of the generated AI model according to the user's emotions.
[0334] Specifically, users upload images and video data of their pet animals to a server from their smartphones, tablets, or PCs. The uploaded data is then analyzed on the server using image analysis models such as TensorFlow to recognize the animal's face, body shape, and movement patterns. If audio data is included, audio analysis tools are also used to analyze the animal's cries.
[0335] The analyzed data is then used to create an AI model that recreates the animal's appearance, movement patterns, and sounds, including a 3D model, movement animations, and audio files.
[0336] To analyze the user's emotional state using the emotion engine, the server uses camera footage and audio data sent from the user's device. It uses facial recognition algorithms to read the user's facial expressions and audio analysis tools to analyze the tone and content of the voice, dynamically detecting whether the user is happy, sad, or expressing other emotions.
[0337] The generated AI model is then recreated in a virtual space, and its behavior is dynamically adjusted according to the user's emotional state. For example, if the user is smiling, the AI model will move around energetically and wag its tail. Conversely, if the user is sad, the AI model will quietly take on a comforting behavior.
[0338] This system allows users to vividly relive their memories of their pet while easing the loss of their pet. It can provide real-time interactions that are tailored to the user's emotional state, resulting in a more emotionally enriching experience.
[0339] Specific examples
[0340] A user launches the virtual store app, takes a video of their pet dog using their smartphone camera, and uploads it to the server. The server then uses TensorFlow models to analyze the images and sounds and generate an AI model. This AI model operates in the virtual space based on the user's emotional state.
[0341] Prompt Sentence Examples
[0342] The system will be provided with an image of a dog and a short video of its barking. It will then analyze the specific behaviors and sounds the dog exhibits and create a 3D model of it. Furthermore, the system will have the dog's movements change in the virtual space depending on the user's emotions.
[0343] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0344] Step 1:
[0345] A user uses a device (smartphone, tablet, PC, etc.) to take or select images and video data of their pet dog and upload them to a server through an application. The input is the image and video data of the pet dog, and the output is the image and video files transferred to the server. This process involves the user selecting the data through the app interface and clicking the upload button.
[0346] Step 2:
[0347] The server receives the uploaded image and video data and begins analyzing it using an image analysis model such as TensorFlow. The input is the uploaded image and video data, and the output is the analyzed animal's features (face recognition, body shape identification, movement patterns, etc.). This process involves the server running an image recognition algorithm to extract the animal's face and body shape from the image and video.
[0348] Step 3:
[0349] The server generates an AI model based on the analysis results. The input is the analyzed animal's feature data, and the output is the generated AI model (3D model, movement animation, audio file, etc.). This process involves the server using a 3D model generation tool to build a 3D model that reproduces the animal's features and creating animations that include movement patterns and sounds.
[0350] Step 4:
[0351] The server uses an emotion engine to analyze the camera footage and audio data sent from the user's device and analyze the user's emotional state in real time. The input is the user's camera footage and audio data, and the output is the user's emotional state (happiness, sadness, etc.). This process involves the emotion analysis algorithm analyzing facial expressions and tone of voice to identify emotions such as joy, anger, sadness, and happiness.
[0352] Step 5:
[0353] The server dynamically adjusts the behavior of the generated AI model according to the user's emotional state. The input is the user's emotional state and the generated AI model, and the output is the adjusted AI model (a behavior pattern according to the emotional state). This process involves the server changing the behavior of the AI model based on the emotion analysis results and executing the corresponding behavior (for example, if the user is smiling, the AI model behaves cheerfully).
[0354] Step 6:
[0355] The server sends the generated AI model and adjusted behavior patterns to the user's device. The input is the adjusted AI model data, and the output is the AI model displayed on the user's device. This process includes the server encoding the data and sending it in a format suitable for the user's VR headset or AR app.
[0356] Step 7:
[0357] The user uses a device (such as a VR headset, smart glasses, or smartphone) to recreate the transmitted AI model in a virtual space and interact with the AI model according to the user's emotional state. The input is the adjusted AI model data and the user's emotional state, and the output is the user's interaction with an emotive virtual pet. This process involves the user operating the device and interacting with the AI model in the virtual space.
[0358] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0359] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0360] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0361] [Second embodiment]
[0362] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0363] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0364] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0365] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0366] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0367] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0368] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0369] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0370] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0371] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0372] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0373] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0374] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses image and video data of pets and analyzes them to generate an AI model that reproduces the appearance and behavior of the pet, and then recreates it in a virtual space.
[0375] Explanation of system program processing
[0376] 1. Data collection and upload
[0377] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[0378] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0379] 2. Data processing and AI model generation
[0380] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[0381] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[0382] The server generates an AI model that replicates the pet's appearance and behavior based on the analysis results, and the generated AI model is tuned to replicate the pet's characteristics as faithfully as possible.
[0383] Example: The server creates a 3D model based on identified face, body shape, and fur color data, reproduces animal sounds using audio data, and generates animations based on movement patterns.
[0384] 3. Providing AI models and verifying their operation
[0385] The server sends the generated AI model to the user's device, which must support a VR headset or AR app.
[0386] Example: A server sends the generated AI model, including the 3D model and audio files, in a format compatible with the user's VR headset.
[0387] Users can then load the AI model using VR or AR devices, reunite with their pet in a virtual space, and enjoy interacting with it.
[0388] Example: A user puts on a VR headset, launches a system-specific app, loads a 3D model of their pet dog, and plays with it in a virtual garden.
[0389] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of their pet, providing users with a highly realistic and moving experience.
[0390] The processing flow will be explained below.
[0391] Step 1:
[0392] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[0393] Step 2:
[0394] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[0395] Step 3:
[0396] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[0397] Step 4:
[0398] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[0399] Step 5:
[0400] The server begins analyzing the stored image and video data, which includes:
[0401] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[0402] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[0403] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[0404] Step 6:
[0405] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[0406] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0407] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0408] Movement animation: Generates animation based on the extracted movement patterns.
[0409] Step 7:
[0410] The server sends the generated AI model to the user's device. The data sent must be in a format compatible with VR or AR.
[0411] Step 8:
[0412] The user launches a VR headset or AR app and loads the sent AI model, recreating the experience of interacting with a pet in a virtual space.
[0413] Step 9:
[0414] The device operates the AI model in response to user input within the virtual space, allowing interaction with the pet. This recreates the natural presence and vitality of dinosaurs. This scenario is translated into English.
[0415] Example 1
[0416] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0417] Pet loss is a serious psychological issue for many people. The present invention aims to alleviate pet loss by using pet image and video data to create an AI model that faithfully reproduces a pet's appearance and behavior. However, conventional technologies have difficulty accurately capturing a pet's characteristics and lack the means to allow users to fully enjoy interacting with a pet recreated in a virtual space. The purpose of the present invention is to solve these problems.
[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0419] In this invention, the server includes means for uploading image and video data of the pet from the terminal to the server, means for analyzing the uploaded data and generating an AI model that captures the characteristics of the pet, means for reproducing the generated AI model in a virtual space and simulating the life of the pet, and means for the user to display the generated AI model in the virtual space using a VR or AR device and interact with it. This allows the generation of a highly accurate AI model that faithfully reproduces the characteristics of the pet, allowing the user to enjoy interacting with the pet in the virtual space.
[0420] A "server" is a computing device that receives and analyzes data and generates and serves AI models.
[0421] A "terminal" is a computer device that is operated by a user and that uploads and receives data.
[0422] A "user" is an individual who uses a terminal to provide image and video data of their pet to the server and interacts with the generated AI model.
[0423] "Image and video data" refers to still image and video files that record the appearance and behavior of your pet.
[0424] "Upload" is the act of sending data from a user's terminal to a server.
[0425] "Analysis" is the process by which the server processes the received data to extract and identify the pet's characteristics.
[0426] An "AI model" is an artificial intelligence model that is generated based on analyzed data and is used to reproduce the appearance and behavior of a pet.
[0427] "Virtual space" is a virtual three-dimensional space generated by computer simulation.
[0428] "Simulation" is the process of recreating the life and behavior of a real pet in a virtual space.
[0429] "VR equipment" refers to equipment for experiencing virtual reality, including devices such as headsets.
[0430] An "AR device" is a device for experiencing augmented reality, and includes devices such as smartphones and AR glasses.
[0431] "Interaction" refers to two-way communication and manipulation between the user and a pet recreated in a virtual space.
[0432] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses images and video data of the user's pet and analyzes them to generate an AI model that reproduces the pet's appearance and behavior, then recreates it in a virtual space.
[0433] First, users upload image and video data of their pets to the server from their own devices. Devices can be smartphones, tablets, or PCs, and data can be easily selected and sent via a dedicated upload interface. For example, a user can select 10 photos of dogs and three videos from their smartphone gallery and click the "Upload" button.
[0434] The server then receives the uploaded data and analyzes it. It uses a TensorFlow-based image recognition model to recognize the animal's face, body shape, and coat color from the image. For videos, it uses a time-series data analysis tool to analyze movement patterns, and an audio analysis tool to analyze the animal's barks from the audio data. For example, the server uses an image analysis engine to identify the dog's face and body shape, and a video analysis engine to extract movement patterns. It also uses an audio analysis engine to extract bark characteristics.
[0435] Based on this analytical data, the server generates a highly accurate AI model that reproduces the pet's appearance and behavior. The server uses 3D modeling tools to create a three-dimensional model of the pet and applies movement animations to it. It also adds audio data to reproduce its cries. For example, the server creates a 3D model based on the identified face, body shape, and fur color data, implements movement animations based on movement patterns, and adds cries data.
[0436] The generated AI model is then sent back to the device. The user's device must support a VR headset or AR application. The server checks file compatibility, converts the format if necessary, and then sends the AI model to the device. The user can then load the AI model using a VR or AR device, reunite with their pet in a virtual space, and enjoy interacting with it. For example, a user can put on a VR headset, launch a dedicated app, load a 3D model of their pet, and play with it in a virtual garden.
[0437] An example of a prompt is as follows:
[0438] "I've uploaded photos and videos of my dog. Please generate a realistic 3D model based on them."
[0439] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of a pet, providing users with a highly realistic and moving experience.
[0440] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0441] Step 1:
[0442] The user selects and uploads pet image and video data from the device. Specifically, the user opens the gallery or file manager on the device and selects multiple images and videos. For example, the user selects 10 dog photos and 3 videos from the smartphone gallery and clicks the "Upload" button. The image and video data selected by the user are used as input, and this data is sent to the server as output.
[0443] Step 2:
[0444] The server receives the image and video data sent by the user. Specifically, the server checks the integrity of the data and verifies that all files have been received correctly. The input is the image and video data from the user, and the output is the storage and verification of the data.
[0445] Step 3:
[0446] The server analyzes the received image data. It uses a TensorFlow-based image recognition model to perform facial recognition, body shape identification, and coat color analysis. Specifically, the image analysis engine extracts the facial position, body shape outline, and coat color characteristics for each image. For example, it identifies the facial feature points (eyes, nose, and mouth) of a dog and stores that information in a database. Using the received image data as input, facial recognition and body shape identification data are generated as output.
[0447] Step 4:
[0448] The server analyzes the received video data. It uses a time-series data analysis tool to analyze movement patterns and extract the pet's barks from the audio data. To determine specific behaviors, the video analysis engine analyzes each frame and extracts behavioral patterns while detecting the continuity of movement. The audio analysis engine also extracts sound wave characteristics and detects bark patterns and frequencies. For example, a dog's behaviors such as running, eating, and sleeping are recorded along a time axis. The received video data is used as input, and movement patterns and bark data are generated as output.
[0449] Step 5:
[0450] The server generates an AI model based on the analysis data. 3D modeling tools are used to create a three-dimensional model of the pet, and movement animations are applied to that model. Analyzed audio data is then added to recreate the animal's barks. Specific movements are then converted by the server into a 3D model of the face, body shape, and coat color, and the movement patterns obtained from video analysis are incorporated into the animation using motion capture technology. For example, a dog's running motion and barking sounds are realistically recreated. The analysis results are used as input, and a complete AI model is generated as output.
[0451] Step 6:
[0452] The server sends the generated AI model to the user's device. The transmission format is compatible with VR or AR. Specifically, the server checks file compatibility and converts the format if necessary. The data is then seamlessly transferred to the user's device. The generated AI model is used as input, and the AI model is sent to the user's device as output.
[0453] Step 7:
[0454] The user loads the sent AI model using a VR or AR device. Using a dedicated application, the user interacts with their pet in a virtual space. Specifically, the user puts on a VR headset, launches a dedicated application, and loads the AI model. The user can then play with their pet dog in a virtual garden. The AI model sent from the server is used as input, and interaction in a virtual space is realized as output.
[0455] (Application example 1)
[0456] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0457] To alleviate the psychological shock and sadness caused by pet loss, there is a need for a system that can accurately recreate memories of past pets and allow users to re-interact with their pets in a virtual space. Furthermore, there is a need for a method to enable such a system to provide a moving experience without users having to physically visit a store.
[0458] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0459] In this invention, the server includes means for uploading image and video data of the animal from the terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for reproducing the generated AI model in a virtual reality space and simulating the life of the animal, and means for the user to access the virtual space on the terminal and enjoy interacting with the animal's AI model. This allows the user to ease the loss of their pet and have a moving experience through interacting with the pet in the virtual space.
[0460] 1. "Domestic animals" refers to all animals kept as pets in homes.
[0461] 2. "Image and video data" refers to digital data consisting of photographs and videos of pets.
[0462] 3. "Device" refers to electronic devices that can connect to the Internet, such as smartphones, tablets, and PCs.
[0463] 4. "Server" refers to a central processing unit used for data analysis and AI model generation.
[0464] 5. "Upload" refers to the act of sending digital data from a device to a server.
[0465] 6. “Analysis” refers to the act of extracting information based on image or video data.
[0466] 7. "Feature-capturing AI model" refers to an artificial intelligence model that faithfully reproduces the appearance and behavior of a pet.
[0467] 8. "Virtual reality space" refers to a virtual environment created using VR technology.
[0468] 9. "Simulate" refers to the act of recreating real-world actions or environments.
[0469] 10. "Accessing a virtual space" refers to the act of a user connecting to a virtual reality environment using a VR headset or device.
[0470] 11. “Interaction” refers to the act of a user interacting with an AI model.
[0471] This invention is a system aimed at alleviating the loss of a pet. It analyzes image and video data of pets to generate an AI model of the pet, which is then recreated in a virtual space, allowing users to relive memories with their pet and providing a moving experience.
[0472] This system consists of the following main elements:
[0473] 1. Terminal
[0474] Users use devices such as smartphones or tablets to collect images and video data of their pets and upload them to a server. A dedicated application is installed on the device, allowing users to easily select and send data.
[0475] 2. Server
[0476] The server receives the uploaded image and video data and analyzes it, which includes the following processes:
[0477] Image analysis uses image recognition software such as TensorFlow to recognize pet faces, identify body shapes, and analyze fur color.
[0478] Video analysis uses video analysis software such as OpenCV to extract pet movements and behavior patterns.
[0479] Based on the analysis results, an AI model is generated using a 3D modeling tool such as Unity. This AI model is generated by integrating the analysis data from TensorFlow and OpenCV to faithfully reproduce the appearance and behavior of the pet.
[0480] 3. Virtual Reality Space
[0481] The AI model generated by the server is reproduced in a virtual reality space, which can be accessed through a VR headset or an AR app. The generated AI model is configured to operate in real time in the virtual reality space using VR development tools such as Unity.
[0482] By wearing a VR headset, users can access a virtual reality space and enjoy interacting with their pets.
[0483] 4. User Interaction
[0484] Users can interact with the AI model in real time through a VR headset. For example, if a user calls a pet's name in the virtual space, the AI model will respond by turning toward the user or performing a specific action, providing an experience similar to interacting with a real pet.
[0485] A specific example of this system is shown below.
[0486] Specific examples
[0487] Users upload using their smartphones
[0488] Users select 10 photos of their dog and three videos from their smartphone gallery and upload them to the server using a dedicated app.
[0489] Server analysis process
[0490] The server uses TensorFlow to analyze images and OpenCV to analyze videos, which allows for facial recognition, body shape identification, sound analysis, and movement pattern extraction of pets.
[0491] Generating AI models
[0492] Based on the analysis results, the server uses Unity to generate a 3D model and create an AI model.
[0493] Reproduction in virtual reality space
[0494] Users put on a VR headset and access the virtual reality space. They then launch the system's dedicated app, load a 3D model of their pet dog, and play with it in a virtual garden.
[0495] Prompt Sentence Examples
[0496] AI model generation prompts
[0497] Please generate a 3D model that is as faithful as possible to my pet's appearance based on past images and videos of my pet.
[0498] Interaction prompts
[0499] Place the generated 3D model in your virtual garden and simulate its movements as you play with your pet.
[0500] In this way, users can ease the loss of their pet and relive their memories with their pet through an emotional experience in a virtual reality space.
[0501] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0502] Step 1:
[0503] The user uploads pet image and video data using a device. Specifically, the user launches a dedicated application on the device, selects pet photos and videos from the gallery, and clicks the "Upload" button. This sends the selected data to the server. The input is the pet image and video data, and the output is the data transfer to the server.
[0504] Step 2:
[0505] The server receives the uploaded image and video data. The received data is temporarily stored in the server's storage. The input is the image and video data sent by the user, and the output is the data saved in the server's storage.
[0506] Step 3:
[0507] The server uses TensorFlow to analyze the received image data. First, the server preprocesses the image data and performs facial recognition, body shape identification, and fur color analysis. Specifically, the server converts the image data into RGB format and inputs each data into a model to extract features. The input is the saved image data, and the output is feature data related to the pet's facial recognition, body shape, and fur color.
[0508] Step 4:
[0509] The server uses OpenCV to analyze the received video data. The server extracts motion patterns from the video data frame by frame and generates time-series data. Specifically, it extracts frames from the video data and detects motion changes using methods such as optical flow. The input is the saved video data, and the output is time-series data on motion patterns.
[0510] Step 5:
[0511] The server integrates the analyzed image and video data to generate an AI model that captures the pet's characteristics. To do this, a 3D model is built using Unity based on the analysis results of TensorFlow and OpenCV, and movement patterns are added. The input is feature data and time-series data, and the output is the pet's AI model.
[0512] Step 6:
[0513] The server sends the generated AI model to the user's device. Specifically, the generated AI model is sent as data in a format compatible with VR headsets and AR apps. The input is the generated AI model, and the output is data transferred to the user's device (VR headset or smartphone).
[0514] Step 7:
[0515] The user puts on a VR headset and launches a dedicated app to load the generated AI model. Specifically, the user selects the sent AI model from the VR headset app and accesses the virtual space. To interact with the pet in the VR space, the user interacts with the AI model through voice recognition and gesture control. The input is the AI model and the VR headset, and the output is the interaction with the pet in the virtual space.
[0516] Through these steps, the system can ease the loss of a pet and provide a moving experience for users.
[0517] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0518] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[0519] Explanation of system program processing
[0520] 1. Data collection and upload
[0521] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[0522] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0523] 2. Data processing and AI model generation
[0524] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[0525] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[0526] The server uses the analysis results to generate an AI model that replicates the pet's appearance and behavior. This model includes the following elements:
[0527] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0528] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0529] Movement animation: Generates animation based on the extracted movement patterns.
[0530] 3. Sentiment analysis and AI model tuning
[0531] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[0532] Example: The server uses a facial recognition algorithm to detect a user's smiling or sad facial expressions, and a speech analysis tool to analyze changes in the tone of their voice.
[0533] 4. Providing AI models and verifying their operation
[0534] The server generates an AI model and sends the emotion analysis results to the user's device. The data is in a format compatible with VR or AR.
[0535] Example: A server sends data containing the generated 3D model, audio files, and emotion analysis results to a user's VR headset.
[0536] The user launches a VR headset or AR app and loads the AI model, recreating the experience of interacting with a pet in a virtual space. Furthermore, the AI model's behavior changes depending on the user's emotional state.
[0537] Example: A user puts on a VR headset, launches a dedicated app for the system, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[0538] This system not only helps users ease the shock of losing a pet and recreate memories with their pet, but also allows them to enjoy interactions that are in tune with their emotions. The introduction of an emotion engine allows users to have a more realistic and moving experience, and provides psychological support to help them move on to a new life.
[0539] The processing flow will be explained below.
[0540] Step 1:
[0541] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[0542] Step 2:
[0543] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[0544] Step 3:
[0545] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[0546] Step 4:
[0547] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[0548] Step 5:
[0549] The server begins analyzing the stored image and video data, which includes:
[0550] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[0551] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[0552] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[0553] Step 6:
[0554] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[0555] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0556] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0557] Movement animation: Generates animation based on the extracted movement patterns.
[0558] Step 7:
[0559] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[0560] Example: The server uses a facial recognition algorithm to detect the user's facial expressions (e.g., smiling, sad), and a voice analysis tool to analyze changes in voice tone.
[0561] Step 8:
[0562] The server then sends the generated AI model and emotion analysis results to the user's device in a format compatible with VR or AR.
[0563] Step 9:
[0564] The user launches the VR headset or AR app and loads the AI model, which then begins interacting with the pet in the virtual space.
[0565] Example: A user puts on a VR headset, launches a dedicated app, and loads a 3D model of their pet dog. If the user is smiling, the AI model of the pet will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of the pet will quietly cuddle with the user.
[0566] This series of processes not only helps users ease the shock of losing a pet, but also provides an emotionally responsive, interactive experience.
[0567] Example 2
[0568] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0569] Users who are grieving the loss of their pet need a system that allows them to not only reminisce about their memories with their deceased pet, but also enjoy emotionally rich interactions. However, existing systems lack the functionality to empathize with the user's emotions and are insufficient in providing psychological support. Therefore, a system that dynamically analyzes the user's emotions and adjusts the behavior of the pet's AI model to provide more emotionally rich interactions is desired.
[0570] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0571] In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for recreating the generated AI model in a virtual space and simulating the life of the animal, means for activating an emotion engine that detects a user's emotions and analyzing emotions from the user's facial expressions and voice using a camera and microphone, and means for adjusting the operation of the AI model based on the emotion analysis results, thereby enabling users to ease the shock of losing their pet and enjoy emotionally sensitive interactions.
[0572] The "server" is a central processing unit that receives and analyzes data uploaded from the user's device, and generates and adjusts the AI model.
[0573] A "terminal" is a device used by a user, such as a computer or smartphone, through which data is uploaded or downloaded.
[0574] "Upload" means a function or interface for transmitting image or video data from a terminal to a server.
[0575] "Analysis" refers to the process of processing the data received by the server and extracting the animal's characteristics and behavioral patterns.
[0576] An "AI model" is an artificial intelligence model that is generated based on analyzed data and reproduces the appearance and behavior of animals.
[0577] A "virtual space" is an imaginary environment generated using computer graphics and related technologies.
[0578] The "simulation" means virtually recreating the generated AI model in a virtual space and imitating the animal's life.
[0579] An "emotion engine" is a collection of software and algorithms for analyzing a user's emotions.
[0580] A "camera" is a photographing device for capturing the user's facial expression.
[0581] A "microphone" is a recording device for capturing the user's voice.
[0582] "Emotion analysis" is the process of determining a user's current emotional state based on data captured using a camera or microphone.
[0583] "Behavior adjustment" is the process of changing the behavior and reactions of an AI model based on the results of emotion analysis.
[0584] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[0585] The system of the present invention mainly uses the following hardware and software:
[0586] Hardware: Servers, devices (smartphones, PCs, VR headsets), cameras, microphones
[0587] Software: TensorFlow-based image recognition models, voice analysis tools, time-series data analysis tools, 3D modeling software, voice synthesis technology, emotion recognition algorithms
[0588] Specific embodiments will be described below.
[0589] Data collection and upload
[0590] Users select images and videos of their pets from their devices and send them to the server using the system's upload interface. For example, a user selects 10 dog photos and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0591] Data processing and AI model generation
[0592] The server receives the uploaded images and videos and stores them in the initial data storage. It then uses a TensorFlow-based image recognition model to analyze the images and perform facial recognition, body shape identification, and fur color analysis of the pet. It also analyzes the video to extract the pet's movements and behavior patterns. Based on these analysis results, the server generates an AI model with the following elements:
[0593] 1. 3D model: A 3D model is created using the analyzed face, body shape, and fur color data.
[0594] 2. Audio file: The extracted call data is reproduced using voice synthesis technology.
[0595] 3. Motion animation: Generate animation based on the extracted motion patterns.
[0596] Sentiment analysis and AI model tuning
[0597] The server activates an emotion engine to monitor the user's emotions, using a camera and microphone to capture the user's facial expressions and voice in real time. For example, the camera captures the user's face and uses a facial recognition algorithm to detect smiling or sad expressions. The audio recorded by the microphone is then analyzed for changes in tone using a voice analysis tool. As a result, the server adjusts the behavior of the AI model according to the user's emotional state. For example, if the user is smiling, the AI pet will move around energetically. Conversely, if the user is sad, the AI pet will quietly cuddle.
[0598] Providing AI models and verifying their operation
[0599] The server sends the generated AI model and emotion analysis results to the user's device (e.g., a VR headset). The user then launches the VR headset or AR app and loads the sent AI model, recreating interaction with a pet in a virtual space. The AI pet's behavior changes depending on the user's emotional state. For example, a user puts on a VR headset, launches a system-specific app, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[0600] Prompt Sentence Examples
[0601] "I want to be reunited with my pet. Please use these photos and videos to create an AI model."
[0602] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0603] Step 1:
[0604] User selects data and uploads it on the device
[0605] The user selects images and video data of their pet from their device (smartphone or PC).
[0606] What happens: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0607] Input: Image and video files stored in your phone's gallery.
[0608] Output: Image and video data sent to the server.
[0609] Step 2:
[0610] The server receives and stores the data
[0611] The server stores the image and video data received from the terminal in an initial data storage.
[0612] Specific operation: The server saves the uploaded 10 images and 3 videos to the file system or database.
[0613] Input: Image and video data submitted by the user.
[0614] Output: Image and video files saved to storage.
[0615] Step 3:
[0616] The server analyzes the data
[0617] The server analyzes the stored data using a TensorFlow-based image recognition model.
[0618] How it works: The server uses image recognition models to identify pet faces, body shapes, and coat colors, and uses time-series data analysis tools to extract movement and behavior patterns from video.
[0619] Input: Stored image and video data.
[0620] Output: Face recognition results, body shape identification results, coat color data, movement pattern data.
[0621] Step 4:
[0622] The server generates the AI model
[0623] Based on the analysis results, the server generates an AI model that includes a 3D model, audio files, and motion animations.
[0624] Specific movements: The server uses 3D modeling software to create a 3D model based on the pet's face, body shape, and fur color data. It then uses voice synthesis technology to reproduce the pet's barking sound and generates movement animations based on the movement pattern data.
[0625] Input: Face recognition results, body shape identification results, coat color data, movement pattern data.
[0626] Output: 3D model, reconstructed audio files, movement animations.
[0627] Step 5:
[0628] The server starts the emotion engine
[0629] The server detects the user's emotions by capturing the user's facial expressions and voice in real time using the device's camera and microphone.
[0630] How it works: The server takes a picture of the user's face with the device's camera, uses a facial recognition algorithm to detect smiling or sad expressions, and uses a voice analysis tool to analyze changes in the tone of the voice recorded by the microphone.
[0631] Input: Video and audio recording of the user's face.
[0632] Output: User sentiment analysis results.
[0633] Step 6:
[0634] The server adjusts the AI model
[0635] The server adjusts the behavior of the generated AI model based on the results of emotion analysis.
[0636] Specific behavior: Based on the results of emotion analysis, the AI model's animations and reactions are changed in real time. For example, if the user is smiling, the AI pet will move around energetically and wag its tail. If the user is sad, the AI pet will quietly cuddle.
[0637] Input: User sentiment analysis results, generated AI model.
[0638] Output: The movement animations and reactions of the tuned AI model.
[0639] Step 7:
[0640] The server sends the AI model to the device.
[0641] The server sends the generated AI model and emotion analysis results to the user's device.
[0642] Specific operation: The server sends data including the generated 3D model, audio files, movement animations, and emotion analysis results to the user's VR headset.
[0643] Input: Tuned AI model, sentiment analysis results.
[0644] Output: Data sent to the user's device.
[0645] Step 8:
[0646] The user launches the AI model and checks its operation
[0647] The user launches a VR headset or AR app and displays the sent AI model in a virtual space.
[0648] Specific operation: The user puts on the VR headset, launches the system's dedicated app, loads the 3D model of the pet, enjoys interacting with the AI pet model, and sees its behavior change in real time.
[0649] Input: User's emotional state, transmitted data.
[0650] Output: Interaction with a pet in a virtual space.
[0651] This system allows users to relive memories with their pets while enjoying emotionally rich interactions.
[0652] (Application example 2)
[0653] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0654] There is a need for a system that can vividly recreate memories of pets, empathize with the user's emotional state, and provide real-time interaction while easing the psychological shock of pet loss. However, conventional systems lack the technology to flexibly respond to changes in the user's emotions, making it difficult to achieve more emotionally rich interactions.
[0655] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for the server to use an emotion engine that analyzes the user's emotional state and dynamically adjust the behavior of the generated AI model, and means for reproducing the generated AI model in a virtual space, simulating the life of the animal, and interacting with the animal in accordance with the user's emotional state. This makes it possible to provide more emotionally rich interactions in real time that are in line with the user's emotional state while easing the loss of a pet.
[0656] "Image" is data containing static visual information of an animal.
[0657] "Video data" refers to data that includes dynamic visual and audio information such as animal movements and sounds.
[0658] "Terminal" refers to a device used by a user, such as a smartphone, tablet, or PC.
[0659] A "server" is a computer system that receives, analyzes, stores, and transmits data over a network.
[0660] "Analysis" is the process of extracting necessary information from image and video data and identifying its features.
[0661] An "AI model" is an artificial intelligence software model created to reproduce an animal's shape, movement patterns, sounds, etc.
[0662] An "emotion engine" is a technology for analyzing a user's emotional state from their facial expressions and voice.
[0663] "Dynamic adjustment" refers to changing the behavior of an AI model in response to the user's emotional state in real time.
[0664] A "virtual space" is a virtual environment created using computer technology.
[0665] "Simulation" means virtually reproducing the life and behavior of real animals.
[0666] "Interaction" refers to the interaction between a user and an AI model.
[0667] "Real-time" refers to processing occurring immediately, without delay.
[0668] The embodiment of this invention is a system that allows a user to upload images and video data of animals kept by the user from a terminal to a server, and the server analyzes this data and generates an AI model that captures the characteristics of the animal.Furthermore, the system uses an emotion engine to analyze the user's emotional state and dynamically adjusts the behavior of the generated AI model according to the user's emotions.
[0669] Specifically, users upload images and video data of their pet animals to a server from their smartphones, tablets, or PCs. The uploaded data is then analyzed on the server using image analysis models such as TensorFlow to recognize the animal's face, body shape, and movement patterns. If audio data is included, audio analysis tools are also used to analyze the animal's cries.
[0670] The analyzed data is then used to create an AI model that recreates the animal's appearance, movement patterns, and sounds, including a 3D model, movement animations, and audio files.
[0671] To analyze the user's emotional state using the emotion engine, the server uses camera footage and audio data sent from the user's device. It uses facial recognition algorithms to read the user's facial expressions and audio analysis tools to analyze the tone and content of the voice, dynamically detecting whether the user is happy, sad, or expressing other emotions.
[0672] The generated AI model is then recreated in a virtual space, and its behavior is dynamically adjusted according to the user's emotional state. For example, if the user is smiling, the AI model will move around energetically and wag its tail. Conversely, if the user is sad, the AI model will quietly take on a comforting behavior.
[0673] This system allows users to vividly relive their memories of their pet while easing the loss of their pet. It can provide real-time interactions that are tailored to the user's emotional state, resulting in a more emotionally enriching experience.
[0674] Specific examples
[0675] A user launches the virtual store app, takes a video of their pet dog using their smartphone camera, and uploads it to the server. The server then uses TensorFlow models to analyze the images and sounds and generate an AI model. This AI model operates in the virtual space based on the user's emotional state.
[0676] Prompt Sentence Examples
[0677] The system will be provided with an image of a dog and a short video of its barking. It will then analyze the specific behaviors and sounds the dog exhibits and create a 3D model of it. Furthermore, the system will have the dog's movements change in the virtual space depending on the user's emotions.
[0678] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0679] Step 1:
[0680] A user uses a device (smartphone, tablet, PC, etc.) to take or select images and video data of their pet dog and upload them to a server through an application. The input is the image and video data of the pet dog, and the output is the image and video files transferred to the server. This process involves the user selecting the data through the app interface and clicking the upload button.
[0681] Step 2:
[0682] The server receives the uploaded image and video data and begins analyzing it using an image analysis model such as TensorFlow. The input is the uploaded image and video data, and the output is the analyzed animal's features (face recognition, body shape identification, movement patterns, etc.). This process involves the server running an image recognition algorithm to extract the animal's face and body shape from the image and video.
[0683] Step 3:
[0684] The server generates an AI model based on the analysis results. The input is the analyzed animal's feature data, and the output is the generated AI model (3D model, movement animation, audio file, etc.). This process involves the server using a 3D model generation tool to build a 3D model that reproduces the animal's features and creating animations that include movement patterns and sounds.
[0685] Step 4:
[0686] The server uses an emotion engine to analyze the camera footage and audio data sent from the user's device and analyze the user's emotional state in real time. The input is the user's camera footage and audio data, and the output is the user's emotional state (happiness, sadness, etc.). This process involves the emotion analysis algorithm analyzing facial expressions and tone of voice to identify emotions such as joy, anger, sadness, and happiness.
[0687] Step 5:
[0688] The server dynamically adjusts the behavior of the generated AI model according to the user's emotional state. The input is the user's emotional state and the generated AI model, and the output is the adjusted AI model (a behavior pattern according to the emotional state). This process involves the server changing the behavior of the AI model based on the emotion analysis results and executing the corresponding behavior (for example, if the user is smiling, the AI model behaves cheerfully).
[0689] Step 6:
[0690] The server sends the generated AI model and adjusted behavior patterns to the user's device. The input is the adjusted AI model data, and the output is the AI model displayed on the user's device. This process includes the server encoding the data and sending it in a format suitable for the user's VR headset or AR app.
[0691] Step 7:
[0692] The user uses a device (such as a VR headset, smart glasses, or smartphone) to recreate the transmitted AI model in a virtual space and interact with the AI model according to the user's emotional state. The input is the adjusted AI model data and the user's emotional state, and the output is the user's interaction with an emotive virtual pet. This process involves the user operating the device and interacting with the AI model in the virtual space.
[0693] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0694] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0695] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0696] [Third embodiment]
[0697] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0698] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0699] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0700] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0701] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0702] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0703] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0704] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0705] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0706] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0707] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0708] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0709] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses image and video data of pets and analyzes them to generate an AI model that reproduces the appearance and behavior of the pet, and then recreates it in a virtual space.
[0710] Explanation of system program processing
[0711] 1. Data collection and upload
[0712] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[0713] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0714] 2. Data processing and AI model generation
[0715] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[0716] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[0717] The server generates an AI model that replicates the pet's appearance and behavior based on the analysis results, and the generated AI model is tuned to replicate the pet's characteristics as faithfully as possible.
[0718] Example: The server creates a 3D model based on identified face, body shape, and fur color data, reproduces animal sounds using audio data, and generates animations based on movement patterns.
[0719] 3. Providing AI models and verifying their operation
[0720] The server sends the generated AI model to the user's device, which must support a VR headset or AR app.
[0721] Example: A server sends the generated AI model, including the 3D model and audio files, in a format compatible with the user's VR headset.
[0722] Users can then load the AI model using VR or AR devices, reunite with their pet in a virtual space, and enjoy interacting with it.
[0723] Example: A user puts on a VR headset, launches a system-specific app, loads a 3D model of their pet dog, and plays with it in a virtual garden.
[0724] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of their pet, providing users with a highly realistic and moving experience.
[0725] The processing flow will be explained below.
[0726] Step 1:
[0727] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[0728] Step 2:
[0729] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[0730] Step 3:
[0731] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[0732] Step 4:
[0733] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[0734] Step 5:
[0735] The server begins analyzing the stored image and video data, which includes:
[0736] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[0737] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[0738] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[0739] Step 6:
[0740] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[0741] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0742] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0743] Movement animation: Generates animation based on the extracted movement patterns.
[0744] Step 7:
[0745] The server sends the generated AI model to the user's device. The data sent must be in a format compatible with VR or AR.
[0746] Step 8:
[0747] The user launches a VR headset or AR app and loads the sent AI model, recreating the experience of interacting with a pet in a virtual space.
[0748] Step 9:
[0749] The device operates the AI model in response to user input within the virtual space, allowing interaction with the pet. This recreates the natural presence and vitality of dinosaurs. This scenario is translated into English.
[0750] Example 1
[0751] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0752] Pet loss is a serious psychological issue for many people. The present invention aims to alleviate pet loss by using pet image and video data to create an AI model that faithfully reproduces a pet's appearance and behavior. However, conventional technologies have difficulty accurately capturing a pet's characteristics and lack the means to allow users to fully enjoy interacting with a pet recreated in a virtual space. The purpose of the present invention is to solve these problems.
[0753] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0754] In this invention, the server includes means for uploading image and video data of the pet from the terminal to the server, means for analyzing the uploaded data and generating an AI model that captures the characteristics of the pet, means for reproducing the generated AI model in a virtual space and simulating the life of the pet, and means for the user to display the generated AI model in the virtual space using a VR or AR device and interact with it. This allows the generation of a highly accurate AI model that faithfully reproduces the characteristics of the pet, allowing the user to enjoy interacting with the pet in the virtual space.
[0755] A "server" is a computing device that receives and analyzes data and generates and serves AI models.
[0756] A "terminal" is a computer device that is operated by a user and that uploads and receives data.
[0757] A "user" is an individual who uses a terminal to provide image and video data of their pet to the server and interacts with the generated AI model.
[0758] "Image and video data" refers to still image and video files that record the appearance and behavior of your pet.
[0759] "Upload" is the act of sending data from a user's terminal to a server.
[0760] "Analysis" is the process by which the server processes the received data to extract and identify the pet's characteristics.
[0761] An "AI model" is an artificial intelligence model that is generated based on analyzed data and is used to reproduce the appearance and behavior of a pet.
[0762] "Virtual space" is a virtual three-dimensional space generated by computer simulation.
[0763] "Simulation" is the process of recreating the life and behavior of a real pet in a virtual space.
[0764] "VR equipment" refers to equipment for experiencing virtual reality, including devices such as headsets.
[0765] An "AR device" is a device for experiencing augmented reality, and includes devices such as smartphones and AR glasses.
[0766] "Interaction" refers to two-way communication and manipulation between the user and a pet recreated in a virtual space.
[0767] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses images and video data of the user's pet and analyzes them to generate an AI model that reproduces the pet's appearance and behavior, then recreates it in a virtual space.
[0768] First, users upload image and video data of their pets to the server from their own devices. Devices can be smartphones, tablets, or PCs, and data can be easily selected and sent via a dedicated upload interface. For example, a user can select 10 photos of dogs and three videos from their smartphone gallery and click the "Upload" button.
[0769] The server then receives the uploaded data and analyzes it. It uses a TensorFlow-based image recognition model to recognize the animal's face, body shape, and coat color from the image. For videos, it uses a time-series data analysis tool to analyze movement patterns, and an audio analysis tool to analyze the animal's barks from the audio data. For example, the server uses an image analysis engine to identify the dog's face and body shape, and a video analysis engine to extract movement patterns. It also uses an audio analysis engine to extract bark characteristics.
[0770] Based on this analytical data, the server generates a highly accurate AI model that reproduces the pet's appearance and behavior. The server uses 3D modeling tools to create a three-dimensional model of the pet and applies movement animations to it. It also adds audio data to reproduce its cries. For example, the server creates a 3D model based on the identified face, body shape, and fur color data, implements movement animations based on movement patterns, and adds cries data.
[0771] The generated AI model is then sent back to the device. The user's device must support a VR headset or AR application. The server checks file compatibility, converts the format if necessary, and then sends the AI model to the device. The user can then load the AI model using a VR or AR device, reunite with their pet in a virtual space, and enjoy interacting with it. For example, a user can put on a VR headset, launch a dedicated app, load a 3D model of their pet, and play with it in a virtual garden.
[0772] An example of a prompt is as follows:
[0773] "I've uploaded photos and videos of my dog. Please generate a realistic 3D model based on them."
[0774] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of a pet, providing users with a highly realistic and moving experience.
[0775] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0776] Step 1:
[0777] The user selects and uploads pet image and video data from the device. Specifically, the user opens the gallery or file manager on the device and selects multiple images and videos. For example, the user selects 10 dog photos and 3 videos from the smartphone gallery and clicks the "Upload" button. The image and video data selected by the user are used as input, and this data is sent to the server as output.
[0778] Step 2:
[0779] The server receives the image and video data sent by the user. Specifically, the server checks the integrity of the data and verifies that all files have been received correctly. The input is the image and video data from the user, and the output is the storage and verification of the data.
[0780] Step 3:
[0781] The server analyzes the received image data. It uses a TensorFlow-based image recognition model to perform facial recognition, body shape identification, and coat color analysis. Specifically, the image analysis engine extracts the facial position, body shape outline, and coat color characteristics for each image. For example, it identifies the facial feature points (eyes, nose, and mouth) of a dog and stores that information in a database. Using the received image data as input, facial recognition and body shape identification data are generated as output.
[0782] Step 4:
[0783] The server analyzes the received video data. It uses a time-series data analysis tool to analyze movement patterns and extract the pet's barks from the audio data. To determine specific behaviors, the video analysis engine analyzes each frame and extracts behavioral patterns while detecting the continuity of movement. The audio analysis engine also extracts sound wave characteristics and detects bark patterns and frequencies. For example, a dog's behaviors such as running, eating, and sleeping are recorded along a time axis. The received video data is used as input, and movement patterns and bark data are generated as output.
[0784] Step 5:
[0785] The server generates an AI model based on the analysis data. 3D modeling tools are used to create a three-dimensional model of the pet, and movement animations are applied to that model. Analyzed audio data is then added to recreate the animal's barks. Specific movements are then converted by the server into a 3D model of the face, body shape, and coat color, and the movement patterns obtained from video analysis are incorporated into the animation using motion capture technology. For example, a dog's running motion and barking sounds are realistically recreated. The analysis results are used as input, and a complete AI model is generated as output.
[0786] Step 6:
[0787] The server sends the generated AI model to the user's device. The transmission format is compatible with VR or AR. Specifically, the server checks file compatibility and converts the format if necessary. The data is then seamlessly transferred to the user's device. The generated AI model is used as input, and the AI model is sent to the user's device as output.
[0788] Step 7:
[0789] The user loads the sent AI model using a VR or AR device. Using a dedicated application, the user interacts with their pet in a virtual space. Specifically, the user puts on a VR headset, launches a dedicated application, and loads the AI model. The user can then play with their pet dog in a virtual garden. The AI model sent from the server is used as input, and interaction in a virtual space is realized as output.
[0790] (Application example 1)
[0791] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0792] To alleviate the psychological shock and sadness caused by pet loss, there is a need for a system that can accurately recreate memories of past pets and allow users to re-interact with their pets in a virtual space. Furthermore, there is a need for a method to enable such a system to provide a moving experience without users having to physically visit a store.
[0793] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0794] In this invention, the server includes means for uploading image and video data of the animal from the terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for reproducing the generated AI model in a virtual reality space and simulating the life of the animal, and means for the user to access the virtual space on the terminal and enjoy interacting with the animal's AI model. This allows the user to ease the loss of their pet and have a moving experience through interacting with the pet in the virtual space.
[0795] 1. "Domestic animals" refers to all animals kept as pets in homes.
[0796] 2. "Image and video data" refers to digital data consisting of photographs and videos of pets.
[0797] 3. "Device" refers to electronic devices that can connect to the Internet, such as smartphones, tablets, and PCs.
[0798] 4. "Server" refers to a central processing unit used for data analysis and AI model generation.
[0799] 5. "Upload" refers to the act of sending digital data from a device to a server.
[0800] 6. “Analysis” refers to the act of extracting information based on image or video data.
[0801] 7. "Feature-capturing AI model" refers to an artificial intelligence model that faithfully reproduces the appearance and behavior of a pet.
[0802] 8. "Virtual reality space" refers to a virtual environment created using VR technology.
[0803] 9. "Simulate" refers to the act of recreating real-world actions or environments.
[0804] 10. "Accessing a virtual space" refers to the act of a user connecting to a virtual reality environment using a VR headset or device.
[0805] 11. “Interaction” refers to the act of a user interacting with an AI model.
[0806] This invention is a system aimed at alleviating the loss of a pet. It analyzes image and video data of pets to generate an AI model of the pet, which is then recreated in a virtual space, allowing users to relive memories with their pet and providing a moving experience.
[0807] This system consists of the following main elements:
[0808] 1. Terminal
[0809] Users use devices such as smartphones or tablets to collect images and video data of their pets and upload them to a server. A dedicated application is installed on the device, allowing users to easily select and send data.
[0810] 2. Server
[0811] The server receives the uploaded image and video data and analyzes it, which includes the following processes:
[0812] Image analysis uses image recognition software such as TensorFlow to recognize pet faces, identify body shapes, and analyze fur color.
[0813] Video analysis uses video analysis software such as OpenCV to extract pet movements and behavior patterns.
[0814] Based on the analysis results, an AI model is generated using a 3D modeling tool such as Unity. This AI model is generated by integrating the analysis data from TensorFlow and OpenCV to faithfully reproduce the appearance and behavior of the pet.
[0815] 3. Virtual Reality Space
[0816] The AI model generated by the server is reproduced in a virtual reality space, which can be accessed through a VR headset or an AR app. The generated AI model is configured to operate in real time in the virtual reality space using VR development tools such as Unity.
[0817] By wearing a VR headset, users can access a virtual reality space and enjoy interacting with their pets.
[0818] 4. User Interaction
[0819] Users can interact with the AI model in real time through a VR headset. For example, if a user calls a pet's name in the virtual space, the AI model will respond by turning toward the user or performing a specific action, providing an experience similar to interacting with a real pet.
[0820] A specific example of this system is shown below.
[0821] Specific examples
[0822] Users upload using their smartphones
[0823] Users select 10 photos of their dog and three videos from their smartphone gallery and upload them to the server using a dedicated app.
[0824] Server analysis process
[0825] The server uses TensorFlow to analyze images and OpenCV to analyze videos, which allows for facial recognition, body shape identification, sound analysis, and movement pattern extraction of pets.
[0826] Generating AI models
[0827] Based on the analysis results, the server uses Unity to generate a 3D model and create an AI model.
[0828] Reproduction in virtual reality space
[0829] Users put on a VR headset and access the virtual reality space. They then launch the system's dedicated app, load a 3D model of their pet dog, and play with it in a virtual garden.
[0830] Prompt Sentence Examples
[0831] AI model generation prompts
[0832] Please generate a 3D model that is as faithful as possible to my pet's appearance based on past images and videos of my pet.
[0833] Interaction prompts
[0834] Place the generated 3D model in your virtual garden and simulate its movements as you play with your pet.
[0835] In this way, users can ease the loss of their pet and relive their memories with their pet through an emotional experience in a virtual reality space.
[0836] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0837] Step 1:
[0838] The user uploads pet image and video data using a device. Specifically, the user launches a dedicated application on the device, selects pet photos and videos from the gallery, and clicks the "Upload" button. This sends the selected data to the server. The input is the pet image and video data, and the output is the data transfer to the server.
[0839] Step 2:
[0840] The server receives the uploaded image and video data. The received data is temporarily stored in the server's storage. The input is the image and video data sent by the user, and the output is the data saved in the server's storage.
[0841] Step 3:
[0842] The server uses TensorFlow to analyze the received image data. First, the server preprocesses the image data and performs facial recognition, body shape identification, and fur color analysis. Specifically, the server converts the image data into RGB format and inputs each data into a model to extract features. The input is the saved image data, and the output is feature data related to the pet's facial recognition, body shape, and fur color.
[0843] Step 4:
[0844] The server uses OpenCV to analyze the received video data. The server extracts motion patterns from the video data frame by frame and generates time-series data. Specifically, it extracts frames from the video data and detects motion changes using methods such as optical flow. The input is the saved video data, and the output is time-series data on motion patterns.
[0845] Step 5:
[0846] The server integrates the analyzed image and video data to generate an AI model that captures the pet's characteristics. To do this, a 3D model is built using Unity based on the analysis results of TensorFlow and OpenCV, and movement patterns are added. The input is feature data and time-series data, and the output is the pet's AI model.
[0847] Step 6:
[0848] The server sends the generated AI model to the user's device. Specifically, the generated AI model is sent as data in a format compatible with VR headsets and AR apps. The input is the generated AI model, and the output is data transferred to the user's device (VR headset or smartphone).
[0849] Step 7:
[0850] The user puts on a VR headset and launches a dedicated app to load the generated AI model. Specifically, the user selects the sent AI model from the VR headset app and accesses the virtual space. To interact with the pet in the VR space, the user interacts with the AI model through voice recognition and gesture control. The input is the AI model and the VR headset, and the output is the interaction with the pet in the virtual space.
[0851] Through these steps, the system can ease the loss of a pet and provide a moving experience for users.
[0852] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0853] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[0854] Explanation of system program processing
[0855] 1. Data collection and upload
[0856] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[0857] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0858] 2. Data processing and AI model generation
[0859] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[0860] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[0861] The server uses the analysis results to generate an AI model that replicates the pet's appearance and behavior. This model includes the following elements:
[0862] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0863] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0864] Movement animation: Generates animation based on the extracted movement patterns.
[0865] 3. Sentiment analysis and AI model tuning
[0866] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[0867] Example: The server uses a facial recognition algorithm to detect a user's smiling or sad facial expressions, and a speech analysis tool to analyze changes in the tone of their voice.
[0868] 4. Providing AI models and verifying their operation
[0869] The server generates an AI model and sends the emotion analysis results to the user's device. The data is in a format compatible with VR or AR.
[0870] Example: A server sends data containing the generated 3D model, audio files, and emotion analysis results to a user's VR headset.
[0871] The user launches a VR headset or AR app and loads the AI model, recreating the experience of interacting with a pet in a virtual space. Furthermore, the AI model's behavior changes depending on the user's emotional state.
[0872] Example: A user puts on a VR headset, launches a dedicated app for the system, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[0873] This system not only helps users ease the shock of losing a pet and recreate memories with their pet, but also allows them to enjoy interactions that are in tune with their emotions. The introduction of an emotion engine allows users to have a more realistic and moving experience, and provides psychological support to help them move on to a new life.
[0874] The processing flow will be explained below.
[0875] Step 1:
[0876] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[0877] Step 2:
[0878] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[0879] Step 3:
[0880] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[0881] Step 4:
[0882] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[0883] Step 5:
[0884] The server begins analyzing the stored image and video data, which includes:
[0885] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[0886] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[0887] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[0888] Step 6:
[0889] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[0890] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[0891] Audio file: The extracted call data is reproduced using voice synthesis technology.
[0892] Movement animation: Generates animation based on the extracted movement patterns.
[0893] Step 7:
[0894] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[0895] Example: The server uses a facial recognition algorithm to detect the user's facial expressions (e.g., smiling, sad), and a voice analysis tool to analyze changes in voice tone.
[0896] Step 8:
[0897] The server then sends the generated AI model and emotion analysis results to the user's device in a format compatible with VR or AR.
[0898] Step 9:
[0899] The user launches the VR headset or AR app and loads the AI model, which then begins interacting with the pet in the virtual space.
[0900] Example: A user puts on a VR headset, launches a dedicated app, and loads a 3D model of their pet dog. If the user is smiling, the AI model of the pet will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of the pet will quietly cuddle with the user.
[0901] This series of processes not only helps users ease the shock of losing a pet, but also provides an emotionally responsive, interactive experience.
[0902] Example 2
[0903] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0904] Users who are grieving the loss of their pet need a system that allows them to not only reminisce about their memories with their deceased pet, but also enjoy emotionally rich interactions. However, existing systems lack the functionality to empathize with the user's emotions and are insufficient in providing psychological support. Therefore, a system that dynamically analyzes the user's emotions and adjusts the behavior of the pet's AI model to provide more emotionally rich interactions is desired.
[0905] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0906] In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for recreating the generated AI model in a virtual space and simulating the life of the animal, means for activating an emotion engine that detects a user's emotions and analyzing emotions from the user's facial expressions and voice using a camera and microphone, and means for adjusting the operation of the AI model based on the emotion analysis results, thereby enabling users to ease the shock of losing their pet and enjoy emotionally sensitive interactions.
[0907] The "server" is a central processing unit that receives and analyzes data uploaded from the user's device, and generates and adjusts the AI model.
[0908] A "terminal" is a device used by a user, such as a computer or smartphone, through which data is uploaded or downloaded.
[0909] "Upload" means a function or interface for transmitting image or video data from a terminal to a server.
[0910] "Analysis" refers to the process of processing the data received by the server and extracting the animal's characteristics and behavioral patterns.
[0911] An "AI model" is an artificial intelligence model that is generated based on analyzed data and reproduces the appearance and behavior of animals.
[0912] A "virtual space" is an imaginary environment generated using computer graphics and related technologies.
[0913] The "simulation" means virtually recreating the generated AI model in a virtual space and imitating the animal's life.
[0914] An "emotion engine" is a collection of software and algorithms for analyzing a user's emotions.
[0915] A "camera" is a photographing device for capturing the user's facial expression.
[0916] A "microphone" is a recording device for capturing the user's voice.
[0917] "Emotion analysis" is the process of determining a user's current emotional state based on data captured using a camera or microphone.
[0918] "Behavior adjustment" is the process of changing the behavior and reactions of an AI model based on the results of emotion analysis.
[0919] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[0920] The system of the present invention mainly uses the following hardware and software:
[0921] Hardware: Servers, devices (smartphones, PCs, VR headsets), cameras, microphones
[0922] Software: TensorFlow-based image recognition models, voice analysis tools, time-series data analysis tools, 3D modeling software, voice synthesis technology, emotion recognition algorithms
[0923] Specific embodiments will be described below.
[0924] Data collection and upload
[0925] Users select images and videos of their pets from their devices and send them to the server using the system's upload interface. For example, a user selects 10 dog photos and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0926] Data processing and AI model generation
[0927] The server receives the uploaded images and videos and stores them in the initial data storage. It then uses a TensorFlow-based image recognition model to analyze the images and perform facial recognition, body shape identification, and fur color analysis of the pet. It also analyzes the video to extract the pet's movements and behavior patterns. Based on these analysis results, the server generates an AI model with the following elements:
[0928] 1. 3D model: A 3D model is created using the analyzed face, body shape, and fur color data.
[0929] 2. Audio file: The extracted call data is reproduced using voice synthesis technology.
[0930] 3. Motion animation: Generate animation based on the extracted motion patterns.
[0931] Sentiment analysis and AI model tuning
[0932] The server activates an emotion engine to monitor the user's emotions, using a camera and microphone to capture the user's facial expressions and voice in real time. For example, the camera captures the user's face and uses a facial recognition algorithm to detect smiling or sad expressions. The audio recorded by the microphone is then analyzed for changes in tone using a voice analysis tool. As a result, the server adjusts the behavior of the AI model according to the user's emotional state. For example, if the user is smiling, the AI pet will move around energetically. Conversely, if the user is sad, the AI pet will quietly cuddle.
[0933] Providing AI models and verifying their operation
[0934] The server sends the generated AI model and emotion analysis results to the user's device (e.g., a VR headset). The user then launches the VR headset or AR app and loads the sent AI model, recreating interaction with a pet in a virtual space. The AI pet's behavior changes depending on the user's emotional state. For example, a user puts on a VR headset, launches a system-specific app, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[0935] Prompt Sentence Examples
[0936] "I want to be reunited with my pet. Please use these photos and videos to create an AI model."
[0937] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0938] Step 1:
[0939] User selects data and uploads it on the device
[0940] The user selects images and video data of their pet from their device (smartphone or PC).
[0941] What happens: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[0942] Input: Image and video files stored in your phone's gallery.
[0943] Output: Image and video data sent to the server.
[0944] Step 2:
[0945] The server receives and stores the data
[0946] The server stores the image and video data received from the terminal in an initial data storage.
[0947] Specific operation: The server saves the uploaded 10 images and 3 videos to the file system or database.
[0948] Input: Image and video data submitted by the user.
[0949] Output: Image and video files saved to storage.
[0950] Step 3:
[0951] The server analyzes the data
[0952] The server analyzes the stored data using a TensorFlow-based image recognition model.
[0953] How it works: The server uses image recognition models to identify pet faces, body shapes, and coat colors, and uses time-series data analysis tools to extract movement and behavior patterns from video.
[0954] Input: Stored image and video data.
[0955] Output: Face recognition results, body shape identification results, coat color data, movement pattern data.
[0956] Step 4:
[0957] The server generates the AI model
[0958] Based on the analysis results, the server generates an AI model that includes a 3D model, audio files, and motion animations.
[0959] Specific movements: The server uses 3D modeling software to create a 3D model based on the pet's face, body shape, and fur color data. It then uses voice synthesis technology to reproduce the pet's barking sound and generates movement animations based on the movement pattern data.
[0960] Input: Face recognition results, body shape identification results, coat color data, movement pattern data.
[0961] Output: 3D model, reconstructed audio files, movement animations.
[0962] Step 5:
[0963] The server starts the emotion engine
[0964] The server detects the user's emotions by capturing the user's facial expressions and voice in real time using the device's camera and microphone.
[0965] How it works: The server takes a picture of the user's face with the device's camera, uses a facial recognition algorithm to detect smiling or sad expressions, and uses a voice analysis tool to analyze changes in the tone of the voice recorded by the microphone.
[0966] Input: Video and audio recording of the user's face.
[0967] Output: User sentiment analysis results.
[0968] Step 6:
[0969] The server adjusts the AI model
[0970] The server adjusts the behavior of the generated AI model based on the results of emotion analysis.
[0971] Specific behavior: Based on the results of emotion analysis, the AI model's animations and reactions are changed in real time. For example, if the user is smiling, the AI pet will move around energetically and wag its tail. If the user is sad, the AI pet will quietly cuddle.
[0972] Input: User sentiment analysis results, generated AI model.
[0973] Output: The movement animations and reactions of the tuned AI model.
[0974] Step 7:
[0975] The server sends the AI model to the device.
[0976] The server sends the generated AI model and emotion analysis results to the user's device.
[0977] Specific operation: The server sends data including the generated 3D model, audio files, movement animations, and emotion analysis results to the user's VR headset.
[0978] Input: Tuned AI model, sentiment analysis results.
[0979] Output: Data sent to the user's device.
[0980] Step 8:
[0981] The user launches the AI model and checks its operation
[0982] The user launches a VR headset or AR app and displays the sent AI model in a virtual space.
[0983] Specific operation: The user puts on the VR headset, launches the system's dedicated app, loads the 3D model of the pet, enjoys interacting with the AI pet model, and sees its behavior change in real time.
[0984] Input: User's emotional state, transmitted data.
[0985] Output: Interaction with a pet in a virtual space.
[0986] This system allows users to relive memories with their pets while enjoying emotionally rich interactions.
[0987] (Application example 2)
[0988] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0989] There is a need for a system that can vividly recreate memories of pets, empathize with the user's emotional state, and provide real-time interaction while easing the psychological shock of pet loss. However, conventional systems lack the technology to flexibly respond to changes in the user's emotions, making it difficult to achieve more emotionally rich interactions.
[0990] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for the server to use an emotion engine that analyzes the user's emotional state and dynamically adjust the behavior of the generated AI model, and means for reproducing the generated AI model in a virtual space, simulating the life of the animal, and interacting with the animal in accordance with the user's emotional state. This makes it possible to provide more emotionally rich interactions in real time that are in line with the user's emotional state while easing the loss of a pet.
[0991] "Image" is data containing static visual information of an animal.
[0992] "Video data" refers to data that includes dynamic visual and audio information such as animal movements and sounds.
[0993] "Terminal" refers to a device used by a user, such as a smartphone, tablet, or PC.
[0994] A "server" is a computer system that receives, analyzes, stores, and transmits data over a network.
[0995] "Analysis" is the process of extracting necessary information from image and video data and identifying its features.
[0996] An "AI model" is an artificial intelligence software model created to reproduce an animal's shape, movement patterns, sounds, etc.
[0997] An "emotion engine" is a technology for analyzing a user's emotional state from their facial expressions and voice.
[0998] "Dynamic adjustment" refers to changing the behavior of an AI model in response to the user's emotional state in real time.
[0999] A "virtual space" is a virtual environment created using computer technology.
[1000] "Simulation" means virtually reproducing the life and behavior of real animals.
[1001] "Interaction" refers to the interaction between a user and an AI model.
[1002] "Real-time" refers to processing occurring immediately, without delay.
[1003] The embodiment of this invention is a system that allows a user to upload images and video data of animals kept by the user from a terminal to a server, and the server analyzes this data and generates an AI model that captures the characteristics of the animal.Furthermore, the system uses an emotion engine to analyze the user's emotional state and dynamically adjusts the behavior of the generated AI model according to the user's emotions.
[1004] Specifically, users upload images and video data of their pet animals to a server from their smartphones, tablets, or PCs. The uploaded data is then analyzed on the server using image analysis models such as TensorFlow to recognize the animal's face, body shape, and movement patterns. If audio data is included, audio analysis tools are also used to analyze the animal's cries.
[1005] The analyzed data is then used to create an AI model that recreates the animal's appearance, movement patterns, and sounds, including a 3D model, movement animations, and audio files.
[1006] To analyze the user's emotional state using the emotion engine, the server uses camera footage and audio data sent from the user's device. It uses facial recognition algorithms to read the user's facial expressions and audio analysis tools to analyze the tone and content of the voice, dynamically detecting whether the user is happy, sad, or expressing other emotions.
[1007] The generated AI model is then recreated in a virtual space, and its behavior is dynamically adjusted according to the user's emotional state. For example, if the user is smiling, the AI model will move around energetically and wag its tail. Conversely, if the user is sad, the AI model will quietly take on a comforting behavior.
[1008] This system allows users to vividly relive their memories of their pet while easing the loss of their pet. It can provide real-time interactions that are tailored to the user's emotional state, resulting in a more emotionally enriching experience.
[1009] Specific examples
[1010] A user launches the virtual store app, takes a video of their pet dog using their smartphone camera, and uploads it to the server. The server then uses TensorFlow models to analyze the images and sounds and generate an AI model. This AI model operates in the virtual space based on the user's emotional state.
[1011] Prompt Sentence Examples
[1012] The system will be provided with an image of a dog and a short video of its barking. It will then analyze the specific behaviors and sounds the dog exhibits and create a 3D model of it. Furthermore, the system will have the dog's movements change in the virtual space depending on the user's emotions.
[1013] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1014] Step 1:
[1015] A user uses a device (smartphone, tablet, PC, etc.) to take or select images and video data of their pet dog and upload them to a server through an application. The input is the image and video data of the pet dog, and the output is the image and video files transferred to the server. This process involves the user selecting the data through the app interface and clicking the upload button.
[1016] Step 2:
[1017] The server receives the uploaded image and video data and begins analyzing it using an image analysis model such as TensorFlow. The input is the uploaded image and video data, and the output is the analyzed animal's features (face recognition, body shape identification, movement patterns, etc.). This process involves the server running an image recognition algorithm to extract the animal's face and body shape from the image and video.
[1018] Step 3:
[1019] The server generates an AI model based on the analysis results. The input is the analyzed animal's feature data, and the output is the generated AI model (3D model, movement animation, audio file, etc.). This process involves the server using a 3D model generation tool to build a 3D model that reproduces the animal's features and creating animations that include movement patterns and sounds.
[1020] Step 4:
[1021] The server uses an emotion engine to analyze the camera footage and audio data sent from the user's device and analyze the user's emotional state in real time. The input is the user's camera footage and audio data, and the output is the user's emotional state (happiness, sadness, etc.). This process involves the emotion analysis algorithm analyzing facial expressions and tone of voice to identify emotions such as joy, anger, sadness, and happiness.
[1022] Step 5:
[1023] The server dynamically adjusts the behavior of the generated AI model according to the user's emotional state. The input is the user's emotional state and the generated AI model, and the output is the adjusted AI model (a behavior pattern according to the emotional state). This process involves the server changing the behavior of the AI model based on the emotion analysis results and executing the corresponding behavior (for example, if the user is smiling, the AI model behaves cheerfully).
[1024] Step 6:
[1025] The server sends the generated AI model and adjusted behavior patterns to the user's device. The input is the adjusted AI model data, and the output is the AI model displayed on the user's device. This process includes the server encoding the data and sending it in a format suitable for the user's VR headset or AR app.
[1026] Step 7:
[1027] The user uses a device (such as a VR headset, smart glasses, or smartphone) to recreate the transmitted AI model in a virtual space and interact with the AI model according to the user's emotional state. The input is the adjusted AI model data and the user's emotional state, and the output is the user's interaction with an emotive virtual pet. This process involves the user operating the device and interacting with the AI model in the virtual space.
[1028] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1029] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1030] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1031] [Fourth embodiment]
[1032] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1033] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1035] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1036] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1037] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1039] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1040] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1041] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1043] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1045] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses image and video data of pets and analyzes them to generate an AI model that reproduces the appearance and behavior of the pet, and then recreates it in a virtual space.
[1046] Explanation of system program processing
[1047] 1. Data collection and upload
[1048] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[1049] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[1050] 2. Data processing and AI model generation
[1051] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[1052] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[1053] The server generates an AI model that replicates the pet's appearance and behavior based on the analysis results, and the generated AI model is tuned to replicate the pet's characteristics as faithfully as possible.
[1054] Example: The server creates a 3D model based on identified face, body shape, and fur color data, reproduces animal sounds using audio data, and generates animations based on movement patterns.
[1055] 3. Providing AI models and verifying their operation
[1056] The server sends the generated AI model to the user's device, which must support a VR headset or AR app.
[1057] Example: A server sends the generated AI model, including the 3D model and audio files, in a format compatible with the user's VR headset.
[1058] Users can then load the AI model using VR or AR devices, reunite with their pet in a virtual space, and enjoy interacting with it.
[1059] Example: A user puts on a VR headset, launches a system-specific app, loads a 3D model of their pet dog, and plays with it in a virtual garden.
[1060] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of their pet, providing users with a highly realistic and moving experience.
[1061] The processing flow will be explained below.
[1062] Step 1:
[1063] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[1064] Step 2:
[1065] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[1066] Step 3:
[1067] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[1068] Step 4:
[1069] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[1070] Step 5:
[1071] The server begins analyzing the stored image and video data, which includes:
[1072] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[1073] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[1074] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[1075] Step 6:
[1076] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[1077] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[1078] Audio file: The extracted call data is reproduced using voice synthesis technology.
[1079] Movement animation: Generates animation based on the extracted movement patterns.
[1080] Step 7:
[1081] The server sends the generated AI model to the user's device. The data sent must be in a format compatible with VR or AR.
[1082] Step 8:
[1083] The user launches a VR headset or AR app and loads the sent AI model, recreating the experience of interacting with a pet in a virtual space.
[1084] Step 9:
[1085] The device operates the AI model in response to user input within the virtual space, allowing interaction with the pet. This recreates the natural presence and vitality of dinosaurs. This scenario is translated into English.
[1086] Example 1
[1087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1088] Pet loss is a serious psychological issue for many people. The present invention aims to alleviate pet loss by using pet image and video data to create an AI model that faithfully reproduces a pet's appearance and behavior. However, conventional technologies have difficulty accurately capturing a pet's characteristics and lack the means to allow users to fully enjoy interacting with a pet recreated in a virtual space. The purpose of the present invention is to solve these problems.
[1089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1090] In this invention, the server includes means for uploading image and video data of the pet from the terminal to the server, means for analyzing the uploaded data and generating an AI model that captures the characteristics of the pet, means for reproducing the generated AI model in a virtual space and simulating the life of the pet, and means for the user to display the generated AI model in the virtual space using a VR or AR device and interact with it. This allows the generation of a highly accurate AI model that faithfully reproduces the characteristics of the pet, allowing the user to enjoy interacting with the pet in the virtual space.
[1091] A "server" is a computing device that receives and analyzes data and generates and serves AI models.
[1092] A "terminal" is a computer device that is operated by a user and that uploads and receives data.
[1093] A "user" is an individual who uses a terminal to provide image and video data of their pet to the server and interacts with the generated AI model.
[1094] "Image and video data" refers to still image and video files that record the appearance and behavior of your pet.
[1095] "Upload" is the act of sending data from a user's terminal to a server.
[1096] "Analysis" is the process by which the server processes the received data to extract and identify the pet's characteristics.
[1097] An "AI model" is an artificial intelligence model that is generated based on analyzed data and is used to reproduce the appearance and behavior of a pet.
[1098] "Virtual space" is a virtual three-dimensional space generated by computer simulation.
[1099] "Simulation" is the process of recreating the life and behavior of a real pet in a virtual space.
[1100] "VR equipment" refers to equipment for experiencing virtual reality, including devices such as headsets.
[1101] An "AR device" is a device for experiencing augmented reality, and includes devices such as smartphones and AR glasses.
[1102] "Interaction" refers to two-way communication and manipulation between the user and a pet recreated in a virtual space.
[1103] This invention is an AI model generation system aimed at alleviating the loss of a pet. This system uses images and video data of the user's pet and analyzes them to generate an AI model that reproduces the pet's appearance and behavior, then recreates it in a virtual space.
[1104] First, users upload image and video data of their pets to the server from their own devices. Devices can be smartphones, tablets, or PCs, and data can be easily selected and sent via a dedicated upload interface. For example, a user can select 10 photos of dogs and three videos from their smartphone gallery and click the "Upload" button.
[1105] The server then receives the uploaded data and analyzes it. It uses a TensorFlow-based image recognition model to recognize the animal's face, body shape, and coat color from the image. For videos, it uses a time-series data analysis tool to analyze movement patterns, and an audio analysis tool to analyze the animal's barks from the audio data. For example, the server uses an image analysis engine to identify the dog's face and body shape, and a video analysis engine to extract movement patterns. It also uses an audio analysis engine to extract bark characteristics.
[1106] Based on this analytical data, the server generates a highly accurate AI model that reproduces the pet's appearance and behavior. The server uses 3D modeling tools to create a three-dimensional model of the pet and applies movement animations to it. It also adds audio data to reproduce its cries. For example, the server creates a 3D model based on the identified face, body shape, and fur color data, implements movement animations based on movement patterns, and adds cries data.
[1107] The generated AI model is then sent back to the device. The user's device must support a VR headset or AR application. The server checks file compatibility, converts the format if necessary, and then sends the AI model to the device. The user can then load the AI model using a VR or AR device, reunite with their pet in a virtual space, and enjoy interacting with it. For example, a user can put on a VR headset, launch a dedicated app, load a 3D model of their pet, and play with it in a virtual garden.
[1108] An example of a prompt is as follows:
[1109] "I've uploaded photos and videos of my dog. Please generate a realistic 3D model based on them."
[1110] Through this system, users can ease the shock of losing their pet and receive support as they move on to a new life by recreating memories of their pet. In addition, the system generates an AI model that accurately reproduces the characteristics of a pet, providing users with a highly realistic and moving experience.
[1111] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1112] Step 1:
[1113] The user selects and uploads pet image and video data from the device. Specifically, the user opens the gallery or file manager on the device and selects multiple images and videos. For example, the user selects 10 dog photos and 3 videos from the smartphone gallery and clicks the "Upload" button. The image and video data selected by the user are used as input, and this data is sent to the server as output.
[1114] Step 2:
[1115] The server receives the image and video data sent by the user. Specifically, the server checks the integrity of the data and verifies that all files have been received correctly. The input is the image and video data from the user, and the output is the storage and verification of the data.
[1116] Step 3:
[1117] The server analyzes the received image data. It uses a TensorFlow-based image recognition model to perform facial recognition, body shape identification, and coat color analysis. Specifically, the image analysis engine extracts the facial position, body shape outline, and coat color characteristics for each image. For example, it identifies the facial feature points (eyes, nose, and mouth) of a dog and stores that information in a database. Using the received image data as input, facial recognition and body shape identification data are generated as output.
[1118] Step 4:
[1119] The server analyzes the received video data. It uses a time-series data analysis tool to analyze movement patterns and extract the pet's barks from the audio data. To determine specific behaviors, the video analysis engine analyzes each frame and extracts behavioral patterns while detecting the continuity of movement. The audio analysis engine also extracts sound wave characteristics and detects bark patterns and frequencies. For example, a dog's behaviors such as running, eating, and sleeping are recorded along a time axis. The received video data is used as input, and movement patterns and bark data are generated as output.
[1120] Step 5:
[1121] The server generates an AI model based on the analysis data. 3D modeling tools are used to create a three-dimensional model of the pet, and movement animations are applied to that model. Analyzed audio data is then added to recreate the animal's barks. Specific movements are then converted by the server into a 3D model of the face, body shape, and coat color, and the movement patterns obtained from video analysis are incorporated into the animation using motion capture technology. For example, a dog's running motion and barking sounds are realistically recreated. The analysis results are used as input, and a complete AI model is generated as output.
[1122] Step 6:
[1123] The server sends the generated AI model to the user's device. The transmission format is compatible with VR or AR. Specifically, the server checks file compatibility and converts the format if necessary. The data is then seamlessly transferred to the user's device. The generated AI model is used as input, and the AI model is sent to the user's device as output.
[1124] Step 7:
[1125] The user loads the sent AI model using a VR or AR device. Using a dedicated application, the user interacts with their pet in a virtual space. Specifically, the user puts on a VR headset, launches a dedicated application, and loads the AI model. The user can then play with their pet dog in a virtual garden. The AI model sent from the server is used as input, and interaction in a virtual space is realized as output.
[1126] (Application example 1)
[1127] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1128] To alleviate the psychological shock and sadness caused by pet loss, there is a need for a system that can accurately recreate memories of past pets and allow users to re-interact with their pets in a virtual space. Furthermore, there is a need for a method to enable such a system to provide a moving experience without users having to physically visit a store.
[1129] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1130] In this invention, the server includes means for uploading image and video data of the animal from the terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for reproducing the generated AI model in a virtual reality space and simulating the life of the animal, and means for the user to access the virtual space on the terminal and enjoy interacting with the animal's AI model. This allows the user to ease the loss of their pet and have a moving experience through interacting with the pet in the virtual space.
[1131] 1. "Domestic animals" refers to all animals kept as pets in homes.
[1132] 2. "Image and video data" refers to digital data consisting of photographs and videos of pets.
[1133] 3. "Device" refers to electronic devices that can connect to the Internet, such as smartphones, tablets, and PCs.
[1134] 4. "Server" refers to a central processing unit used for data analysis and AI model generation.
[1135] 5. "Upload" refers to the act of sending digital data from a device to a server.
[1136] 6. “Analysis” refers to the act of extracting information based on image or video data.
[1137] 7. "Feature-capturing AI model" refers to an artificial intelligence model that faithfully reproduces the appearance and behavior of a pet.
[1138] 8. "Virtual reality space" refers to a virtual environment created using VR technology.
[1139] 9. "Simulate" refers to the act of recreating real-world actions or environments.
[1140] 10. "Accessing a virtual space" refers to the act of a user connecting to a virtual reality environment using a VR headset or device.
[1141] 11. “Interaction” refers to the act of a user interacting with an AI model.
[1142] This invention is a system aimed at alleviating the loss of a pet. It analyzes image and video data of pets to generate an AI model of the pet, which is then recreated in a virtual space, allowing users to relive memories with their pet and providing a moving experience.
[1143] This system consists of the following main elements:
[1144] 1. Terminal
[1145] Users use devices such as smartphones or tablets to collect images and video data of their pets and upload them to a server. A dedicated application is installed on the device, allowing users to easily select and send data.
[1146] 2. Server
[1147] The server receives the uploaded image and video data and analyzes it, which includes the following processes:
[1148] Image analysis uses image recognition software such as TensorFlow to recognize pet faces, identify body shapes, and analyze fur color.
[1149] Video analysis uses video analysis software such as OpenCV to extract pet movements and behavior patterns.
[1150] Based on the analysis results, an AI model is generated using a 3D modeling tool such as Unity. This AI model is generated by integrating the analysis data from TensorFlow and OpenCV to faithfully reproduce the appearance and behavior of the pet.
[1151] 3. Virtual Reality Space
[1152] The AI model generated by the server is reproduced in a virtual reality space, which can be accessed through a VR headset or an AR app. The generated AI model is configured to operate in real time in the virtual reality space using VR development tools such as Unity.
[1153] By wearing a VR headset, users can access a virtual reality space and enjoy interacting with their pets.
[1154] 4. User Interaction
[1155] Users can interact with the AI model in real time through a VR headset. For example, if a user calls a pet's name in the virtual space, the AI model will respond by turning toward the user or performing a specific action, providing an experience similar to interacting with a real pet.
[1156] A specific example of this system is shown below.
[1157] Specific examples
[1158] Users upload using their smartphones
[1159] Users select 10 photos of their dog and three videos from their smartphone gallery and upload them to the server using a dedicated app.
[1160] Server analysis process
[1161] The server uses TensorFlow to analyze images and OpenCV to analyze videos, which allows for facial recognition, body shape identification, sound analysis, and movement pattern extraction of pets.
[1162] Generating AI models
[1163] Based on the analysis results, the server uses Unity to generate a 3D model and create an AI model.
[1164] Reproduction in virtual reality space
[1165] Users put on a VR headset and access the virtual reality space. They then launch the system's dedicated app, load a 3D model of their pet dog, and play with it in a virtual garden.
[1166] Prompt Sentence Examples
[1167] AI model generation prompts
[1168] Please generate a 3D model that is as faithful as possible to my pet's appearance based on past images and videos of my pet.
[1169] Interaction prompts
[1170] Place the generated 3D model in your virtual garden and simulate its movements as you play with your pet.
[1171] In this way, users can ease the loss of their pet and relive their memories with their pet through an emotional experience in a virtual reality space.
[1172] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1173] Step 1:
[1174] The user uploads pet image and video data using a device. Specifically, the user launches a dedicated application on the device, selects pet photos and videos from the gallery, and clicks the "Upload" button. This sends the selected data to the server. The input is the pet image and video data, and the output is the data transfer to the server.
[1175] Step 2:
[1176] The server receives the uploaded image and video data. The received data is temporarily stored in the server's storage. The input is the image and video data sent by the user, and the output is the data saved in the server's storage.
[1177] Step 3:
[1178] The server uses TensorFlow to analyze the received image data. First, the server preprocesses the image data and performs facial recognition, body shape identification, and fur color analysis. Specifically, the server converts the image data into RGB format and inputs each data into a model to extract features. The input is the saved image data, and the output is feature data related to the pet's facial recognition, body shape, and fur color.
[1179] Step 4:
[1180] The server uses OpenCV to analyze the received video data. The server extracts motion patterns from the video data frame by frame and generates time-series data. Specifically, it extracts frames from the video data and detects motion changes using methods such as optical flow. The input is the saved video data, and the output is time-series data on motion patterns.
[1181] Step 5:
[1182] The server integrates the analyzed image and video data to generate an AI model that captures the pet's characteristics. To do this, a 3D model is built using Unity based on the analysis results of TensorFlow and OpenCV, and movement patterns are added. The input is feature data and time-series data, and the output is the pet's AI model.
[1183] Step 6:
[1184] The server sends the generated AI model to the user's device. Specifically, the generated AI model is sent as data in a format compatible with VR headsets and AR apps. The input is the generated AI model, and the output is data transferred to the user's device (VR headset or smartphone).
[1185] Step 7:
[1186] The user puts on a VR headset and launches a dedicated app to load the generated AI model. Specifically, the user selects the sent AI model from the VR headset app and accesses the virtual space. To interact with the pet in the VR space, the user interacts with the AI model through voice recognition and gesture control. The input is the AI model and the VR headset, and the output is the interaction with the pet in the virtual space.
[1187] Through these steps, the system can ease the loss of a pet and provide a moving experience for users.
[1188] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1189] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[1190] Explanation of system program processing
[1191] 1. Data collection and upload
[1192] The user selects images and videos of their pet and sends them to the server using the device's upload interface.
[1193] Example: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[1194] 2. Data processing and AI model generation
[1195] The server receives the uploaded data and first performs image analysis, which includes facial recognition, body shape recognition, and fur color analysis. It also analyzes the pet's movements and behavior patterns from the video.
[1196] Example: The server uses TensorFlow-based image recognition models to identify pet faces and body shapes, audio analysis tools to analyze sounds, and time series data analysis tools to extract movements.
[1197] The server uses the analysis results to generate an AI model that replicates the pet's appearance and behavior. This model includes the following elements:
[1198] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[1199] Audio file: The extracted call data is reproduced using voice synthesis technology.
[1200] Movement animation: Generates animation based on the extracted movement patterns.
[1201] 3. Sentiment analysis and AI model tuning
[1202] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[1203] Example: The server uses a facial recognition algorithm to detect a user's smiling or sad facial expressions, and a speech analysis tool to analyze changes in the tone of their voice.
[1204] 4. Providing AI models and verifying their operation
[1205] The server generates an AI model and sends the emotion analysis results to the user's device. The data is in a format compatible with VR or AR.
[1206] Example: A server sends data containing the generated 3D model, audio files, and emotion analysis results to a user's VR headset.
[1207] The user launches a VR headset or AR app and loads the AI model, recreating the experience of interacting with a pet in a virtual space. Furthermore, the AI model's behavior changes depending on the user's emotional state.
[1208] Example: A user puts on a VR headset, launches a dedicated app for the system, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[1209] This system not only helps users ease the shock of losing a pet and recreate memories with their pet, but also allows them to enjoy interactions that are in tune with their emotions. The introduction of an emotion engine allows users to have a more realistic and moving experience, and provides psychological support to help them move on to a new life.
[1210] The processing flow will be explained below.
[1211] Step 1:
[1212] The user selects images and videos of the pet from the terminal's file browser and prepares them for transmission to the server through the system's upload interface.
[1213] Step 2:
[1214] The device compresses the images and videos selected by the user into the appropriate format and prepares them for transmission to the server: images into JPEG format, and videos into MP4 format.
[1215] Step 3:
[1216] The device sends the compressed data to the server over the internet using a POST request to the server's upload API endpoint.
[1217] Step 4:
[1218] The server stores the received image and video data in secure storage, specifically in a cloud storage service, and records the metadata in a database.
[1219] Step 5:
[1220] The server begins analyzing the stored image and video data, which includes:
[1221] Image analysis: Use image recognition algorithms (e.g., TensorFlow) to identify your pet's face, body shape, and coat color.
[1222] Audio analysis: Pet sounds are extracted from video data and analyzed using an audio analysis tool.
[1223] Motion analysis: Pet movements and behavior patterns are extracted from video data using time series data analysis tools.
[1224] Step 6:
[1225] The server generates an AI model based on the analysis results, capturing the characteristics of the pet. This model includes the following elements:
[1226] 3D model: A 3D model is created using analyzed face, body shape, and fur color data.
[1227] Audio file: The extracted call data is reproduced using voice synthesis technology.
[1228] Movement animation: Generates animation based on the extracted movement patterns.
[1229] Step 7:
[1230] The server activates an emotion engine that detects the user's emotions. The emotion engine uses a camera and microphone to analyze the user's emotions from their facial expressions and voice.
[1231] Example: The server uses a facial recognition algorithm to detect the user's facial expressions (e.g., smiling, sad), and a voice analysis tool to analyze changes in voice tone.
[1232] Step 8:
[1233] The server then sends the generated AI model and emotion analysis results to the user's device in a format compatible with VR or AR.
[1234] Step 9:
[1235] The user launches the VR headset or AR app and loads the AI model, which then begins interacting with the pet in the virtual space.
[1236] Example: A user puts on a VR headset, launches a dedicated app, and loads a 3D model of their pet dog. If the user is smiling, the AI model of the pet will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of the pet will quietly cuddle with the user.
[1237] This series of processes not only helps users ease the shock of losing a pet, but also provides an emotionally responsive, interactive experience.
[1238] Example 2
[1239] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1240] Users who are grieving the loss of their pet need a system that allows them to not only reminisce about their memories with their deceased pet, but also enjoy emotionally rich interactions. However, existing systems lack the functionality to empathize with the user's emotions and are insufficient in providing psychological support. Therefore, a system that dynamically analyzes the user's emotions and adjusts the behavior of the pet's AI model to provide more emotionally rich interactions is desired.
[1241] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1242] In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for recreating the generated AI model in a virtual space and simulating the life of the animal, means for activating an emotion engine that detects a user's emotions and analyzing emotions from the user's facial expressions and voice using a camera and microphone, and means for adjusting the operation of the AI model based on the emotion analysis results, thereby enabling users to ease the shock of losing their pet and enjoy emotionally sensitive interactions.
[1243] The "server" is a central processing unit that receives and analyzes data uploaded from the user's device, and generates and adjusts the AI model.
[1244] A "terminal" is a device used by a user, such as a computer or smartphone, through which data is uploaded or downloaded.
[1245] "Upload" means a function or interface for transmitting image or video data from a terminal to a server.
[1246] "Analysis" refers to the process of processing the data received by the server and extracting the animal's characteristics and behavioral patterns.
[1247] An "AI model" is an artificial intelligence model that is generated based on analyzed data and reproduces the appearance and behavior of animals.
[1248] A "virtual space" is an imaginary environment generated using computer graphics and related technologies.
[1249] The "simulation" means virtually recreating the generated AI model in a virtual space and imitating the animal's life.
[1250] An "emotion engine" is a collection of software and algorithms for analyzing a user's emotions.
[1251] A "camera" is a photographing device for capturing the user's facial expression.
[1252] A "microphone" is a recording device for capturing the user's voice.
[1253] "Emotion analysis" is the process of determining a user's current emotional state based on data captured using a camera or microphone.
[1254] "Behavior adjustment" is the process of changing the behavior and reactions of an AI model based on the results of emotion analysis.
[1255] This invention combines an AI model generation system with an emotion engine that recognizes the user's emotions, aiming to alleviate the loss of a pet. This system dynamically analyzes the user's emotional state and adjusts the behavior of the pet's AI model, making interactions with the user more emotionally enriching.
[1256] The system of the present invention mainly uses the following hardware and software:
[1257] Hardware: Servers, devices (smartphones, PCs, VR headsets), cameras, microphones
[1258] Software: TensorFlow-based image recognition models, voice analysis tools, time-series data analysis tools, 3D modeling software, voice synthesis technology, emotion recognition algorithms
[1259] Specific embodiments will be described below.
[1260] Data collection and upload
[1261] Users select images and videos of their pets from their devices and send them to the server using the system's upload interface. For example, a user selects 10 dog photos and 3 videos from their smartphone gallery and clicks the "Upload" button.
[1262] Data processing and AI model generation
[1263] The server receives the uploaded images and videos and stores them in the initial data storage. It then uses a TensorFlow-based image recognition model to analyze the images and perform facial recognition, body shape identification, and fur color analysis of the pet. It also analyzes the video to extract the pet's movements and behavior patterns. Based on these analysis results, the server generates an AI model with the following elements:
[1264] 1. 3D model: A 3D model is created using the analyzed face, body shape, and fur color data.
[1265] 2. Audio file: The extracted call data is reproduced using voice synthesis technology.
[1266] 3. Motion animation: Generate animation based on the extracted motion patterns.
[1267] Sentiment analysis and AI model tuning
[1268] The server activates an emotion engine to monitor the user's emotions, using a camera and microphone to capture the user's facial expressions and voice in real time. For example, the camera captures the user's face and uses a facial recognition algorithm to detect smiling or sad expressions. The audio recorded by the microphone is then analyzed for changes in tone using a voice analysis tool. As a result, the server adjusts the behavior of the AI model according to the user's emotional state. For example, if the user is smiling, the AI pet will move around energetically. Conversely, if the user is sad, the AI pet will quietly cuddle.
[1269] Providing AI models and verifying their operation
[1270] The server sends the generated AI model and emotion analysis results to the user's device (e.g., a VR headset). The user then launches the VR headset or AR app and loads the sent AI model, recreating interaction with a pet in a virtual space. The AI pet's behavior changes depending on the user's emotional state. For example, a user puts on a VR headset, launches a system-specific app, and loads a 3D model of their beloved dog. If the user is smiling, the AI model of their beloved dog will react by moving around energetically and wagging its tail. Conversely, if the user is sad, the AI model of their beloved dog will quietly cuddle with them.
[1271] Prompt Sentence Examples
[1272] "I want to be reunited with my pet. Please use these photos and videos to create an AI model."
[1273] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1274] Step 1:
[1275] User selects data and uploads it on the device
[1276] The user selects images and video data of their pet from their device (smartphone or PC).
[1277] What happens: A user selects 10 photos of dogs and 3 videos from their smartphone gallery and clicks the "Upload" button.
[1278] Input: Image and video files stored in your phone's gallery.
[1279] Output: Image and video data sent to the server.
[1280] Step 2:
[1281] The server receives and stores the data
[1282] The server stores the image and video data received from the terminal in an initial data storage.
[1283] Specific operation: The server saves the uploaded 10 images and 3 videos to the file system or database.
[1284] Input: Image and video data submitted by the user.
[1285] Output: Image and video files saved to storage.
[1286] Step 3:
[1287] The server analyzes the data
[1288] The server analyzes the stored data using a TensorFlow-based image recognition model.
[1289] How it works: The server uses image recognition models to identify pet faces, body shapes, and coat colors, and uses time-series data analysis tools to extract movement and behavior patterns from video.
[1290] Input: Stored image and video data.
[1291] Output: Face recognition results, body shape identification results, coat color data, movement pattern data.
[1292] Step 4:
[1293] The server generates the AI model
[1294] Based on the analysis results, the server generates an AI model that includes a 3D model, audio files, and motion animations.
[1295] Specific movements: The server uses 3D modeling software to create a 3D model based on the pet's face, body shape, and fur color data. It then uses voice synthesis technology to reproduce the pet's barking sound and generates movement animations based on the movement pattern data.
[1296] Input: Face recognition results, body shape identification results, coat color data, movement pattern data.
[1297] Output: 3D model, reconstructed audio files, movement animations.
[1298] Step 5:
[1299] The server starts the emotion engine
[1300] The server detects the user's emotions by capturing the user's facial expressions and voice in real time using the device's camera and microphone.
[1301] How it works: The server takes a picture of the user's face with the device's camera, uses a facial recognition algorithm to detect smiling or sad expressions, and uses a voice analysis tool to analyze changes in the tone of the voice recorded by the microphone.
[1302] Input: Video and audio recording of the user's face.
[1303] Output: User sentiment analysis results.
[1304] Step 6:
[1305] The server adjusts the AI model
[1306] The server adjusts the behavior of the generated AI model based on the results of emotion analysis.
[1307] Specific behavior: Based on the results of emotion analysis, the AI model's animations and reactions are changed in real time. For example, if the user is smiling, the AI pet will move around energetically and wag its tail. If the user is sad, the AI pet will quietly cuddle.
[1308] Input: User sentiment analysis results, generated AI model.
[1309] Output: The movement animations and reactions of the tuned AI model.
[1310] Step 7:
[1311] The server sends the AI model to the device.
[1312] The server sends the generated AI model and emotion analysis results to the user's device.
[1313] Specific operation: The server sends data including the generated 3D model, audio files, movement animations, and emotion analysis results to the user's VR headset.
[1314] Input: Tuned AI model, sentiment analysis results.
[1315] Output: Data sent to the user's device.
[1316] Step 8:
[1317] The user launches the AI model and checks its operation
[1318] The user launches a VR headset or AR app and displays the sent AI model in a virtual space.
[1319] Specific operation: The user puts on the VR headset, launches the system's dedicated app, loads the 3D model of the pet, enjoys interacting with the AI pet model, and sees its behavior change in real time.
[1320] Input: User's emotional state, transmitted data.
[1321] Output: Interaction with a pet in a virtual space.
[1322] This system allows users to relive memories with their pets while enjoying emotionally rich interactions.
[1323] (Application example 2)
[1324] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1325] There is a need for a system that can vividly recreate memories of pets, empathize with the user's emotional state, and provide real-time interaction while easing the psychological shock of pet loss. However, conventional systems lack the technology to flexibly respond to changes in the user's emotions, making it difficult to achieve more emotionally rich interactions.
[1326] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading image and video data of a pet animal from a terminal to the server, means for the server to analyze the uploaded data and generate an AI model that captures the characteristics of the animal, means for the server to use an emotion engine that analyzes the user's emotional state and dynamically adjust the behavior of the generated AI model, and means for reproducing the generated AI model in a virtual space, simulating the life of the animal, and interacting with the animal in accordance with the user's emotional state. This makes it possible to provide more emotionally rich interactions in real time that are in line with the user's emotional state while easing the loss of a pet.
[1327] "Image" is data containing static visual information of an animal.
[1328] "Video data" refers to data that includes dynamic visual and audio information such as animal movements and sounds.
[1329] "Terminal" refers to a device used by a user, such as a smartphone, tablet, or PC.
[1330] A "server" is a computer system that receives, analyzes, stores, and transmits data over a network.
[1331] "Analysis" is the process of extracting necessary information from image and video data and identifying its features.
[1332] An "AI model" is an artificial intelligence software model created to reproduce an animal's shape, movement patterns, sounds, etc.
[1333] An "emotion engine" is a technology for analyzing a user's emotional state from their facial expressions and voice.
[1334] "Dynamic adjustment" refers to changing the behavior of an AI model in response to the user's emotional state in real time.
[1335] A "virtual space" is a virtual environment created using computer technology.
[1336] "Simulation" means virtually reproducing the life and behavior of real animals.
[1337] "Interaction" refers to the interaction between a user and an AI model.
[1338] "Real-time" refers to processing occurring immediately, without delay.
[1339] The embodiment of this invention is a system that allows a user to upload images and video data of animals kept by the user from a terminal to a server, and the server analyzes this data and generates an AI model that captures the characteristics of the animal.Furthermore, the system uses an emotion engine to analyze the user's emotional state and dynamically adjusts the behavior of the generated AI model according to the user's emotions.
[1340] Specifically, users upload images and video data of their pet animals to a server from their smartphones, tablets, or PCs. The uploaded data is then analyzed on the server using image analysis models such as TensorFlow to recognize the animal's face, body shape, and movement patterns. If audio data is included, audio analysis tools are also used to analyze the animal's cries.
[1341] The analyzed data is then used to create an AI model that recreates the animal's appearance, movement patterns, and sounds, including a 3D model, movement animations, and audio files.
[1342] To analyze the user's emotional state using the emotion engine, the server uses camera footage and audio data sent from the user's device. It uses facial recognition algorithms to read the user's facial expressions and audio analysis tools to analyze the tone and content of the voice, dynamically detecting whether the user is happy, sad, or expressing other emotions.
[1343] The generated AI model is then recreated in a virtual space, and its behavior is dynamically adjusted according to the user's emotional state. For example, if the user is smiling, the AI model will move around energetically and wag its tail. Conversely, if the user is sad, the AI model will quietly take on a comforting behavior.
[1344] This system allows users to vividly relive their memories of their pet while easing the loss of their pet. It can provide real-time interactions that are tailored to the user's emotional state, resulting in a more emotionally enriching experience.
[1345] Specific examples
[1346] A user launches the virtual store app, takes a video of their pet dog using their smartphone camera, and uploads it to the server. The server then uses TensorFlow models to analyze the images and sounds and generate an AI model. This AI model operates in the virtual space based on the user's emotional state.
[1347] Prompt Sentence Examples
[1348] The system will be provided with an image of a dog and a short video of its barking. It will then analyze the specific behaviors and sounds the dog exhibits and create a 3D model of it. Furthermore, the system will have the dog's movements change in the virtual space depending on the user's emotions.
[1349] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1350] Step 1:
[1351] A user uses a device (smartphone, tablet, PC, etc.) to take or select images and video data of their pet dog and upload them to a server through an application. The input is the image and video data of the pet dog, and the output is the image and video files transferred to the server. This process involves the user selecting the data through the app interface and clicking the upload button.
[1352] Step 2:
[1353] The server receives the uploaded image and video data and begins analyzing it using an image analysis model such as TensorFlow. The input is the uploaded image and video data, and the output is the analyzed animal's features (face recognition, body shape identification, movement patterns, etc.). This process involves the server running an image recognition algorithm to extract the animal's face and body shape from the image and video.
[1354] Step 3:
[1355] The server generates an AI model based on the analysis results. The input is the analyzed animal's feature data, and the output is the generated AI model (3D model, movement animation, audio file, etc.). This process involves the server using a 3D model generation tool to build a 3D model that reproduces the animal's features and creating animations that include movement patterns and sounds.
[1356] Step 4:
[1357] The server uses an emotion engine to analyze the camera footage and audio data sent from the user's device and analyze the user's emotional state in real time. The input is the user's camera footage and audio data, and the output is the user's emotional state (happiness, sadness, etc.). This process involves the emotion analysis algorithm analyzing facial expressions and tone of voice to identify emotions such as joy, anger, sadness, and happiness.
[1358] Step 5:
[1359] The server dynamically adjusts the behavior of the generated AI model according to the user's emotional state. The input is the user's emotional state and the generated AI model, and the output is the adjusted AI model (a behavior pattern according to the emotional state). This process involves the server changing the behavior of the AI model based on the emotion analysis results and executing the corresponding behavior (for example, if the user is smiling, the AI model behaves cheerfully).
[1360] Step 6:
[1361] The server sends the generated AI model and adjusted behavior patterns to the user's device. The input is the adjusted AI model data, and the output is the AI model displayed on the user's device. This process includes the server encoding the data and sending it in a format suitable for the user's VR headset or AR app.
[1362] Step 7:
[1363] The user uses a device (such as a VR headset, smart glasses, or smartphone) to recreate the transmitted AI model in a virtual space and interact with the AI model according to the user's emotional state. The input is the adjusted AI model data and the user's emotional state, and the output is the user's interaction with an emotive virtual pet. This process involves the user operating the device and interacting with the AI model in the virtual space.
[1364] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1365] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1366] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1367] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1368] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1369] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1370] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1371] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1372] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1373] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1374] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1375] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1376] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1377] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1378] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1379] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1380] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1381] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1382] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1383] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1384] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1385] The following is further disclosed regarding the above embodiment.
[1386] (Claim 1)
[1387] A means for uploading image and video data of the animals to be kept from the terminal to a server;
[1388] The server analyzes the uploaded data and generates an AI model that captures the characteristics of the animal;
[1389] means for reproducing the generated AI model in a virtual space and simulating the life of the animal;
[1390] A system including:
[1391] (Claim 2)
[1392] 2. The system of claim 1, wherein the server includes means for extracting animal facial recognition, body shape identification, and movement patterns from the image and video data.
[1393] (Claim 3)
[1394] 2. The system according to claim 1, further comprising means for causing the AI model to operate and react in response to user operations in the virtual space.
[1395] "Example 1"
[1396] (Claim 1)
[1397] A means for uploading image and video data of the animals to be kept from the terminal to a server;
[1398] The server analyzes the uploaded data and generates an AI model that captures the characteristics of the animal;
[1399] means for reproducing the generated AI model in a virtual space and simulating the life of the animal;
[1400] A means for a user to display the generated AI model in a virtual space using a VR or AR device and interact with it;
[1401] A system including:
[1402] (Claim 2)
[1403] 2. The system according to claim 1, wherein the server includes means for extracting animal facial recognition, body shape identification, movement patterns, and sounds from audio data from the image and video data.
[1404] (Claim 3)
[1405] 2. The system according to claim 1, further comprising means for causing the AI model to operate and react in response to user operations in the virtual space.
[1406] "Application Example 1"
[1407] (Claim 1)
[1408] A means for uploading image and video data of the animals to be kept from the terminal to a server;
[1409] The server analyzes the uploaded data and generates an AI model that captures the characteristics of the animal;
[1410] A means for reproducing the generated AI model in a virtual reality space and simulating the life of the animal;
[1411] A means for users to access the virtual space on their devices and enjoy interacting with AI animal models,
[1412] A system including:
[1413] (Claim 2)
[1414] 2. The system of claim 1, wherein the server includes means for extracting animal facial recognition, body shape identification, and movement patterns from the image and video data.
[1415] (Claim 3)
[1416] 2. The system according to claim 1, further comprising means for causing the AI model to operate and react in response to user operations in the virtual reality space.
[1417] "Example 2: Combining Emotion Engines"
[1418] (Claim 1)
[1419] A means for uploading image and video data of the animals to be kept from the terminal to a server;
[1420] The server analyzes the uploaded data and generates an AI model that captures the characteristics of the animal;
[1421] means for reproducing the generated AI model in a virtual space and simulating the life of the animal;
[1422] A means for activating an emotion engine that detects the user's emotions and analyzing the emotions from the user's facial expressions and voice using a camera and a microphone;
[1423] means for adjusting the operation of the AI model based on the emotion analysis result;
[1424] A system including:
[1425] (Claim 2)
[1426] 2. The system of claim 1, wherein the server includes means for extracting animal facial recognition, body shape identification, and movement patterns from the image and video data.
[1427] (Claim 3)
[1428] 2. The system according to claim 1, further comprising means for causing the AI model to operate and react in response to user operations in the virtual space.
[1429] "Application example 2 when combining emotion engines"
[1430] (Claim 1)
[1431] A means for uploading image and video data of the animals to be kept from the terminal to a server;
[1432] The server analyzes the uploaded data and generates an AI model that captures the characteristics of the animal;
[1433] The server uses an emotion engine that analyzes the user's emotional state to dynamically adjust the behavior of the generated AI model;
[1434] means for reproducing the generated AI model in a virtual space, simulating the life of the animal, and interacting with the animal according to the emotional state of a user;
[1435] A system including:
[1436] (Claim 2)
[1437] 2. The system of claim 1, wherein the server includes means for extracting animal facial recognition, body shape identification, and movement patterns from the image and video data.
[1438] (Claim 3)
[1439] 2. The system according to claim 1, further comprising means for the AI model to operate and react in real time in the virtual space in response to the user's operations and emotional state. [Explanation of symbols]
[1440] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for uploading image and video data of the animals to be kept from the terminal to a server; The server analyzes the uploaded data and generates an AI model that captures the characteristics of the animal; means for reproducing the generated AI model in a virtual space and simulating the life of the animal; A system including:
2. 2. The system of claim 1, wherein the server includes means for extracting animal facial recognition, body shape identification, and movement patterns from the image and video data.
3. 2. The system according to claim 1, further comprising means for causing said AI model to operate and react in response to a user's operation in said virtual space.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A