system
The system recreates pet avatars using photos, videos, and social media data, integrating image and multimodal analysis with AR/VR for natural conversations, addressing the lack of emotional support and interaction in existing technologies.
Patent Information
- Application Number
- JP2024130404
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Current technology lacks the ability to accurately recreate the appearance and behavior of deceased or desired pets in a virtual space, enabling natural conversations and diverse interactions, failing to provide emotional support and healing for pet owners.
A system that uses photos, videos, and social media posting data to generate a pet avatar, incorporating image generation, multimodal analysis for gestures and sounds, and a large-scale language model for natural conversations, integrated with AR and VR environments for interactive experiences.
Enables realistic and interactive pet avatars that provide emotional healing and a rich communication experience, allowing users to interact with their pets in multiple dimensions.
Smart Images

Figure 2026028106000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Pet owners suffering from pet loss and people facing difficulties in caring for pets often lack emotional support and healing. There is also a need to recreate memories with pets in a more interactive and realistic way than just photos and videos. However, current technology lacks a way to accurately recreate the appearance and behavior of deceased or desired pets in a virtual space, enabling natural conversations and diverse interactions. [Means for solving the problem]
[0005] The present invention is a system for realistically recreating a pet avatar in a virtual space using photos, videos, and social media posting data provided by the user. The system receives photo and video data provided by the user and generates a pet avatar using an image generation means. Furthermore, a multimodal means is used to extract the pet's gestures and cries from the video data and reflect them in the avatar, enabling a realistic reproduction. Additionally, a large-scale language model is used as a conversation generation means, and a chatbot adjusted based on the user's social media posting data enables the pet avatar to have natural conversations with the user.
[0006] In addition, by including avatar display means, nurturing means, and interaction means in AR and VR environments, users will be able to interact with pet avatars in multiple dimensions. This system will provide emotional healing for owners suffering from pet loss and those who find it difficult to keep pets, and will realize a rich communication experience.
[0007] "User" refers to a person who uses the system of the present invention.
[0008] "Photo" refers to still image data provided by the user for avatar generation.
[0009] "Video" refers to video data provided by a user to reproduce the movements and sounds of a pet.
[0010] "SNS posting data" refers to historical data of text and messages posted by users on social networking services.
[0011] "Image generation means" refers to a technology that generates a pet avatar based on a photo provided by the user.
[0012] "Multimodal means" refers to technology that extracts a pet's gestures and cries from video data and reflects them in an avatar.
[0013] "Conversation generation means" refers to technology that uses a large-scale language model to analyze users' SNS posting data and engage in natural conversations with the users.
[0014] "Avatar" refers to a virtual 3D model of a pet generated based on data provided by the user.
[0015] "Training means" refers to the functionality that allows users to interact with avatars within the app.
[0016] "AR environment" refers to an environment that uses augmented reality technology to overlay virtual objects onto real space.
[0017] "VR environment" refers to an environment that uses virtual reality technology to provide an immersive virtual space. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This system uses photos, videos, and social media posts provided by users to realistically recreate pet avatars in a virtual space. This system consists of three major components: a server, a device, and users.
[0040] Server-side processing
[0041] 1. Receipt and storage of data
[0042] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[0043] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[0044] 2. Image Generation and Multimodal Analysis
[0045] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[0046] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[0047] 3. Adjusting the ChatGPT model
[0048] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[0049] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[0050] 4. Submission of Avatar Data
[0051] The server transmits the data of the generated pet avatar to the device, which enables the device to display the avatar.
[0052] Example: The server sends a 3D avatar, gesture data, sound data, and a conversation script to the device.
[0053] Terminal side processing
[0054] 1. Displaying Avatars
[0055] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[0056] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[0057] 2. Providing training functions
[0058] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[0059] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[0060] 3. Providing AR and VR functionality
[0061] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[0062] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[0063] User Behavior
[0064] 1. Upload your data
[0065] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[0066] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[0067] 2. Interacting with avatars
[0068] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[0069] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[0070] 3. Using AR / VR mode
[0071] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[0072] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[0073] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[0077] Step 2:
[0078] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[0079] Step 3:
[0080] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[0081] Step 4:
[0082] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[0083] Step 5:
[0084] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[0085] Step 6:
[0086] The server sends the generated pet avatar data (3D model, movement, voice, conversation script) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[0087] Step 7:
[0088] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[0089] Step 8:
[0090] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[0091] Step 9:
[0092] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[0093] Step 10:
[0094] When users talk to their pet avatars in the app, the avatars respond naturally using ChatGPT, which converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[0095] Step 11:
[0096] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[0097] Example 1
[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0099] In recent years, the number of people who are unable to keep pets has increased. This has led to a growing demand for systems that allow users to interact with pets in virtual spaces. However, existing systems struggle to generate realistic 3D avatars using photos and videos of pets, to reflect the pet's movements and sounds, and to enable natural conversations with the user. Furthermore, there is a lack of systems that can effectively interact with real-world spaces in AR and VR environments. There is a need for a system that can address these technical challenges and provide a more realistic and intimate virtual pet experience.
[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0101] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, and multimodal means for extracting the pet's movements and cries from the video data and reflecting them in the avatar. This realizes a system that generates a realistic 3D avatar of the pet based on the data provided by the user, enabling natural interaction with the user.
[0102] "Photo data" is still image data provided by the user.
[0103] "Video data" refers to data of moving images provided by the user.
[0104] "Image generation means" refers to the technology or method for generating a 3D avatar of a pet in a virtual space based on photographic data.
[0105] "Multimodal means" refers to technologies and methods for extracting pet gestures and sounds from video data and reflecting these characteristics in an avatar.
[0106] "Conversation generation means" refers to techniques and methods for analyzing users' SNS posting data and engaging in natural conversations with users.
[0107] "SNS posting data" refers to information such as text, images, and videos uploaded by users to social networking services.
[0108] A "generative AI model" is a model that is tuned to perform a specific task using AI techniques such as large-scale language models or image generation models.
[0109] A "terminal" is a hardware device that allows a user to run applications and receive and display data sent from a server.
[0110] An "augmented reality environment" is a technology or system that displays computer-generated information overlaid on the real world.
[0111] A "virtual reality environment" is a technology or system that allows users to have an immersive experience in a computer-generated virtual space.
[0112] A "3D avatar" is a three-dimensional model of a pet generated based on data provided by the user.
[0113] This invention is a system that realistically recreates a pet avatar in a virtual space based on digital data provided by the user. This system is composed of three major elements: a server, a terminal, and a user.
[0114] Server-side processing
[0115] The server receives photos and video data of pets uploaded by users through the app. This includes direct uploads from devices such as smartphones and tablets. The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data, while motion analysis software (e.g., OpenPose) is used to analyze the video data.
[0116] The server uses 3D modeling software (e.g., Blender) to generate a 3D avatar of your pet based on the photo data. This 3D avatar incorporates features from the photos and videos provided by the user. Additionally, the server extracts your pet's movements and sounds from the video data and integrates these attributes into the 3D avatar.
[0117] The server also analyzes users' social media posting data and adjusts generative AI models such as ChatGPT based on the information obtained. This allows the pet avatar to have natural conversations with the user. Finally, the generated pet avatar data is sent to the device. This includes the 3D avatar model data, movement data, voice data, and the adjusted conversation model.
[0118] Terminal side processing
[0119] The device uses a 3D rendering engine (e.g., Unity or Unreal Engine) to display the pet avatar based on the data received from the server. The user can view this avatar through the app. The device also provides an interface for controlling the pet's care (e.g., feeding, walking, etc.).
[0120] The device also has the ability to provide AR and VR environments. When a user selects AR mode, the device activates the camera and uses an AR kit (e.g., ARCore, ARKit) to overlay a pet avatar in real space. Additionally, by using a VR headset, users can interact with the pet avatar in a virtual reality environment.
[0121] User Behavior
[0122] Users begin by uploading photos, videos, and social media posts of their pet through the app. Specifically, they select photos and videos of their pet from their smartphone's gallery and upload the social media posts as well. This accumulates the basic data that the system uses to generate a pet avatar.
[0123] Once the avatar is created, users can talk to the pet avatar within the app and perform care operations (e.g., feeding, walking, etc.) Users can also enjoy realistic interactions with virtual pets by selecting AR or VR mode.
[0124] Specific examples
[0125] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server analyzes the received photos and generates a 3D avatar of their pet dog. It also extracts the dog's running movements and barks from the video data and reflects them in the 3D avatar. It then analyzes the social media posting data and adjusts ChatGPT based on the dog's frequently used phrases and characteristic behaviors. The model generated in this way allows the pet avatar to converse with the user in a familiar manner. The server then sends the 3D avatar, gesture data, bark data, and conversation script to the device.
[0126] The device uses a 3D rendering engine to display a 3D avatar of the dog based on the data received from the server. When the user clicks the food button in the app, the device displays the avatar eating food and showing a happy expression. When the user selects AR mode, the device activates the camera and overlays the dog's 3D avatar in real space.
[0127] Prompt Sentence Examples
[0128] "Generate a 3D avatar of your pet dog and adjust it so that it can have a natural conversation with the user. Also, provide a function to display it in real space in AR mode."
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] Users upload photos, videos, and social media posts of their pets through the app.
[0132] Specifically, users select photos or videos from their smartphone's gallery and tap the upload button. Users also link their social media accounts in the app's settings screen and allow access to the posted data.
[0133] Input: Pet photos, videos, and SNS posting data
[0134] Output: Photos, videos, and SNS posting data uploaded to the server
[0135] Step 2:
[0136] The server receives the photo and video data provided by the user and stores this data in dedicated storage.
[0137] Specifically, the server receives the uploaded data and stores the photos in a photo folder, the videos in a video folder, and the text data in a text folder.
[0138] Input: User-uploaded photos, videos, and social media posting data
[0139] Output: Data stored in the server storage
[0140] Step 3:
[0141] The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data.
[0142] Specifically, the server analyzes the photo data and detects facial features, body contours, etc.
[0143] Input: Photo data stored in the server storage
[0144] Output: Facial features and body contour information
[0145] Step 4:
[0146] The server uses the analyzed information to generate a 3D avatar of the pet using 3D modeling software (e.g., Blender).
[0147] Specifically, the server uses Blender to create a 3D model of the pet's face and body.
[0148] Input: Facial features and body contour information
[0149] Output: 3D avatar of your pet
[0150] Step 5:
[0151] The server analyzes the video data using motion analysis software (e.g., OpenPose) to extract the pet's movements and sounds.
[0152] Specifically, the server extracts frames from the video and detects the pet's movements and sounds.
[0153] Input: Video data stored in the server storage
[0154] Output: Pet movement patterns and barking data
[0155] Step 6:
[0156] The server integrates the extracted movement patterns and sounds into a 3D avatar.
[0157] In terms of specific movements, the server reflects the movement data and sound data into the 3D model to create a more realistic avatar.
[0158] Input: Pet 3D avatar, movement patterns, and sound data
[0159] Output: 3D avatar of your pet, with movements and sounds reflected
[0160] Step 7:
[0161] The server analyzes the SNS post data and adjusts the generative AI model (e.g., ChatGPT) based on the information obtained.
[0162] Specifically, the server extracts language patterns and behavioral characteristics related to pets from social media data and retrains ChatGPT.
[0163] Input: Social media post data stored in server storage
[0164] Output: Pet-optimized ChatGPT model
[0165] Step 8:
[0166] The server sends the generated pet avatar data to the device, including a 3D model, movement data, voice data, and a conversation script.
[0167] As a specific operation, the server divides the generated data into packets and transmits them to the terminal.
[0168] Input: 3D pet avatar, movement data, voice data, conversation script
[0169] Output: Avatar data sent to the device
[0170] Step 9:
[0171] The device displays the avatar using a 3D rendering engine (e.g., Unity, Unreal Engine) based on the pet avatar data received from the server.
[0172] Specifically, the device loads the received data into a rendering engine and displays the 3D avatar.
[0173] Input: Avatar data received from the server
[0174] Output: 3D avatar of your pet displayed on your device
[0175] Step 10:
[0176] The user performs training operations within the app, and the device displays the corresponding movements on the 3D avatar.
[0177] Specifically, when the user presses the rice button, the device executes the action of the avatar eating rice.
[0178] Input: User interaction
[0179] Output: Avatar movement
[0180] Step 11:
[0181] The device switches to AR mode or VR mode depending on the user's selection.
[0182] Specifically, when a user selects AR mode, the device activates the camera and overlays an avatar onto the real world.
[0183] Input: User mode selection
[0184] Output: Avatar displayed in real space (AR) or avatar operating in virtual space (VR)
[0185] (Application example 1)
[0186] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0187] Today's pet owners want to consider pet products through direct interaction with physical products and services. However, with the spread of online shopping, physical contact and trying on products has become difficult. In particular, pet products require more detailed interaction to confirm size, fit, and usability. Furthermore, services that allow users to record and utilize their pet memories as digital data are not yet widely available. Therefore, there is a need to provide an environment where users can virtually recreate the characteristics and behavior of their pets and use that data to try out and select the best products.
[0188] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0189] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, and means for using the pet avatar to try out products in a virtual store. This allows the user to try out products in the virtual store through an avatar that reflects the characteristics of their pet and select the product that best suits them.
[0190] "Means for receiving photo and video data provided by users" refers to a function that allows users to upload photos and videos to a server via the Internet and receive that data on the server side.
[0191] The "image generation means for generating a pet avatar using photographic data" is a function for recreating the appearance of a pet as a digital 3D model based on photographic data provided by the user.
[0192] "Multimodal means for extracting pet movements and sounds from video data and reflecting them in the avatar" is a function for analyzing the pet's movements and sounds from video data provided by the user and reflecting them in the avatar.
[0193] "A conversation generation means for analyzing user SNS posting data and engaging in natural conversation with the user" is a function for analyzing text data posted by the user on the SNS and generating natural conversation with the user based on that information.
[0194] The "means for using the pet avatar in the virtual store to try out products" is a function that allows a user to try out products such as pet supplies using a pet avatar in a virtual space.
[0195] The "means for displaying an avatar" is a display function that allows the user to visually confirm the generated pet avatar.
[0196] The "fostering means for interaction" is a function that allows the user to interact with the pet avatar through various responses and actions.
[0197] "Means for interacting with the avatar in AR and VR environments" refers to functionality for interacting with a pet avatar in an augmented reality (AR) or virtual reality (VR) environment.
[0198] "Means for displaying a pet avatar in real space using a smartphone, smart glasses, or head-mounted display" refers to a function that uses various devices to display a pet avatar superimposed on real-world scenery, enabling interaction.
[0199] "A conversation generation means adjusted based on user SNS posting data using a large-scale language model" is a function that uses a large-scale natural language processing model to generate conversation content based on information obtained from user SNS posts.
[0200] This invention is a system that uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and uses the avatars to try out products in a virtual store. Below, we will explain an embodiment of this system.
[0201] Server-side processing
[0202] The server includes the following means for receiving data provided by the user and generating and adjusting a pet avatar based on the data:
[0203] Receiving and storing data
[0204] Users upload photos, videos, and social media posts of their pets to the server via their smartphones or other devices. The server receives this data and uses frameworks such as AWS S3 and Django to store it appropriately.
[0205] Image Generation and Multimodal Analysis
[0206] The server uses tools such as TensorFlow, PyTorch, and Blender to generate a 3D avatar of the pet based on the photo data provided by the user. It also extracts the pet's movements and sounds from the video data and applies multimodal analysis to the avatar.
[0207] Adjusting the ChatGPT model
[0208] The server analyzes users' social media posting data and adjusts the ChatGPT model based on the information obtained. This allows the pet avatar to have a natural conversation with the user. The tools used are PyTorch and Huggingface Transformers.
[0209] Product trial in a virtual store
[0210] The server then sends the generated avatar to the terminal as data for trying out products in a virtual store, allowing the user to try out products in the virtual space.
[0211] Terminal side processing
[0212] The terminal side includes the following means for allowing the user to interact with the pet avatar based on the data received from the server.
[0213] Display avatar
[0214] The device displays the pet avatar data received from the server. Users can view the avatar using devices such as smartphones, smart glasses, and head-mounted displays. Specific tools used include Unity and ARKit / ARCore.
[0215] Providing training functions
[0216] The device provides a nurturing element for users to interact with their pet avatar. For example, when a user clicks the food button in the app, the avatar will appear eating food.
[0217] Providing AR and VR functionality
[0218] The device switches to AR / VR mode based on the user's selection, displaying a pet avatar in real space, allowing users to play with their virtual pet in the living room.
[0219] User Behavior
[0220] Users use the system through the following actions:
[0221] Uploading data
[0222] Users use their smartphones or other devices to upload photos, videos, and social media posts of their pets to the server.
[0223] Interacting with avatars
[0224] Users can talk to the avatar and perform training operations to enjoy the avatar's reactions in real time.
[0225] Virtual product trials
[0226] Users can use their pet avatars in the virtual store to try out products, for example, by having their pet's 3D avatar try on a new collar.
[0227] Prompt Sentence Examples
[0228] Below are some examples of prompt sentences.
[0229] "I'll upload photos and videos of my dog and generate a 3D avatar for you in the virtual store."
[0230] "Try having this 3D avatar hold this ball."
[0231] In this way, by implementing the system based on the present invention, users can try out products in a virtual store through their pet avatar, providing a more realistic experience.
[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0233] Step 1:
[0234] A server receives photo and video data from a user.
[0235] Input: Users upload photos and video data of their pets using their smartphones or other devices.
[0236] How it works: The server receives this data via HTTP requests and stores it in AWS S3 or a database.
[0237] Output: Saved photo and video data.
[0238] Step 2:
[0239] The server uses the received photo data to generate a 3D avatar of your pet.
[0240] Input: Saved photo data.
[0241] How it works: The server uses TensorFlow and PyTorch to process images and Blender to generate 3D avatars.
[0242] Output: A 3D avatar of the generated pet.
[0243] Step 3:
[0244] The server extracts the pet's movements and sounds from the video data it receives and reflects them on the avatar.
[0245] Input: Saved video data.
[0246] Motion: The server uses multimodal analysis techniques to analyze the audio and motion data from the video, which is then applied to an avatar in Blender or another 3D modeling tool.
[0247] Output: A 3D avatar of your pet, with movements and sounds.
[0248] Step 4:
[0249] The server analyzes the user's SNS posting data and adjusts the conversation generation model for the pet avatar.
[0250] Input: User's social media posting data.
[0251] How it works: The server uses Huggingface Transformers to train ChatGPT by feeding social media post data into a natural language processing model, allowing the avatar to have a natural conversation with the user.
[0252] Output: A tuned speech generation model.
[0253] Step 5:
[0254] The server sends the data of the generated pet avatar to the terminal.
[0255] Input: 3D pet avatar, gestures, sound data, conversation generation model.
[0256] Operation: The server sends this data to the device using a communication protocol such as HTTP or WebSocket.
[0257] Output: Avatar data sent to the device.
[0258] Step 6:
[0259] The device displays the pet avatar data received from the server.
[0260] Input: Avatar data sent from the server.
[0261] How it works: The device uses Unity and / or ARKit / ARCore to visually display the pet avatar to the user.
[0262] Output: A 3D avatar of the pet displayed on the user's screen.
[0263] Step 7:
[0264] The terminal provides a pet avatar-raising element in response to a user's interaction request.
[0265] Input: User input (e.g. clicking the rice button).
[0266] Action: The device receives the user's input and makes the pet avatar perform the specified action (e.g., eating food).
[0267] Output: Display of the pet avatar's behavior and pet breeding scene.
[0268] Step 8:
[0269] The device displays the pet avatar in an AR / VR environment according to the user's selection.
[0270] Input: The AR / VR mode selected by the user.
[0271] How it works: The device uses its camera and sensors to capture the real world and overlays a pet avatar on top of it.
[0272] Output: A pet avatar displayed in real space.
[0273] Step 9:
[0274] A user tries out products using a pet avatar in a virtual store.
[0275] Input: A user prompt (e.g., "Try letting him hold this ball").
[0276] Operation: The terminal receives instructions from the user and causes the pet avatar to try out the specified product.
[0277] Output: A pet avatar trying out products in a virtual store.
[0278] In this way, users can try out products in a virtual store through their pet avatar, providing a more realistic shopping experience.
[0279] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0280] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[0281] Server-side processing
[0282] 1. Receipt and storage of data
[0283] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[0284] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[0285] 2. Image Generation and Multimodal Analysis
[0286] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[0287] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[0288] 3. Adjusting the ChatGPT model
[0289] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[0290] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[0291] 4. Implementing the Emotion Engine
[0292] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions, and dynamically adjusts the avatar's behavior and conversation content based on the user's emotional state.
[0293] Example: The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. For example, if the user has a sad expression, the pet avatar will be adjusted to send a comforting message.
[0294] 5. Submission of Avatar Data
[0295] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[0296] Example: The server transmits a 3D avatar, gesture data, sound data, conversation scripts, and emotional response patterns to the terminal.
[0297] Terminal side processing
[0298] 1. Displaying Avatars
[0299] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[0300] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[0301] 2. Providing training functions
[0302] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[0303] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[0304] 3. Collaboration with emotion engine
[0305] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[0306] Example: While a user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will provide playful behavior and fun topics.
[0307] 4. Providing AR and VR functionality
[0308] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[0309] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[0310] User Behavior
[0311] 1. Upload your data
[0312] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[0313] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[0314] 2. Interacting with avatars
[0315] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[0316] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[0317] 3. Feeling a response through emotional connection
[0318] The pet avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[0319] Example: If the user shows signs of fatigue, the avatar will offer a kind message such as, "Take it easy today."
[0320] 4. Using AR / VR mode
[0321] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[0322] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[0323] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments. In addition, by combining it with an emotion engine, the avatar can respond and behave in accordance with the user's emotions, providing an even more realistic and moving experience.
[0324] The processing flow will be explained below.
[0325] Step 1:
[0326] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[0327] Step 2:
[0328] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[0329] Step 3:
[0330] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[0331] Step 4:
[0332] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[0333] Step 5:
[0334] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[0335] Step 6:
[0336] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. Specifically, it analyzes the tone and pitch of the voice and changes in facial expressions in real time to identify the user's emotional state.
[0337] Step 7:
[0338] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[0339] Step 8:
[0340] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[0341] Step 9:
[0342] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[0343] Step 10:
[0344] The device uses a camera and microphone to collect the user's voice and facial expression data and transmits it to the emotion engine, which then analyzes the collected voice and video data to determine the user's emotional state in real time.
[0345] Step 11:
[0346] Based on the results from the emotion engine, the device adjusts the behavior and responses of the pet avatar. For example, if the user is sad, the avatar will display comforting behavior and conversations.
[0347] Step 12:
[0348] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[0349] Step 13:
[0350] When users talk to their pet avatars or perform pet-raising operations within the app, the avatars respond naturally using ChatGPT. Specifically, the app converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[0351] Step 14:
[0352] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[0353] Example 2
[0354] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0355] Previous pet avatar systems lacked the realism of generated avatars due to insufficient analysis of user-provided photos and video data. Furthermore, emotion recognition was not performed during user interaction, making it impossible to realize behaviors and responses that correspond to the emotions of individual users. Furthermore, interaction in AR / VR environments was limited, making it difficult to seamlessly integrate reality and virtuality.
[0356] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0357] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, emotion recognition means for analyzing the user's voice and facial expression data and adjusting the avatar's behavior and conversation content based on the user's emotional state, and means for transmitting the generated pet avatar data to the terminal. This allows the server to analyze a variety of user data to generate a realistic and responsive pet avatar, enabling interaction based on the user's emotions.
[0358] The "means for receiving photo and video data provided by the user" refers to a method for uploading photo and video files to a server via the Internet from a terminal owned by the user.
[0359] "Image generation means" refers to the algorithms and technology that analyzes uploaded photo data and generates a 3D model based on it.
[0360] "Multimodal methods" are methods that integrate multiple pieces of information (gestures, sounds, etc.) extracted from video data and reflect them in a 3D avatar.
[0361] "Conversation generation means" is a technology that analyzes users' SNS posting data and generates natural conversations based on the information obtained.
[0362] "Emotion recognition means" refers to the algorithms and techniques utilized to analyze a user's voice and facial expression data and recognize the user's emotional state.
[0363] "Means for transmitting data of the generated pet avatar to the terminal" refers to a communication method for transmitting the 3D model and related data generated by the server to the user's terminal.
[0364] "Means for displaying an avatar" refers to a technique for displaying avatar data received from a server on a user's terminal.
[0365] "Raising means" refers to a method of providing various functions (e.g., feeding, walking, etc.) that allow users to interact with their pet avatar within the app.
[0366] "Means of interacting in AR and VR environments" refers to methods that use augmented reality and virtual reality technologies to combine the real world with the virtual world and allow users to interact with avatars.
[0367] "Means for recording user interactions" refers to technology that records all operations and conversations that a user has with an avatar as a log.
[0368] A "large-scale language model" is an artificial intelligence algorithm that learns from large amounts of text data and enables natural language generation.
[0369] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[0370] Server-side processing
[0371] Receiving and storing data
[0372] The server receives photos, videos, and social media posting data provided by users through the smartphone app. The received data is appropriately saved in the photo folder, video folder, and text folder, respectively.
[0373] Examples:
[0374] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in various folders.
[0375] Image Generation and Multimodal Analysis
[0376] The server analyzes the received photo data using deep learning technology (e.g., TensorFlow, PyTorch) to generate a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the 3D avatar.
[0377] Examples:
[0378] The server analyzes photos of pet dogs uploaded by users and generates a 3D avatar. It also extracts the dog's running movements and barks from video data and reflects them in the avatar.
[0379] Adjusting the ChatGPT model
[0380] The server analyzes users' social media posting data and adjusts a large-scale language model (e.g., OpenAI GPT-4), enabling the pet avatar to have natural conversations with the user.
[0381] Examples:
[0382] The server analyzes users' social media posting data, extracts their frequently used phrases and characteristic behaviors, and adjusts the ChatGPT model, allowing the pet avatar to converse with the user in a familiar way.
[0383] Implementing the Emotion Engine
[0384] The server implements an emotion engine that analyzes the user's emotions using voice recognition and facial expression recognition technologies (e.g., Microsoft Azure Face API), and adjusts the avatar's behavior and conversation content based on the user's emotional state.
[0385] Examples:
[0386] The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. If the user looks sad, the avatar will send a comforting message.
[0387] Sending avatar data
[0388] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[0389] Examples:
[0390] The server transmits the 3D avatar, gesture data, sound data, conversation script, and emotional response patterns to the terminal.
[0391] Terminal side processing
[0392] Display avatar
[0393] The device displays the pet avatar data received from the server, and this avatar is shown to the user through the app.
[0394] Examples:
[0395] When a user launches the app, the device displays a 3D avatar of their beloved dog based on the received data.
[0396] Providing training functions
[0397] The terminal provides elements for raising the pet avatar (e.g., feeding it, taking it for a walk) that allow the user to interact with the pet avatar.
[0398] Examples:
[0399] When a user clicks the meal button in the app, the device displays an avatar eating the meal and showing a happy expression.
[0400] Collaboration with emotion engine
[0401] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[0402] Examples:
[0403] While the user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will display playful behavior and provide fun topics.
[0404] Providing AR and VR functionality
[0405] The device switches to AR (augmented reality) / VR (virtual reality) mode depending on the user's selection, and displays the pet avatar in real space.
[0406] Examples:
[0407] When a user selects AR mode, the device activates the camera and displays a 3D avatar of their beloved dog overlaid on the real world, allowing the user to play with the virtual pet in the living room.
[0408] User Behavior
[0409] Uploading data
[0410] Users can upload photos, videos, and social media posts of their beloved dogs through the app.
[0411] Examples:
[0412] Users select photos and videos of their pet dog and upload them to the server through the app. Past social media posting data is also sent to the server.
[0413] Interacting with avatars
[0414] Users can talk to their pet avatars and control their care within the app, and the avatars respond in real time.
[0415] Examples:
[0416] Users can talk to their pet dog using the chat function within the app, and the pet dog avatar will respond in natural conversation using ChatGPT.
[0417] Feeling a response through emotional connection
[0418] The avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[0419] Examples:
[0420] If the user shows signs of fatigue, the avatar will offer kind words such as, "Take it easy and rest today."
[0421] Using AR / VR mode
[0422] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[0423] Examples:
[0424] Users can activate the AR mode and have their pet's avatar displayed in the actual room, making it feel as if the pet is actually there.
[0425] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0426] Step 1:
[0427] Users upload photo and video data
[0428] Specific operation: The user selects photos and videos of their pet dog from the smartphone gallery and clicks the "Upload" button. This action sends the selected data to the server.
[0429] Input: User-selected photo and video data
[0430] Output: Photo and video data sent to the server
[0431] Step 2:
[0432] The server receives and stores the photo and video data.
[0433] Specific operation: The server classifies and saves the received data in the appropriate folder. Photos are saved in the "Photos folder" and videos in the "Videos folder."
[0434] Input: Photo and video data submitted by the user
[0435] Output: Photo and video data stored on the server
[0436] Step 3:
[0437] The server analyzes the photo data and generates a 3D avatar
[0438] Specific operation: The server uses deep learning techniques (e.g., TensorFlow, PyTorch) to analyze the photo data and generate a 3D model based on the pet's appearance.
[0439] Input: Photo data stored on the server
[0440] Output: Generated 3D avatar
[0441] Step 4:
[0442] The server analyzes the video data and extracts behaviors and sounds.
[0443] Specific movements: The server analyzes the video frame by frame to extract the pet's movement patterns and sounds, which then adds realistic movements and sounds to the avatar.
[0444] Input: Video data stored on the server
[0445] Output: Extracted movement patterns and call data
[0446] Step 5:
[0447] The server analyzes the user's SNS posting data and adjusts ChatGPT
[0448] How it works: The server analyzes social media posting data, extracts the user's linguistic expressions and preferences, and adjusts the ChatGPT model, allowing the avatar to engage in natural and personalized conversations.
[0449] Input: User's SNS post data
[0450] Output: Adjusted ChatGPT model
[0451] Step 6:
[0452] The server implements the emotion recognition engine.
[0453] Specific operation: The server uses voice recognition technology and facial expression recognition technology (e.g., Microsoft Azure Face API) to analyze the user's voice and facial expression data to recognize emotions. The avatar's behavior and conversation content are adjusted based on the emotions.
[0454] Input: User's voice and facial expression data
[0455] Output: Avatar response pattern based on the user's emotional state
[0456] Step 7:
[0457] The server sends the generated avatar data to the device.
[0458] Specific operation: The server sends the 3D avatar, movement data, sound data, conversation script, and emotional response patterns to the device. The user's device receives these and uses them within the app.
[0459] Input: Generated avatar data (3D model, movement data, sounds, conversation script, emotional response patterns)
[0460] Output: Avatar data sent to the device
[0461] Step 8:
[0462] The device displays an avatar
[0463] Specific operation: When a user launches the app, the device displays a 3D avatar based on the data received from the server.
[0464] Input: Avatar data sent from the server
[0465] Output: 3D avatar displayed on the device
[0466] Step 9:
[0467] The device provides training functions
[0468] Specific operation: The device provides functions for raising pets, such as feeding and walking, so that users can interact with their pet avatars.
[0469] Input: User interaction operations
[0470] Output: Real-time avatar response (e.g. eating behavior)
[0471] Step 10:
[0472] The device will work with the emotion engine
[0473] Specific operation: The device transmits the user's voice and facial expression data to the server's emotion recognition engine in real time, and adjusts the avatar's behavior and conversation.
[0474] Input: Real-time collected voice and facial expression data
[0475] Output: Adjusted avatar behavior and dialogue
[0476] Step 11:
[0477] The device offers AR and VR capabilities
[0478] Specific operation: When a user selects AR / VR mode within the app, the device activates the corresponding hardware (e.g., camera, VR goggles) and displays the pet avatar in real or virtual space.
[0479] Input: User mode selection operation
[0480] Output: Avatar displayed in real or virtual space
[0481] (Application example 2)
[0482] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0483] Although systems that realistically recreate pet avatars in virtual spaces have existed, they have struggled to recognize the user's emotions and provide interactions that respond to the user's real-time mood. Furthermore, previous systems have limited the functionality that users can use to interact with their pet avatars, making the experience unrealistic and restrictive, especially in augmented reality (AR) and virtual reality (VR) environments. Furthermore, the pet avatars' conversational capabilities are limited, and systems that can provide more natural conversations based on the user's social networking service (SNS) posting data have yet to be fully realized.
[0484] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving image data, video data, and data posted to a social networking service provided by a user; image generation means for generating an animal avatar using the image data; multimodal means for extracting animal behavior and sound from the video data and reflecting them in the avatar; dialogue generation means for analyzing the user's social networking service post data and engaging in natural dialogue with the user; and emotion recognition means for analyzing the user's emotions in real time. This allows the avatar's behavior and dialogue content to be dynamically adjusted based on the user's emotional state, enabling more natural and realistic interactions. Furthermore, it is possible to enhance interaction with the avatar in augmented reality and virtual reality environments, improving the user experience.
[0485] "Image data" is digital data that contains visual information, such as photographs and illustrations.
[0486] "Moving image data" refers to digital data that includes visual information that changes over time, and includes moving images and animations.
[0487] "Social networking service posted data" refers to digital data including text, images, videos, and other content posted by users to social networking services.
[0488] An "animal avatar" is an animal character in a virtual space that is generated based on digital data provided by the user.
[0489] "Image generation means" means a technique or method for generating a visual representation, such as a 3D model, based on provided image data.
[0490] A "multimodal means" is a technique or method that integrates different types of data (e.g., images, audio, text) for analysis and processing.
[0491] The "dialogue generation means" is a technique or method for generating conversation content based on provided data in order to realize natural dialogue with the user.
[0492] "Emotion recognition means" refers to a technique or method for analyzing the user's voice and facial expressions in real time and recognizing the user's emotional state.
[0493] An "augmented reality environment" is a technology or method that overlays virtual visual information onto real visual information.
[0494] A "virtual reality environment" is a technology or method that provides a user with an immersive virtual space by generating completely virtual visual and audio information.
[0495] "Dynamic adjustment" means changing the system's operation or behavior in real time according to the situation or conditions.
[0496] The present invention is a system that generates realistic animal avatars in a virtual space using image data, video data, and data posted on social networking services provided by users, and further recognizes the user's emotions in real time and dynamically adjusts interactions with the avatars. Specific program processing and the hardware and software used are described below.
[0497] Server-side processing
[0498] 1. Receipt and storage of data
[0499] The server receives image data, video data, and data posted to social networking services that users have uploaded through the application, and stores this data in the appropriate folders. This process uses folders for images, video, and text data.
[0500] 2. Image Generation and Multimodal Analysis
[0501] The server analyzes the received image data and generates an animal avatar using image generation technology. It also extracts the animal's behavior and sounds from the video data and reflects them in the avatar. Specifically, the image generation technology uses 3D modeling software and AI-based image generation algorithms.
[0502] 3. Dialogue Generation Method
[0503] The server analyzes data posted on social networking services and generates conversations based on this data using large-scale language models (e.g., GPT-3 or GPT-4). This conversation generation model is adjusted based on the user's past posting data, enabling realistic and natural conversations.
[0504] 4. Emotion recognition means
[0505] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. The emotion engine uses Google's Emotion API and Microsoft's AI for Emotional Analysis. The avatar's behavior and dialogue content are dynamically adjusted according to the user's emotional state.
[0506] 5. Submission of Avatar Data
[0507] The server transmits the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the user's device.
[0508] Terminal side processing
[0509] 1. Displaying Avatars
[0510] The device displays the animal avatar data received from the server, and can be a smartphone with a high-performance graphics processor, smart glasses, or a head-mounted display.
[0511] 2. Providing training functions
[0512] The device provides a nurturing function that allows users to interact with the animal avatar, including the ability to virtually feed it and provide simple training.
[0513] 3. Collaboration with emotion engine
[0514] The device sends the user's voice and facial expression data to an emotion engine in real time, which adjusts the avatar's movements and speech based on the user's emotional state.
[0515] 4. Providing AR and VR functionality
[0516] The device switches to AR / VR mode based on the user's selection and displays an animal avatar in real space, allowing users to play with a virtual pet in a real space such as their living room.
[0517] Specific examples
[0518] As a concrete example, when a user launches a virtual pet shop app, a login screen is displayed. After logging in, the user enters the virtual shop, and the camera on their smartphone or smart glasses captures the surrounding environment to create a virtual pet shop interior. When the user approaches a pet avatar, detailed information about that pet is displayed. The detailed information reflects the actual characteristics of the pet using photos, videos, and data posted on social networking services uploaded by the user.
[0519] Prompt Sentence Examples
[0520] "I'm creating a virtual pet shop app. I want to create an app that allows users to interact with multiple virtual pets. I want the app to detect when the user is smiling or sad and change the avatar's reaction accordingly. I also want to display detailed information (photos, videos, social media posts) about the pet the user has selected. How can I do this?"
[0521] The above is a specific embodiment for carrying out the present invention, which allows users to enjoy dynamic interactions according to their emotional state.
[0522] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0523] Step 1:
[0524] Image data, video data, and data posted to social networking services provided by the user are uploaded to the server via the application. This involves the user selecting data from their smartphone or PC and pressing the send button. The input is the various data provided by the user, and the output is the state in which that data is received on the server side.
[0525] Step 2:
[0526] The server stores the received image data, video data, and data posted to social networking services in the appropriate folders. Specifically, the data is classified and stored in folders for images, video, and text data. The input is the uploaded raw data, and the output is the data saved in each folder.
[0527] Step 3:
[0528] The server analyzes the image data and generates animal avatars using image generation technology, utilizing AI-based image generation algorithms and 3D modeling software. The input is the image data, and the output is the generated 3D model of the animal avatar.
[0529] Step 4:
[0530] The server extracts the animal's behavior and sound from the video data and reflects it in the avatar. Specifically, it uses video analysis technology to analyze the pet's movements and cries, and incorporates these behaviors and sounds into the animal avatar. The input is video data, and the output is an animal avatar that reflects the behavior and sound.
[0531] Step 5:
[0532] The server analyzes data posted on social networking services and adjusts a dialogue generation model using a large-scale language model (such as GPT-3 or GPT-4). This model enables natural dialogue with users. The input is text data from the social networking service, and the output is the adjusted dialogue generation model.
[0533] Step 6:
[0534] The server analyzes the user's voice data and facial expression data and executes emotion recognition using tools such as Google's Emotion API and Microsoft's AI for Emotional Analysis. The input is voice data and facial expression data, and the output is the user's emotional state.
[0535] Step 7:
[0536] The server sends the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the device. The input is a series of avatar data generated on the server side, and the output is the avatar data received on the device side.
[0537] Step 8:
[0538] The device displays the received animal avatar data and allows the user to interact with the avatar. For example, the avatar appears on the screen and performs actions in response to user commands. The input is the avatar data received from the server, and the output is an interactive avatar displayed on the device screen.
[0539] Step 9:
[0540] The terminal sends the user's voice and facial expression data to the emotion recognition means in real time, and dynamically adjusts the avatar's behavior and dialogue based on the analysis results. The input is the user's emotion data collected in real time, and the output is the dynamically adjusted avatar's behavior and dialogue content.
[0541] Step 10:
[0542] The device switches to AR / VR mode according to the user's selection and displays an animal avatar in the real space, allowing the user to play with a virtual pet in a real room or environment. The input is the user's mode selection and camera data, and the output is the animal avatar displayed in the AR / VR environment.
[0543] The above are the specific processing steps for carrying out the present invention.
[0544] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0545] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0546] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0547] [Second embodiment]
[0548] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0549] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0550] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0551] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0552] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0553] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0554] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0555] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0556] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0557] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0558] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0559] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0560] This system uses photos, videos, and social media posts provided by users to realistically recreate pet avatars in a virtual space. This system consists of three major components: a server, a device, and users.
[0561] Server-side processing
[0562] 1. Receipt and storage of data
[0563] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[0564] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[0565] 2. Image Generation and Multimodal Analysis
[0566] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[0567] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[0568] 3. Adjusting the ChatGPT model
[0569] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[0570] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[0571] 4. Submission of Avatar Data
[0572] The server transmits the data of the generated pet avatar to the device, which enables the device to display the avatar.
[0573] Example: The server sends a 3D avatar, gesture data, sound data, and a conversation script to the device.
[0574] Terminal side processing
[0575] 1. Displaying Avatars
[0576] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[0577] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[0578] 2. Providing training functions
[0579] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[0580] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[0581] 3. Providing AR and VR functionality
[0582] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[0583] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[0584] User Behavior
[0585] 1. Upload your data
[0586] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[0587] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[0588] 2. Interacting with avatars
[0589] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[0590] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[0591] 3. Using AR / VR mode
[0592] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[0593] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[0594] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments.
[0595] The processing flow will be explained below.
[0596] Step 1:
[0597] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[0598] Step 2:
[0599] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[0600] Step 3:
[0601] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[0602] Step 4:
[0603] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[0604] Step 5:
[0605] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[0606] Step 6:
[0607] The server sends the generated pet avatar data (3D model, movement, voice, conversation script) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[0608] Step 7:
[0609] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[0610] Step 8:
[0611] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[0612] Step 9:
[0613] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[0614] Step 10:
[0615] When users talk to their pet avatars in the app, the avatars respond naturally using ChatGPT, which converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[0616] Step 11:
[0617] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[0618] Example 1
[0619] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0620] In recent years, the number of people who are unable to keep pets has increased. This has led to a growing demand for systems that allow users to interact with pets in virtual spaces. However, existing systems struggle to generate realistic 3D avatars using photos and videos of pets, to reflect the pet's movements and sounds, and to enable natural conversations with the user. Furthermore, there is a lack of systems that can effectively interact with real-world spaces in AR and VR environments. There is a need for a system that can address these technical challenges and provide a more realistic and intimate virtual pet experience.
[0621] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0622] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, and multimodal means for extracting the pet's movements and cries from the video data and reflecting them in the avatar. This realizes a system that generates a realistic 3D avatar of the pet based on the data provided by the user, enabling natural interaction with the user.
[0623] "Photo data" is still image data provided by the user.
[0624] "Video data" refers to data of moving images provided by the user.
[0625] "Image generation means" refers to the technology or method for generating a 3D avatar of a pet in a virtual space based on photographic data.
[0626] "Multimodal means" refers to technologies and methods for extracting pet gestures and sounds from video data and reflecting these characteristics in an avatar.
[0627] "Conversation generation means" refers to techniques and methods for analyzing users' SNS posting data and engaging in natural conversations with users.
[0628] "SNS posting data" refers to information such as text, images, and videos uploaded by users to social networking services.
[0629] A "generative AI model" is a model that is tuned to perform a specific task using AI techniques such as large-scale language models or image generation models.
[0630] A "terminal" is a hardware device that allows a user to run applications and receive and display data sent from a server.
[0631] An "augmented reality environment" is a technology or system that displays computer-generated information overlaid on the real world.
[0632] A "virtual reality environment" is a technology or system that allows users to have an immersive experience in a computer-generated virtual space.
[0633] A "3D avatar" is a three-dimensional model of a pet generated based on data provided by the user.
[0634] This invention is a system that realistically recreates a pet avatar in a virtual space based on digital data provided by the user. This system is composed of three major elements: a server, a terminal, and a user.
[0635] Server-side processing
[0636] The server receives photos and video data of pets uploaded by users through the app. This includes direct uploads from devices such as smartphones and tablets. The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data, while motion analysis software (e.g., OpenPose) is used to analyze the video data.
[0637] The server uses 3D modeling software (e.g., Blender) to generate a 3D avatar of your pet based on the photo data. This 3D avatar incorporates features from the photos and videos provided by the user. Additionally, the server extracts your pet's movements and sounds from the video data and integrates these attributes into the 3D avatar.
[0638] The server also analyzes users' social media posting data and adjusts generative AI models such as ChatGPT based on the information obtained. This allows the pet avatar to have natural conversations with the user. Finally, the generated pet avatar data is sent to the device. This includes the 3D avatar model data, movement data, voice data, and the adjusted conversation model.
[0639] Terminal side processing
[0640] The device uses a 3D rendering engine (e.g., Unity or Unreal Engine) to display the pet avatar based on the data received from the server. The user can view this avatar through the app. The device also provides an interface for controlling the pet's care (e.g., feeding, walking, etc.).
[0641] The device also has the ability to provide AR and VR environments. When a user selects AR mode, the device activates the camera and uses an AR kit (e.g., ARCore, ARKit) to overlay a pet avatar in real space. Additionally, by using a VR headset, users can interact with the pet avatar in a virtual reality environment.
[0642] User Behavior
[0643] Users begin by uploading photos, videos, and social media posts of their pet through the app. Specifically, they select photos and videos of their pet from their smartphone's gallery and upload the social media posts as well. This accumulates the basic data that the system uses to generate a pet avatar.
[0644] Once the avatar is created, users can talk to the pet avatar within the app and perform care operations (e.g., feeding, walking, etc.) Users can also enjoy realistic interactions with virtual pets by selecting AR or VR mode.
[0645] Specific examples
[0646] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server analyzes the received photos and generates a 3D avatar of their pet dog. It also extracts the dog's running movements and barks from the video data and reflects them in the 3D avatar. It then analyzes the social media posting data and adjusts ChatGPT based on the dog's frequently used phrases and characteristic behaviors. The model generated in this way allows the pet avatar to converse with the user in a familiar manner. The server then sends the 3D avatar, gesture data, bark data, and conversation script to the device.
[0647] The device uses a 3D rendering engine to display a 3D avatar of the dog based on the data received from the server. When the user clicks the food button in the app, the device displays the avatar eating food and showing a happy expression. When the user selects AR mode, the device activates the camera and overlays the dog's 3D avatar in real space.
[0648] Prompt Sentence Examples
[0649] "Generate a 3D avatar of your pet dog and adjust it so that it can have a natural conversation with the user. Also, provide a function to display it in real space in AR mode."
[0650] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0651] Step 1:
[0652] Users upload photos, videos, and social media posts of their pets through the app.
[0653] Specifically, users select photos or videos from their smartphone's gallery and tap the upload button. Users also link their social media accounts in the app's settings screen and allow access to the posted data.
[0654] Input: Pet photos, videos, and SNS posting data
[0655] Output: Photos, videos, and SNS posting data uploaded to the server
[0656] Step 2:
[0657] The server receives the photo and video data provided by the user and stores this data in dedicated storage.
[0658] Specifically, the server receives the uploaded data and stores the photos in a photo folder, the videos in a video folder, and the text data in a text folder.
[0659] Input: User-uploaded photos, videos, and social media posting data
[0660] Output: Data stored in the server storage
[0661] Step 3:
[0662] The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data.
[0663] Specifically, the server analyzes the photo data and detects facial features, body contours, etc.
[0664] Input: Photo data stored in the server storage
[0665] Output: Facial features and body contour information
[0666] Step 4:
[0667] The server uses the analyzed information to generate a 3D avatar of the pet using 3D modeling software (e.g., Blender).
[0668] Specifically, the server uses Blender to create a 3D model of the pet's face and body.
[0669] Input: Facial features and body contour information
[0670] Output: 3D avatar of your pet
[0671] Step 5:
[0672] The server analyzes the video data using motion analysis software (e.g., OpenPose) to extract the pet's movements and sounds.
[0673] Specifically, the server extracts frames from the video and detects the pet's movements and sounds.
[0674] Input: Video data stored in the server storage
[0675] Output: Pet movement patterns and barking data
[0676] Step 6:
[0677] The server integrates the extracted movement patterns and sounds into a 3D avatar.
[0678] In terms of specific movements, the server reflects the movement data and sound data into the 3D model to create a more realistic avatar.
[0679] Input: Pet 3D avatar, movement patterns, and sound data
[0680] Output: 3D avatar of your pet, with movements and sounds reflected
[0681] Step 7:
[0682] The server analyzes the SNS post data and adjusts the generative AI model (e.g., ChatGPT) based on the information obtained.
[0683] Specifically, the server extracts language patterns and behavioral characteristics related to pets from social media data and retrains ChatGPT.
[0684] Input: Social media post data stored in server storage
[0685] Output: Pet-optimized ChatGPT model
[0686] Step 8:
[0687] The server sends the generated pet avatar data to the device, including a 3D model, movement data, voice data, and a conversation script.
[0688] As a specific operation, the server divides the generated data into packets and transmits them to the terminal.
[0689] Input: 3D pet avatar, movement data, voice data, conversation script
[0690] Output: Avatar data sent to the device
[0691] Step 9:
[0692] The device displays the avatar using a 3D rendering engine (e.g., Unity, Unreal Engine) based on the pet avatar data received from the server.
[0693] Specifically, the device loads the received data into a rendering engine and displays the 3D avatar.
[0694] Input: Avatar data received from the server
[0695] Output: 3D avatar of your pet displayed on your device
[0696] Step 10:
[0697] The user performs training operations within the app, and the device displays the corresponding movements on the 3D avatar.
[0698] Specifically, when the user presses the rice button, the device executes the action of the avatar eating rice.
[0699] Input: User interaction
[0700] Output: Avatar movement
[0701] Step 11:
[0702] The device switches to AR mode or VR mode depending on the user's selection.
[0703] Specifically, when a user selects AR mode, the device activates the camera and overlays an avatar onto the real world.
[0704] Input: User mode selection
[0705] Output: Avatar displayed in real space (AR) or avatar operating in virtual space (VR)
[0706] (Application example 1)
[0707] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0708] Today's pet owners want to consider pet products through direct interaction with physical products and services. However, with the spread of online shopping, physical contact and trying on products has become difficult. In particular, pet products require more detailed interaction to confirm size, fit, and usability. Furthermore, services that allow users to record and utilize their pet memories as digital data are not yet widely available. Therefore, there is a need to provide an environment where users can virtually recreate the characteristics and behavior of their pets and use that data to try out and select the best products.
[0709] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0710] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, and means for using the pet avatar to try out products in a virtual store. This allows the user to try out products in the virtual store through an avatar that reflects the characteristics of their pet and select the product that best suits them.
[0711] "Means for receiving photo and video data provided by users" refers to a function that allows users to upload photos and videos to a server via the Internet and receive that data on the server side.
[0712] The "image generation means for generating a pet avatar using photographic data" is a function for recreating the appearance of a pet as a digital 3D model based on photographic data provided by the user.
[0713] "Multimodal means for extracting pet movements and sounds from video data and reflecting them in the avatar" is a function for analyzing the pet's movements and sounds from video data provided by the user and reflecting them in the avatar.
[0714] "A conversation generation means for analyzing user SNS posting data and engaging in natural conversation with the user" is a function for analyzing text data posted by the user on the SNS and generating natural conversation with the user based on that information.
[0715] The "means for using the pet avatar in the virtual store to try out products" is a function that allows a user to try out products such as pet supplies using a pet avatar in a virtual space.
[0716] The "means for displaying an avatar" is a display function that allows the user to visually confirm the generated pet avatar.
[0717] The "fostering means for interaction" is a function that allows the user to interact with the pet avatar through various responses and actions.
[0718] "Means for interacting with the avatar in AR and VR environments" refers to functionality for interacting with a pet avatar in an augmented reality (AR) or virtual reality (VR) environment.
[0719] "Means for displaying a pet avatar in real space using a smartphone, smart glasses, or head-mounted display" refers to a function that uses various devices to display a pet avatar superimposed on real-world scenery, enabling interaction.
[0720] "A conversation generation means adjusted based on user SNS posting data using a large-scale language model" is a function that uses a large-scale natural language processing model to generate conversation content based on information obtained from user SNS posts.
[0721] This invention is a system that uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and uses the avatars to try out products in a virtual store. Below, we will explain an embodiment of this system.
[0722] Server-side processing
[0723] The server includes the following means for receiving data provided by the user and generating and adjusting a pet avatar based on the data:
[0724] Receiving and storing data
[0725] Users upload photos, videos, and social media posts of their pets to the server via their smartphones or other devices. The server receives this data and uses frameworks such as AWS S3 and Django to store it appropriately.
[0726] Image Generation and Multimodal Analysis
[0727] The server uses tools such as TensorFlow, PyTorch, and Blender to generate a 3D avatar of the pet based on the photo data provided by the user. It also extracts the pet's movements and sounds from the video data and applies multimodal analysis to the avatar.
[0728] Adjusting the ChatGPT model
[0729] The server analyzes users' social media posting data and adjusts the ChatGPT model based on the information obtained. This allows the pet avatar to have a natural conversation with the user. The tools used are PyTorch and Huggingface Transformers.
[0730] Product trial in a virtual store
[0731] The server then sends the generated avatar to the terminal as data for trying out products in a virtual store, allowing the user to try out products in the virtual space.
[0732] Terminal side processing
[0733] The terminal side includes the following means for allowing the user to interact with the pet avatar based on the data received from the server.
[0734] Display avatar
[0735] The device displays the pet avatar data received from the server. Users can view the avatar using devices such as smartphones, smart glasses, and head-mounted displays. Specific tools used include Unity and ARKit / ARCore.
[0736] Providing training functions
[0737] The device provides a nurturing element for users to interact with their pet avatar. For example, when a user clicks the food button in the app, the avatar will appear eating food.
[0738] Providing AR and VR functionality
[0739] The device switches to AR / VR mode based on the user's selection, displaying a pet avatar in real space, allowing users to play with their virtual pet in the living room.
[0740] User Behavior
[0741] Users use the system through the following actions:
[0742] Uploading data
[0743] Users use their smartphones or other devices to upload photos, videos, and social media posts of their pets to the server.
[0744] Interacting with avatars
[0745] Users can talk to the avatar and perform training operations to enjoy the avatar's reactions in real time.
[0746] Virtual product trials
[0747] Users can use their pet avatars in the virtual store to try out products, for example, by having their pet's 3D avatar try on a new collar.
[0748] Prompt Sentence Examples
[0749] Below are some examples of prompt sentences.
[0750] "I'll upload photos and videos of my dog and generate a 3D avatar for you in the virtual store."
[0751] "Try having this 3D avatar hold this ball."
[0752] In this way, by implementing the system based on the present invention, users can try out products in a virtual store through their pet avatar, providing a more realistic experience.
[0753] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0754] Step 1:
[0755] A server receives photo and video data from a user.
[0756] Input: Users upload photos and video data of their pets using their smartphones or other devices.
[0757] How it works: The server receives this data via HTTP requests and stores it in AWS S3 or a database.
[0758] Output: Saved photo and video data.
[0759] Step 2:
[0760] The server uses the received photo data to generate a 3D avatar of your pet.
[0761] Input: Saved photo data.
[0762] How it works: The server uses TensorFlow and PyTorch to process images and Blender to generate 3D avatars.
[0763] Output: A 3D avatar of the generated pet.
[0764] Step 3:
[0765] The server extracts the pet's movements and sounds from the video data it receives and reflects them on the avatar.
[0766] Input: Saved video data.
[0767] Motion: The server uses multimodal analysis techniques to analyze the audio and motion data from the video, which is then applied to an avatar in Blender or another 3D modeling tool.
[0768] Output: A 3D avatar of your pet, with movements and sounds.
[0769] Step 4:
[0770] The server analyzes the user's SNS posting data and adjusts the conversation generation model for the pet avatar.
[0771] Input: User's social media posting data.
[0772] How it works: The server uses Huggingface Transformers to train ChatGPT by feeding social media post data into a natural language processing model, allowing the avatar to have a natural conversation with the user.
[0773] Output: A tuned speech generation model.
[0774] Step 5:
[0775] The server sends the data of the generated pet avatar to the terminal.
[0776] Input: 3D pet avatar, gestures, sound data, conversation generation model.
[0777] Operation: The server sends this data to the device using a communication protocol such as HTTP or WebSocket.
[0778] Output: Avatar data sent to the device.
[0779] Step 6:
[0780] The device displays the pet avatar data received from the server.
[0781] Input: Avatar data sent from the server.
[0782] How it works: The device uses Unity and / or ARKit / ARCore to visually display the pet avatar to the user.
[0783] Output: A 3D avatar of the pet displayed on the user's screen.
[0784] Step 7:
[0785] The terminal provides a pet avatar-raising element in response to a user's interaction request.
[0786] Input: User input (e.g. clicking the rice button).
[0787] Action: The device receives the user's input and makes the pet avatar perform the specified action (e.g., eating food).
[0788] Output: Display of the pet avatar's behavior and pet breeding scene.
[0789] Step 8:
[0790] The device displays the pet avatar in an AR / VR environment according to the user's selection.
[0791] Input: The AR / VR mode selected by the user.
[0792] How it works: The device uses its camera and sensors to capture the real world and overlays a pet avatar on top of it.
[0793] Output: A pet avatar displayed in real space.
[0794] Step 9:
[0795] A user tries out products using a pet avatar in a virtual store.
[0796] Input: A user prompt (e.g., "Try letting him hold this ball").
[0797] Operation: The terminal receives instructions from the user and causes the pet avatar to try out the specified product.
[0798] Output: A pet avatar trying out products in a virtual store.
[0799] In this way, users can try out products in a virtual store through their pet avatar, providing a more realistic shopping experience.
[0800] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0801] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[0802] Server-side processing
[0803] 1. Receipt and storage of data
[0804] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[0805] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[0806] 2. Image Generation and Multimodal Analysis
[0807] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[0808] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[0809] 3. Adjusting the ChatGPT model
[0810] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[0811] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[0812] 4. Implementing the Emotion Engine
[0813] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions, and dynamically adjusts the avatar's behavior and conversation content based on the user's emotional state.
[0814] Example: The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. For example, if the user has a sad expression, the pet avatar will be adjusted to send a comforting message.
[0815] 5. Submission of Avatar Data
[0816] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[0817] Example: The server transmits a 3D avatar, gesture data, sound data, conversation scripts, and emotional response patterns to the terminal.
[0818] Terminal side processing
[0819] 1. Displaying Avatars
[0820] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[0821] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[0822] 2. Providing training functions
[0823] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[0824] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[0825] 3. Collaboration with emotion engine
[0826] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[0827] Example: While a user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will provide playful behavior and fun topics.
[0828] 4. Providing AR and VR functionality
[0829] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[0830] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[0831] User Behavior
[0832] 1. Upload your data
[0833] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[0834] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[0835] 2. Interacting with avatars
[0836] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[0837] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[0838] 3. Feeling a response through emotional connection
[0839] The pet avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[0840] Example: If the user shows signs of fatigue, the avatar will offer a kind message such as, "Take it easy today."
[0841] 4. Using AR / VR mode
[0842] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[0843] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[0844] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments. In addition, by combining it with an emotion engine, the avatar can respond and behave in accordance with the user's emotions, providing an even more realistic and moving experience.
[0845] The processing flow will be explained below.
[0846] Step 1:
[0847] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[0848] Step 2:
[0849] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[0850] Step 3:
[0851] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[0852] Step 4:
[0853] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[0854] Step 5:
[0855] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[0856] Step 6:
[0857] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. Specifically, it analyzes the tone and pitch of the voice and changes in facial expressions in real time to identify the user's emotional state.
[0858] Step 7:
[0859] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[0860] Step 8:
[0861] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[0862] Step 9:
[0863] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[0864] Step 10:
[0865] The device uses a camera and microphone to collect the user's voice and facial expression data and transmits it to the emotion engine, which then analyzes the collected voice and video data to determine the user's emotional state in real time.
[0866] Step 11:
[0867] Based on the results from the emotion engine, the device adjusts the behavior and responses of the pet avatar. For example, if the user is sad, the avatar will display comforting behavior and conversations.
[0868] Step 12:
[0869] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[0870] Step 13:
[0871] When users talk to their pet avatars or perform pet-raising operations within the app, the avatars respond naturally using ChatGPT. Specifically, the app converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[0872] Step 14:
[0873] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[0874] Example 2
[0875] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0876] Previous pet avatar systems lacked the realism of generated avatars due to insufficient analysis of user-provided photos and video data. Furthermore, emotion recognition was not performed during user interaction, making it impossible to realize behaviors and responses that correspond to the emotions of individual users. Furthermore, interaction in AR / VR environments was limited, making it difficult to seamlessly integrate reality and virtuality.
[0877] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0878] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, emotion recognition means for analyzing the user's voice and facial expression data and adjusting the avatar's behavior and conversation content based on the user's emotional state, and means for transmitting the generated pet avatar data to the terminal. This allows the server to analyze a variety of user data to generate a realistic and responsive pet avatar, enabling interaction based on the user's emotions.
[0879] The "means for receiving photo and video data provided by the user" refers to a method for uploading photo and video files to a server via the Internet from a terminal owned by the user.
[0880] "Image generation means" refers to the algorithms and technology that analyzes uploaded photo data and generates a 3D model based on it.
[0881] "Multimodal methods" are methods that integrate multiple pieces of information (gestures, sounds, etc.) extracted from video data and reflect them in a 3D avatar.
[0882] "Conversation generation means" is a technology that analyzes users' SNS posting data and generates natural conversations based on the information obtained.
[0883] "Emotion recognition means" refers to the algorithms and techniques utilized to analyze a user's voice and facial expression data and recognize the user's emotional state.
[0884] "Means for transmitting data of the generated pet avatar to the terminal" refers to a communication method for transmitting the 3D model and related data generated by the server to the user's terminal.
[0885] "Means for displaying an avatar" refers to a technique for displaying avatar data received from a server on a user's terminal.
[0886] "Raising means" refers to a method of providing various functions (e.g., feeding, walking, etc.) that allow users to interact with their pet avatar within the app.
[0887] "Means of interacting in AR and VR environments" refers to methods that use augmented reality and virtual reality technologies to combine the real world with the virtual world and allow users to interact with avatars.
[0888] "Means for recording user interactions" refers to technology that records all operations and conversations that a user has with an avatar as a log.
[0889] A "large-scale language model" is an artificial intelligence algorithm that learns from large amounts of text data and enables natural language generation.
[0890] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[0891] Server-side processing
[0892] Receiving and storing data
[0893] The server receives photos, videos, and social media posting data provided by users through the smartphone app. The received data is appropriately saved in the photo folder, video folder, and text folder, respectively.
[0894] Examples:
[0895] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in various folders.
[0896] Image Generation and Multimodal Analysis
[0897] The server analyzes the received photo data using deep learning technology (e.g., TensorFlow, PyTorch) to generate a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the 3D avatar.
[0898] Examples:
[0899] The server analyzes photos of pet dogs uploaded by users and generates a 3D avatar. It also extracts the dog's running movements and barks from video data and reflects them in the avatar.
[0900] Adjusting the ChatGPT model
[0901] The server analyzes users' social media posting data and adjusts a large-scale language model (e.g., OpenAI GPT-4), enabling the pet avatar to have natural conversations with the user.
[0902] Examples:
[0903] The server analyzes users' social media posting data, extracts their frequently used phrases and characteristic behaviors, and adjusts the ChatGPT model, allowing the pet avatar to converse with the user in a familiar way.
[0904] Implementing the Emotion Engine
[0905] The server implements an emotion engine that analyzes the user's emotions using voice recognition and facial expression recognition technologies (e.g., Microsoft Azure Face API), and adjusts the avatar's behavior and conversation content based on the user's emotional state.
[0906] Examples:
[0907] The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. If the user looks sad, the avatar will send a comforting message.
[0908] Sending avatar data
[0909] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[0910] Examples:
[0911] The server transmits the 3D avatar, gesture data, sound data, conversation script, and emotional response patterns to the terminal.
[0912] Terminal side processing
[0913] Display avatar
[0914] The device displays the pet avatar data received from the server, and this avatar is shown to the user through the app.
[0915] Examples:
[0916] When a user launches the app, the device displays a 3D avatar of their beloved dog based on the received data.
[0917] Providing training functions
[0918] The terminal provides elements for raising the pet avatar (e.g., feeding it, taking it for a walk) that allow the user to interact with the pet avatar.
[0919] Examples:
[0920] When a user clicks the meal button in the app, the device displays an avatar eating the meal and showing a happy expression.
[0921] Collaboration with emotion engine
[0922] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[0923] Examples:
[0924] While the user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will display playful behavior and provide fun topics.
[0925] Providing AR and VR functionality
[0926] The device switches to AR (augmented reality) / VR (virtual reality) mode depending on the user's selection, and displays the pet avatar in real space.
[0927] Examples:
[0928] When a user selects AR mode, the device activates the camera and displays a 3D avatar of their beloved dog overlaid on the real world, allowing the user to play with the virtual pet in the living room.
[0929] User Behavior
[0930] Uploading data
[0931] Users can upload photos, videos, and social media posts of their beloved dogs through the app.
[0932] Examples:
[0933] Users select photos and videos of their pet dog and upload them to the server through the app. Past social media posting data is also sent to the server.
[0934] Interacting with avatars
[0935] Users can talk to their pet avatars and control their care within the app, and the avatars respond in real time.
[0936] Examples:
[0937] Users can talk to their pet dog using the chat function within the app, and the pet dog avatar will respond in natural conversation using ChatGPT.
[0938] Feeling a response through emotional connection
[0939] The avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[0940] Examples:
[0941] If the user shows signs of fatigue, the avatar will offer kind words such as, "Take it easy and rest today."
[0942] Using AR / VR mode
[0943] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[0944] Examples:
[0945] Users can activate the AR mode and have their pet's avatar displayed in the actual room, making it feel as if the pet is actually there.
[0946] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0947] Step 1:
[0948] Users upload photo and video data
[0949] Specific operation: The user selects photos and videos of their pet dog from the smartphone gallery and clicks the "Upload" button. This action sends the selected data to the server.
[0950] Input: User-selected photo and video data
[0951] Output: Photo and video data sent to the server
[0952] Step 2:
[0953] The server receives and stores the photo and video data.
[0954] Specific operation: The server classifies and saves the received data in the appropriate folder. Photos are saved in the "Photos folder" and videos in the "Videos folder."
[0955] Input: Photo and video data submitted by the user
[0956] Output: Photo and video data stored on the server
[0957] Step 3:
[0958] The server analyzes the photo data and generates a 3D avatar
[0959] Specific operation: The server uses deep learning techniques (e.g., TensorFlow, PyTorch) to analyze the photo data and generate a 3D model based on the pet's appearance.
[0960] Input: Photo data stored on the server
[0961] Output: Generated 3D avatar
[0962] Step 4:
[0963] The server analyzes the video data and extracts behaviors and sounds.
[0964] Specific movements: The server analyzes the video frame by frame to extract the pet's movement patterns and sounds, which then adds realistic movements and sounds to the avatar.
[0965] Input: Video data stored on the server
[0966] Output: Extracted movement patterns and call data
[0967] Step 5:
[0968] The server analyzes the user's SNS posting data and adjusts ChatGPT
[0969] How it works: The server analyzes social media posting data, extracts the user's linguistic expressions and preferences, and adjusts the ChatGPT model, allowing the avatar to engage in natural and personalized conversations.
[0970] Input: User's SNS post data
[0971] Output: Adjusted ChatGPT model
[0972] Step 6:
[0973] The server implements the emotion recognition engine.
[0974] Specific operation: The server uses voice recognition technology and facial expression recognition technology (e.g., Microsoft Azure Face API) to analyze the user's voice and facial expression data to recognize emotions. The avatar's behavior and conversation content are adjusted based on the emotions.
[0975] Input: User's voice and facial expression data
[0976] Output: Avatar response pattern based on the user's emotional state
[0977] Step 7:
[0978] The server sends the generated avatar data to the device.
[0979] Specific operation: The server sends the 3D avatar, movement data, sound data, conversation script, and emotional response patterns to the device. The user's device receives these and uses them within the app.
[0980] Input: Generated avatar data (3D model, movement data, sounds, conversation script, emotional response patterns)
[0981] Output: Avatar data sent to the device
[0982] Step 8:
[0983] The device displays an avatar
[0984] Specific operation: When a user launches the app, the device displays a 3D avatar based on the data received from the server.
[0985] Input: Avatar data sent from the server
[0986] Output: 3D avatar displayed on the device
[0987] Step 9:
[0988] The device provides training functions
[0989] Specific operation: The device provides functions for raising pets, such as feeding and walking, so that users can interact with their pet avatars.
[0990] Input: User interaction operations
[0991] Output: Real-time avatar response (e.g. eating behavior)
[0992] Step 10:
[0993] The device will work with the emotion engine
[0994] Specific operation: The device transmits the user's voice and facial expression data in real time to the server's emotion recognition engine, and adjusts the avatar's behavior and conversation.
[0995] Input: Real-time collected voice and facial expression data
[0996] Output: Adjusted avatar behavior and dialogue
[0997] Step 11:
[0998] The device offers AR and VR capabilities
[0999] Specific operation: When a user selects AR / VR mode within the app, the device activates the corresponding hardware (e.g., camera, VR goggles) and displays the pet avatar in real or virtual space.
[1000] Input: User mode selection operation
[1001] Output: Avatar displayed in real or virtual space
[1002] (Application example 2)
[1003] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1004] Although systems that realistically recreate pet avatars in virtual spaces have existed, they have struggled to recognize the user's emotions and provide interactions that respond to the user's real-time mood. Furthermore, previous systems have limited the functionality that users can use to interact with their pet avatars, making the experience unrealistic and restrictive, especially in augmented reality (AR) and virtual reality (VR) environments. Furthermore, the pet avatars' conversational capabilities are limited, and systems that can provide more natural conversations based on the user's social networking service (SNS) posting data have yet to be fully realized.
[1005] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving image data, video data, and data posted to a social networking service provided by a user; image generation means for generating an animal avatar using the image data; multimodal means for extracting animal behavior and sound from the video data and reflecting them in the avatar; dialogue generation means for analyzing the user's social networking service post data and engaging in natural dialogue with the user; and emotion recognition means for analyzing the user's emotions in real time. This allows the avatar's behavior and dialogue content to be dynamically adjusted based on the user's emotional state, enabling more natural and realistic interactions. Furthermore, it is possible to enhance interaction with the avatar in augmented reality and virtual reality environments, improving the user experience.
[1006] "Image data" is digital data that contains visual information, such as photographs and illustrations.
[1007] "Moving image data" refers to digital data that includes visual information that changes over time, and includes moving images and animations.
[1008] "Social networking service posted data" refers to digital data including text, images, videos, and other content posted by users to social networking services.
[1009] An "animal avatar" is an animal character in a virtual space that is generated based on digital data provided by the user.
[1010] "Image generation means" means a technique or method for generating a visual representation, such as a 3D model, based on provided image data.
[1011] A "multimodal means" is a technique or method that integrates different types of data (e.g., images, audio, text) for analysis and processing.
[1012] The "dialogue generation means" is a technique or method for generating conversation content based on provided data in order to realize natural dialogue with the user.
[1013] "Emotion recognition means" refers to a technique or method for analyzing the user's voice and facial expressions in real time and recognizing the user's emotional state.
[1014] An "augmented reality environment" is a technology or method that overlays virtual visual information onto real visual information.
[1015] A "virtual reality environment" is a technology or method that provides a user with an immersive virtual space by generating completely virtual visual and audio information.
[1016] "Dynamic adjustment" means changing the system's operation or behavior in real time according to the situation or conditions.
[1017] The present invention is a system that generates realistic animal avatars in a virtual space using image data, video data, and data posted on social networking services provided by users, and further recognizes the user's emotions in real time and dynamically adjusts interactions with the avatars. Specific program processing and the hardware and software used are described below.
[1018] Server-side processing
[1019] 1. Receipt and storage of data
[1020] The server receives image data, video data, and data posted to social networking services that users have uploaded through the application, and stores this data in the appropriate folders. This process uses folders for images, video, and text data.
[1021] 2. Image Generation and Multimodal Analysis
[1022] The server analyzes the received image data and generates an animal avatar using image generation technology. It also extracts the animal's behavior and sounds from the video data and reflects them in the avatar. Specifically, the image generation technology uses 3D modeling software and AI-based image generation algorithms.
[1023] 3. Dialogue Generation Method
[1024] The server analyzes data posted on social networking services and generates conversations based on this data using large-scale language models (e.g., GPT-3 or GPT-4). This conversation generation model is adjusted based on the user's past posting data, enabling realistic and natural conversations.
[1025] 4. Emotion recognition means
[1026] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. The emotion engine uses Google's Emotion API and Microsoft's AI for Emotional Analysis. The avatar's behavior and dialogue content are dynamically adjusted according to the user's emotional state.
[1027] 5. Submission of Avatar Data
[1028] The server transmits the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the user's device.
[1029] Terminal side processing
[1030] 1. Displaying Avatars
[1031] The device displays the animal avatar data received from the server, and can be a smartphone with a high-performance graphics processor, smart glasses, or a head-mounted display.
[1032] 2. Providing training functions
[1033] The device provides a nurturing function that allows users to interact with the animal avatar, including the ability to virtually feed it and provide simple training.
[1034] 3. Collaboration with emotion engine
[1035] The device sends the user's voice and facial expression data to an emotion engine in real time, which adjusts the avatar's movements and speech based on the user's emotional state.
[1036] 4. Providing AR and VR functionality
[1037] The device switches to AR / VR mode based on the user's selection and displays an animal avatar in real space, allowing users to play with a virtual pet in a real space such as their living room.
[1038] Specific examples
[1039] As a concrete example, when a user launches a virtual pet shop app, a login screen is displayed. After logging in, the user enters the virtual shop, and the camera on their smartphone or smart glasses captures the surrounding environment to create a virtual pet shop interior. When the user approaches a pet avatar, detailed information about that pet is displayed. The detailed information reflects the actual characteristics of the pet using photos, videos, and data posted on social networking services uploaded by the user.
[1040] Prompt Sentence Examples
[1041] "I'm creating a virtual pet shop app. I want to create an app that allows users to interact with multiple virtual pets. I want the app to detect when the user is smiling or sad and change the avatar's reaction accordingly. I also want to display detailed information (photos, videos, social media posts) about the pet the user has selected. How can I do this?"
[1042] The above is a specific embodiment for carrying out the present invention, which allows users to enjoy dynamic interactions according to their emotional state.
[1043] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1044] Step 1:
[1045] Image data, video data, and data posted to social networking services provided by the user are uploaded to the server via the application. This involves the user selecting data from their smartphone or PC and pressing the send button. The input is the various data provided by the user, and the output is the state in which that data is received on the server side.
[1046] Step 2:
[1047] The server stores the received image data, video data, and data posted to social networking services in the appropriate folders. Specifically, the data is classified and stored in folders for images, video, and text data. The input is the uploaded raw data, and the output is the data saved in each folder.
[1048] Step 3:
[1049] The server analyzes the image data and generates animal avatars using image generation technology, utilizing AI-based image generation algorithms and 3D modeling software. The input is the image data, and the output is the generated 3D model of the animal avatar.
[1050] Step 4:
[1051] The server extracts the animal's behavior and sound from the video data and reflects it in the avatar. Specifically, it uses video analysis technology to analyze the pet's movements and cries, and incorporates these behaviors and sounds into the animal avatar. The input is video data, and the output is an animal avatar that reflects the behavior and sound.
[1052] Step 5:
[1053] The server analyzes data posted on social networking services and adjusts a dialogue generation model using a large-scale language model (such as GPT-3 or GPT-4). This model enables natural dialogue with users. The input is text data from the social networking service, and the output is the adjusted dialogue generation model.
[1054] Step 6:
[1055] The server analyzes the user's voice data and facial expression data and executes emotion recognition using tools such as Google's Emotion API and Microsoft's AI for Emotional Analysis. The input is voice data and facial expression data, and the output is the user's emotional state.
[1056] Step 7:
[1057] The server sends the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the device. The input is a series of avatar data generated on the server side, and the output is the avatar data received on the device side.
[1058] Step 8:
[1059] The device displays the received animal avatar data and allows the user to interact with the avatar. For example, the avatar appears on the screen and performs actions in response to user commands. The input is the avatar data received from the server, and the output is an interactive avatar displayed on the device screen.
[1060] Step 9:
[1061] The terminal sends the user's voice and facial expression data to the emotion recognition means in real time, and dynamically adjusts the avatar's behavior and dialogue based on the analysis results. The input is the user's emotion data collected in real time, and the output is the dynamically adjusted avatar's behavior and dialogue content.
[1062] Step 10:
[1063] The device switches to AR / VR mode according to the user's selection and displays an animal avatar in the real space, allowing the user to play with a virtual pet in a real room or environment. The input is the user's mode selection and camera data, and the output is the animal avatar displayed in the AR / VR environment.
[1064] The above are the specific processing steps for carrying out the present invention.
[1065] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1066] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1067] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1068] [Third embodiment]
[1069] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1070] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1071] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1072] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1073] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1074] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1075] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1076] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1077] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1078] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1079] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1080] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1081] This system uses photos, videos, and social media posts provided by users to realistically recreate pet avatars in a virtual space. This system consists of three major components: a server, a device, and users.
[1082] Server-side processing
[1083] 1. Receipt and storage of data
[1084] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[1085] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[1086] 2. Image Generation and Multimodal Analysis
[1087] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[1088] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[1089] 3. Adjusting the ChatGPT model
[1090] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[1091] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[1092] 4. Submission of Avatar Data
[1093] The server transmits the data of the generated pet avatar to the device, which enables the device to display the avatar.
[1094] Example: The server sends a 3D avatar, gesture data, sound data, and a conversation script to the device.
[1095] Terminal side processing
[1096] 1. Displaying Avatars
[1097] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[1098] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[1099] 2. Providing training functions
[1100] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[1101] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[1102] 3. Providing AR and VR functionality
[1103] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[1104] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[1105] User Behavior
[1106] 1. Upload your data
[1107] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[1108] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[1109] 2. Interacting with avatars
[1110] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[1111] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[1112] 3. Using AR / VR mode
[1113] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[1114] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[1115] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments.
[1116] The processing flow will be explained below.
[1117] Step 1:
[1118] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[1119] Step 2:
[1120] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[1121] Step 3:
[1122] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[1123] Step 4:
[1124] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[1125] Step 5:
[1126] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[1127] Step 6:
[1128] The server sends the generated pet avatar data (3D model, movement, voice, conversation script) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[1129] Step 7:
[1130] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[1131] Step 8:
[1132] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[1133] Step 9:
[1134] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[1135] Step 10:
[1136] When users talk to their pet avatars in the app, the avatars respond naturally using ChatGPT, which converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[1137] Step 11:
[1138] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[1139] Example 1
[1140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1141] In recent years, the number of people who are unable to keep pets has increased. This has led to a growing demand for systems that allow users to interact with pets in virtual spaces. However, existing systems struggle to generate realistic 3D avatars using photos and videos of pets, to reflect the pet's movements and sounds, and to enable natural conversations with the user. Furthermore, there is a lack of systems that can effectively interact with real-world spaces in AR and VR environments. There is a need for a system that can address these technical challenges and provide a more realistic and intimate virtual pet experience.
[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1143] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, and multimodal means for extracting the pet's movements and cries from the video data and reflecting them in the avatar. This realizes a system that generates a realistic 3D avatar of the pet based on the data provided by the user, enabling natural interaction with the user.
[1144] "Photo data" is still image data provided by the user.
[1145] "Video data" refers to data of moving images provided by the user.
[1146] "Image generation means" refers to the technology or method for generating a 3D avatar of a pet in a virtual space based on photographic data.
[1147] "Multimodal means" refers to technologies and methods for extracting pet gestures and sounds from video data and reflecting these characteristics in an avatar.
[1148] "Conversation generation means" refers to techniques and methods for analyzing users' SNS posting data and engaging in natural conversations with users.
[1149] "SNS posting data" refers to information such as text, images, and videos uploaded by users to social networking services.
[1150] A "generative AI model" is a model that is tuned to perform a specific task using AI techniques such as large-scale language models or image generation models.
[1151] A "terminal" is a hardware device that allows a user to run applications and receive and display data sent from a server.
[1152] An "augmented reality environment" is a technology or system that displays computer-generated information overlaid on the real world.
[1153] A "virtual reality environment" is a technology or system that allows users to have an immersive experience in a computer-generated virtual space.
[1154] A "3D avatar" is a three-dimensional model of a pet generated based on data provided by the user.
[1155] This invention is a system that realistically recreates a pet avatar in a virtual space based on digital data provided by the user. This system is composed of three major elements: a server, a terminal, and a user.
[1156] Server-side processing
[1157] The server receives photos and video data of pets uploaded by users through the app. This includes direct uploads from devices such as smartphones and tablets. The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data, while motion analysis software (e.g., OpenPose) is used to analyze the video data.
[1158] The server uses 3D modeling software (e.g., Blender) to generate a 3D avatar of your pet based on the photo data. This 3D avatar incorporates features from the photos and videos provided by the user. Additionally, the server extracts your pet's movements and sounds from the video data and integrates these attributes into the 3D avatar.
[1159] The server also analyzes users' social media posting data and adjusts generative AI models such as ChatGPT based on the information obtained. This allows the pet avatar to have natural conversations with the user. Finally, the generated pet avatar data is sent to the device. This includes the 3D avatar model data, movement data, voice data, and the adjusted conversation model.
[1160] Terminal side processing
[1161] The device uses a 3D rendering engine (e.g., Unity or Unreal Engine) to display the pet avatar based on the data received from the server. The user can view this avatar through the app. The device also provides an interface for controlling the pet's care (e.g., feeding, walking, etc.).
[1162] The device also has the ability to provide AR and VR environments. When a user selects AR mode, the device activates the camera and uses an AR kit (e.g., ARCore, ARKit) to overlay a pet avatar in real space. Additionally, by using a VR headset, users can interact with the pet avatar in a virtual reality environment.
[1163] User Behavior
[1164] Users begin by uploading photos, videos, and social media posts of their pet through the app. Specifically, they select photos and videos of their pet from their smartphone's gallery and upload the social media posts as well. This accumulates the basic data that the system uses to generate a pet avatar.
[1165] Once the avatar is created, users can talk to the pet avatar within the app and perform care operations (e.g., feeding, walking, etc.) Users can also enjoy realistic interactions with virtual pets by selecting AR or VR mode.
[1166] Specific examples
[1167] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server analyzes the received photos and generates a 3D avatar of their pet dog. It also extracts the dog's running movements and barks from the video data and reflects them in the 3D avatar. It then analyzes the social media posting data and adjusts ChatGPT based on the dog's frequently used phrases and characteristic behaviors. The model generated in this way allows the pet avatar to converse with the user in a familiar manner. The server then sends the 3D avatar, gesture data, bark data, and conversation script to the device.
[1168] The device uses a 3D rendering engine to display a 3D avatar of the dog based on the data received from the server. When the user clicks the food button in the app, the device displays the avatar eating food and showing a happy expression. When the user selects AR mode, the device activates the camera and overlays the dog's 3D avatar in real space.
[1169] Prompt Sentence Examples
[1170] "Generate a 3D avatar of your pet dog and adjust it so that it can have a natural conversation with the user. Also, provide a function to display it in real space in AR mode."
[1171] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1172] Step 1:
[1173] Users upload photos, videos, and social media posts of their pets through the app.
[1174] Specifically, users select photos or videos from their smartphone's gallery and tap the upload button. Users also link their social media accounts in the app's settings screen and allow access to the posted data.
[1175] Input: Pet photos, videos, and SNS posting data
[1176] Output: Photos, videos, and SNS posting data uploaded to the server
[1177] Step 2:
[1178] The server receives the photo and video data provided by the user and stores this data in dedicated storage.
[1179] Specifically, the server receives the uploaded data and stores the photos in a photo folder, the videos in a video folder, and the text data in a text folder.
[1180] Input: User-uploaded photos, videos, and social media posting data
[1181] Output: Data stored in the server storage
[1182] Step 3:
[1183] The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data.
[1184] Specifically, the server analyzes the photo data and detects facial features, body contours, etc.
[1185] Input: Photo data stored in the server storage
[1186] Output: Facial features and body contour information
[1187] Step 4:
[1188] The server uses the analyzed information to generate a 3D avatar of the pet using 3D modeling software (e.g., Blender).
[1189] Specifically, the server uses Blender to create a 3D model of the pet's face and body.
[1190] Input: Facial features and body contour information
[1191] Output: 3D avatar of your pet
[1192] Step 5:
[1193] The server analyzes the video data using motion analysis software (e.g., OpenPose) to extract the pet's movements and sounds.
[1194] Specifically, the server extracts frames from the video and detects the pet's movements and sounds.
[1195] Input: Video data stored in the server storage
[1196] Output: Pet movement patterns and barking data
[1197] Step 6:
[1198] The server integrates the extracted movement patterns and sounds into a 3D avatar.
[1199] In terms of specific movements, the server reflects the movement data and sound data into the 3D model to create a more realistic avatar.
[1200] Input: Pet 3D avatar, movement patterns, and sound data
[1201] Output: 3D avatar of your pet, with movements and sounds reflected
[1202] Step 7:
[1203] The server analyzes the SNS post data and adjusts the generative AI model (e.g., ChatGPT) based on the information obtained.
[1204] Specifically, the server extracts language patterns and behavioral characteristics related to pets from social media data and retrains ChatGPT.
[1205] Input: Social media post data stored in server storage
[1206] Output: Pet-optimized ChatGPT model
[1207] Step 8:
[1208] The server sends the generated pet avatar data to the device, including a 3D model, movement data, voice data, and a conversation script.
[1209] As a specific operation, the server divides the generated data into packets and transmits them to the terminal.
[1210] Input: 3D pet avatar, movement data, voice data, conversation script
[1211] Output: Avatar data sent to the device
[1212] Step 9:
[1213] The device displays the avatar using a 3D rendering engine (e.g., Unity, Unreal Engine) based on the pet avatar data received from the server.
[1214] Specifically, the device loads the received data into a rendering engine and displays the 3D avatar.
[1215] Input: Avatar data received from the server
[1216] Output: 3D avatar of your pet displayed on your device
[1217] Step 10:
[1218] The user performs training operations within the app, and the device displays the corresponding movements on the 3D avatar.
[1219] Specifically, when the user presses the rice button, the device executes the action of the avatar eating rice.
[1220] Input: User interaction
[1221] Output: Avatar movement
[1222] Step 11:
[1223] The device switches to AR mode or VR mode depending on the user's selection.
[1224] Specifically, when a user selects AR mode, the device activates the camera and overlays an avatar onto the real world.
[1225] Input: User mode selection
[1226] Output: Avatar displayed in real space (AR) or avatar operating in virtual space (VR)
[1227] (Application example 1)
[1228] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1229] Today's pet owners want to consider pet products through direct interaction with physical products and services. However, with the spread of online shopping, physical contact and trying on products has become difficult. In particular, pet products require more detailed interaction to confirm size, fit, and usability. Furthermore, services that allow users to record and utilize their pet memories as digital data are not yet widely available. Therefore, there is a need to provide an environment where users can virtually recreate the characteristics and behavior of their pets and use that data to try out and select the best products.
[1230] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1231] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, and means for using the pet avatar to try out products in a virtual store. This allows the user to try out products in the virtual store through an avatar that reflects the characteristics of their pet and select the product that best suits them.
[1232] "Means for receiving photo and video data provided by users" refers to a function that allows users to upload photos and videos to a server via the Internet and receive that data on the server side.
[1233] The "image generation means for generating a pet avatar using photographic data" is a function for recreating the appearance of a pet as a digital 3D model based on photographic data provided by the user.
[1234] "Multimodal means for extracting pet movements and sounds from video data and reflecting them in the avatar" is a function for analyzing the pet's movements and sounds from video data provided by the user and reflecting them in the avatar.
[1235] "A conversation generation means for analyzing user SNS posting data and engaging in natural conversation with the user" is a function for analyzing text data posted by the user on the SNS and generating natural conversation with the user based on that information.
[1236] The "means for using the pet avatar in the virtual store to try out products" is a function that allows a user to try out products such as pet supplies using a pet avatar in a virtual space.
[1237] The "means for displaying an avatar" is a display function that allows the user to visually confirm the generated pet avatar.
[1238] The "fostering means for interaction" is a function that allows the user to interact with the pet avatar through various responses and actions.
[1239] "Means for interacting with the avatar in AR and VR environments" refers to functionality for interacting with a pet avatar in an augmented reality (AR) or virtual reality (VR) environment.
[1240] "Means for displaying a pet avatar in real space using a smartphone, smart glasses, or head-mounted display" refers to a function that uses various devices to display a pet avatar superimposed on real-world scenery, enabling interaction.
[1241] "A conversation generation means adjusted based on user SNS posting data using a large-scale language model" is a function that uses a large-scale natural language processing model to generate conversation content based on information obtained from user SNS posts.
[1242] This invention is a system that uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and uses the avatars to try out products in a virtual store. Below, we will explain an embodiment of this system.
[1243] Server-side processing
[1244] The server includes the following means for receiving data provided by the user and generating and adjusting a pet avatar based on the data:
[1245] Receiving and storing data
[1246] Users upload photos, videos, and social media posts of their pets to the server via their smartphones or other devices. The server receives this data and uses frameworks such as AWS S3 and Django to store it appropriately.
[1247] Image Generation and Multimodal Analysis
[1248] The server uses tools such as TensorFlow, PyTorch, and Blender to generate a 3D avatar of the pet based on the photo data provided by the user. It also extracts the pet's movements and sounds from the video data and applies multimodal analysis to the avatar.
[1249] Adjusting the ChatGPT model
[1250] The server analyzes users' social media posting data and adjusts the ChatGPT model based on the information obtained. This allows the pet avatar to have a natural conversation with the user. The tools used are PyTorch and Huggingface Transformers.
[1251] Product trial in a virtual store
[1252] The server then sends the generated avatar to the terminal as data for trying out products in a virtual store, allowing the user to try out products in the virtual space.
[1253] Terminal side processing
[1254] The terminal side includes the following means for allowing the user to interact with the pet avatar based on the data received from the server.
[1255] Display avatar
[1256] The device displays the pet avatar data received from the server. Users can view the avatar using devices such as smartphones, smart glasses, and head-mounted displays. Specific tools used include Unity and ARKit / ARCore.
[1257] Providing training functions
[1258] The device provides a nurturing element for users to interact with their pet avatar. For example, when a user clicks the food button in the app, the avatar will appear eating food.
[1259] Providing AR and VR functionality
[1260] The device switches to AR / VR mode based on the user's selection, displaying a pet avatar in real space, allowing users to play with their virtual pet in the living room.
[1261] User Behavior
[1262] Users use the system through the following actions:
[1263] Uploading data
[1264] Users use their smartphones or other devices to upload photos, videos, and social media posts of their pets to the server.
[1265] Interacting with avatars
[1266] Users can talk to the avatar and perform training operations to enjoy the avatar's reactions in real time.
[1267] Virtual product trials
[1268] Users can use their pet avatars in the virtual store to try out products, for example, by having their pet's 3D avatar try on a new collar.
[1269] Prompt Sentence Examples
[1270] Below are some examples of prompt sentences.
[1271] "I'll upload photos and videos of my dog and generate a 3D avatar for you in the virtual store."
[1272] "Try having this 3D avatar hold this ball."
[1273] In this way, by implementing the system based on the present invention, users can try out products in a virtual store through their pet avatar, providing a more realistic experience.
[1274] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1275] Step 1:
[1276] A server receives photo and video data from a user.
[1277] Input: Users upload photos and video data of their pets using their smartphones or other devices.
[1278] How it works: The server receives this data via HTTP requests and stores it in AWS S3 or a database.
[1279] Output: Saved photo and video data.
[1280] Step 2:
[1281] The server uses the received photo data to generate a 3D avatar of your pet.
[1282] Input: Saved photo data.
[1283] How it works: The server uses TensorFlow and PyTorch to process images and Blender to generate 3D avatars.
[1284] Output: A 3D avatar of the generated pet.
[1285] Step 3:
[1286] The server extracts the pet's movements and sounds from the video data it receives and reflects them on the avatar.
[1287] Input: Saved video data.
[1288] Motion: The server uses multimodal analysis techniques to analyze the audio and motion data from the video, which is then applied to an avatar in Blender or another 3D modeling tool.
[1289] Output: A 3D avatar of your pet, with movements and sounds.
[1290] Step 4:
[1291] The server analyzes the user's SNS posting data and adjusts the conversation generation model for the pet avatar.
[1292] Input: User's social media posting data.
[1293] How it works: The server uses Huggingface Transformers to train ChatGPT by feeding social media post data into a natural language processing model, allowing the avatar to have a natural conversation with the user.
[1294] Output: A tuned speech generation model.
[1295] Step 5:
[1296] The server sends the data of the generated pet avatar to the terminal.
[1297] Input: 3D pet avatar, gestures, sound data, conversation generation model.
[1298] Operation: The server sends this data to the device using a communication protocol such as HTTP or WebSocket.
[1299] Output: Avatar data sent to the device.
[1300] Step 6:
[1301] The device displays the pet avatar data received from the server.
[1302] Input: Avatar data sent from the server.
[1303] How it works: The device uses Unity and / or ARKit / ARCore to visually display the pet avatar to the user.
[1304] Output: A 3D avatar of the pet displayed on the user's screen.
[1305] Step 7:
[1306] The terminal provides a pet avatar-raising element in response to a user's interaction request.
[1307] Input: User input (e.g. clicking the rice button).
[1308] Action: The device receives the user's input and makes the pet avatar perform the specified action (e.g., eating food).
[1309] Output: Display of the pet avatar's behavior and pet breeding scene.
[1310] Step 8:
[1311] The device displays the pet avatar in an AR / VR environment according to the user's selection.
[1312] Input: The AR / VR mode selected by the user.
[1313] How it works: The device uses its camera and sensors to capture the real world and overlays a pet avatar on top of it.
[1314] Output: A pet avatar displayed in real space.
[1315] Step 9:
[1316] A user tries out products using a pet avatar in a virtual store.
[1317] Input: A user prompt (e.g., "Try letting him hold this ball").
[1318] Operation: The terminal receives instructions from the user and causes the pet avatar to try out the specified product.
[1319] Output: A pet avatar trying out products in a virtual store.
[1320] In this way, users can try out products in a virtual store through their pet avatar, providing a more realistic shopping experience.
[1321] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1322] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[1323] Server-side processing
[1324] 1. Receipt and storage of data
[1325] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[1326] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[1327] 2. Image Generation and Multimodal Analysis
[1328] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[1329] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[1330] 3. Adjusting the ChatGPT model
[1331] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[1332] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[1333] 4. Implementing the Emotion Engine
[1334] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions, and dynamically adjusts the avatar's behavior and conversation content based on the user's emotional state.
[1335] Example: The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. For example, if the user has a sad expression, the pet avatar will be adjusted to send a comforting message.
[1336] 5. Submission of Avatar Data
[1337] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[1338] Example: The server transmits a 3D avatar, gesture data, sound data, conversation scripts, and emotional response patterns to the terminal.
[1339] Terminal side processing
[1340] 1. Displaying Avatars
[1341] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[1342] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[1343] 2. Providing training functions
[1344] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[1345] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[1346] 3. Collaboration with emotion engine
[1347] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[1348] Example: While a user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will provide playful behavior and fun topics.
[1349] 4. Providing AR and VR functionality
[1350] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[1351] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[1352] User Behavior
[1353] 1. Upload your data
[1354] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[1355] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[1356] 2. Interacting with avatars
[1357] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[1358] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[1359] 3. Feeling a response through emotional connection
[1360] The pet avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[1361] Example: If the user shows signs of fatigue, the avatar will offer a kind message such as, "Take it easy today."
[1362] 4. Using AR / VR mode
[1363] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[1364] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[1365] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments. In addition, by combining it with an emotion engine, the avatar can respond and behave in accordance with the user's emotions, providing an even more realistic and moving experience.
[1366] The processing flow will be explained below.
[1367] Step 1:
[1368] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[1369] Step 2:
[1370] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[1371] Step 3:
[1372] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[1373] Step 4:
[1374] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[1375] Step 5:
[1376] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[1377] Step 6:
[1378] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. Specifically, it analyzes the tone and pitch of the voice and changes in facial expressions in real time to identify the user's emotional state.
[1379] Step 7:
[1380] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[1381] Step 8:
[1382] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[1383] Step 9:
[1384] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[1385] Step 10:
[1386] The device uses a camera and microphone to collect the user's voice and facial expression data and transmits it to the emotion engine, which then analyzes the collected voice and video data to determine the user's emotional state in real time.
[1387] Step 11:
[1388] Based on the results from the emotion engine, the device adjusts the behavior and responses of the pet avatar. For example, if the user is sad, the avatar will display comforting behavior and conversations.
[1389] Step 12:
[1390] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[1391] Step 13:
[1392] When users talk to their pet avatars or perform pet-raising operations within the app, the avatars respond naturally using ChatGPT. Specifically, the app converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[1393] Step 14:
[1394] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[1395] Example 2
[1396] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1397] Previous pet avatar systems lacked the realism of generated avatars due to insufficient analysis of user-provided photos and video data. Furthermore, emotion recognition was not performed during user interaction, making it impossible to realize behaviors and responses that correspond to the emotions of individual users. Furthermore, interaction in AR / VR environments was limited, making it difficult to seamlessly integrate reality and virtuality.
[1398] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1399] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, emotion recognition means for analyzing the user's voice and facial expression data and adjusting the avatar's behavior and conversation content based on the user's emotional state, and means for transmitting the generated pet avatar data to the terminal. This allows the server to analyze a variety of user data to generate a realistic and responsive pet avatar, enabling interaction based on the user's emotions.
[1400] The "means for receiving photo and video data provided by the user" refers to a method for uploading photo and video files to a server via the Internet from a terminal owned by the user.
[1401] "Image generation means" refers to the algorithms and technology that analyzes uploaded photo data and generates a 3D model based on it.
[1402] "Multimodal methods" are methods that integrate multiple pieces of information (gestures, sounds, etc.) extracted from video data and reflect them in a 3D avatar.
[1403] "Conversation generation means" is a technology that analyzes users' SNS posting data and generates natural conversations based on the information obtained.
[1404] "Emotion recognition means" refers to the algorithms and techniques utilized to analyze a user's voice and facial expression data and recognize the user's emotional state.
[1405] "Means for transmitting data of the generated pet avatar to the terminal" refers to a communication method for transmitting the 3D model and related data generated by the server to the user's terminal.
[1406] "Means for displaying an avatar" refers to a technique for displaying avatar data received from a server on a user's terminal.
[1407] "Raising means" refers to a method of providing various functions (e.g., feeding, walking, etc.) that allow users to interact with their pet avatar within the app.
[1408] "Means of interacting in AR and VR environments" refers to methods that use augmented reality and virtual reality technologies to combine the real world with the virtual world and allow users to interact with avatars.
[1409] "Means for recording user interactions" refers to technology that records all operations and conversations that a user has with an avatar as a log.
[1410] A "large-scale language model" is an artificial intelligence algorithm that learns from large amounts of text data and enables natural language generation.
[1411] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[1412] Server-side processing
[1413] Receiving and storing data
[1414] The server receives photos, videos, and social media posting data provided by users through the smartphone app. The received data is appropriately saved in the photo folder, video folder, and text folder, respectively.
[1415] Examples:
[1416] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in various folders.
[1417] Image Generation and Multimodal Analysis
[1418] The server analyzes the received photo data using deep learning technology (e.g., TensorFlow, PyTorch) to generate a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the 3D avatar.
[1419] Examples:
[1420] The server analyzes photos of pet dogs uploaded by users and generates a 3D avatar. It also extracts the dog's running movements and barks from video data and reflects them in the avatar.
[1421] Adjusting the ChatGPT model
[1422] The server analyzes users' social media posting data and adjusts a large-scale language model (e.g., OpenAI GPT-4), enabling the pet avatar to have natural conversations with the user.
[1423] Examples:
[1424] The server analyzes users' social media posting data, extracts their frequently used phrases and characteristic behaviors, and adjusts the ChatGPT model, allowing the pet avatar to converse with the user in a familiar way.
[1425] Implementing the Emotion Engine
[1426] The server implements an emotion engine that analyzes the user's emotions using voice recognition and facial expression recognition technologies (e.g., Microsoft Azure Face API), and adjusts the avatar's behavior and conversation content based on the user's emotional state.
[1427] Examples:
[1428] The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. If the user looks sad, the avatar will send a comforting message.
[1429] Sending avatar data
[1430] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[1431] Examples:
[1432] The server transmits the 3D avatar, gesture data, sound data, conversation script, and emotional response patterns to the terminal.
[1433] Terminal side processing
[1434] Display avatar
[1435] The device displays the pet avatar data received from the server, and this avatar is shown to the user through the app.
[1436] Examples:
[1437] When a user launches the app, the device displays a 3D avatar of their beloved dog based on the received data.
[1438] Providing training functions
[1439] The terminal provides elements for raising the pet avatar (e.g., feeding it, taking it for a walk) that allow the user to interact with the pet avatar.
[1440] Examples:
[1441] When a user clicks the meal button in the app, the device displays an avatar eating the meal and showing a happy expression.
[1442] Collaboration with emotion engine
[1443] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[1444] Examples:
[1445] While the user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will display playful behavior and provide fun topics.
[1446] Providing AR and VR functionality
[1447] The device switches to AR (augmented reality) / VR (virtual reality) mode depending on the user's selection, and displays the pet avatar in real space.
[1448] Examples:
[1449] When a user selects AR mode, the device activates the camera and displays a 3D avatar of their beloved dog overlaid on the real world, allowing the user to play with the virtual pet in the living room.
[1450] User Behavior
[1451] Uploading data
[1452] Users can upload photos, videos, and social media posts of their beloved dogs through the app.
[1453] Examples:
[1454] Users select photos and videos of their pet dog and upload them to the server through the app. Past social media posting data is also sent to the server.
[1455] Interacting with avatars
[1456] Users can talk to their pet avatars and control their care within the app, and the avatars respond in real time.
[1457] Examples:
[1458] Users can talk to their pet dog using the chat function within the app, and the pet dog avatar will respond in natural conversation using ChatGPT.
[1459] Feeling a response through emotional connection
[1460] The avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[1461] Examples:
[1462] If the user shows signs of fatigue, the avatar will offer kind words such as, "Take it easy and rest today."
[1463] Using AR / VR mode
[1464] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[1465] Examples:
[1466] Users can activate the AR mode and have their pet's avatar displayed in the actual room, making it feel as if the pet is actually there.
[1467] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1468] Step 1:
[1469] Users upload photo and video data
[1470] Specific operation: The user selects photos and videos of their pet dog from the smartphone gallery and clicks the "Upload" button. This action sends the selected data to the server.
[1471] Input: User-selected photo and video data
[1472] Output: Photo and video data sent to the server
[1473] Step 2:
[1474] The server receives and stores the photo and video data.
[1475] Specific operation: The server classifies and saves the received data in the appropriate folder. Photos are saved in the "Photos folder" and videos in the "Videos folder."
[1476] Input: Photo and video data submitted by the user
[1477] Output: Photo and video data stored on the server
[1478] Step 3:
[1479] The server analyzes the photo data and generates a 3D avatar
[1480] Specific operation: The server uses deep learning techniques (e.g., TensorFlow, PyTorch) to analyze the photo data and generate a 3D model based on the pet's appearance.
[1481] Input: Photo data stored on the server
[1482] Output: Generated 3D avatar
[1483] Step 4:
[1484] The server analyzes the video data and extracts behaviors and sounds.
[1485] Specific movements: The server analyzes the video frame by frame to extract the pet's movement patterns and sounds, which then adds realistic movements and sounds to the avatar.
[1486] Input: Video data stored on the server
[1487] Output: Extracted movement patterns and call data
[1488] Step 5:
[1489] The server analyzes the user's SNS posting data and adjusts ChatGPT
[1490] How it works: The server analyzes social media posting data, extracts the user's linguistic expressions and preferences, and adjusts the ChatGPT model, allowing the avatar to engage in natural and personalized conversations.
[1491] Input: User's SNS post data
[1492] Output: Adjusted ChatGPT model
[1493] Step 6:
[1494] The server implements the emotion recognition engine.
[1495] Specific operation: The server uses voice recognition technology and facial expression recognition technology (e.g., Microsoft Azure Face API) to analyze the user's voice and facial expression data to recognize emotions. The avatar's behavior and conversation content are adjusted based on the emotions.
[1496] Input: User's voice and facial expression data
[1497] Output: Avatar response pattern based on the user's emotional state
[1498] Step 7:
[1499] The server sends the generated avatar data to the device.
[1500] Specific operation: The server sends the 3D avatar, movement data, sound data, conversation script, and emotional response patterns to the device. The user's device receives these and uses them within the app.
[1501] Input: Generated avatar data (3D model, movement data, sounds, conversation script, emotional response patterns)
[1502] Output: Avatar data sent to the device
[1503] Step 8:
[1504] The device displays an avatar
[1505] Specific operation: When a user launches the app, the device displays a 3D avatar based on the data received from the server.
[1506] Input: Avatar data sent from the server
[1507] Output: 3D avatar displayed on the device
[1508] Step 9:
[1509] The device provides training functions
[1510] Specific operation: The device provides functions for raising pets, such as feeding and walking, so that users can interact with their pet avatars.
[1511] Input: User interaction operations
[1512] Output: Real-time avatar response (e.g. eating behavior)
[1513] Step 10:
[1514] The device will work with the emotion engine
[1515] Specific operation: The device transmits the user's voice and facial expression data to the server's emotion recognition engine in real time, and adjusts the avatar's behavior and conversation.
[1516] Input: Real-time collected voice and facial expression data
[1517] Output: Adjusted avatar behavior and dialogue
[1518] Step 11:
[1519] The device offers AR and VR capabilities
[1520] Specific operation: When a user selects AR / VR mode within the app, the device activates the corresponding hardware (e.g., camera, VR goggles) and displays the pet avatar in real or virtual space.
[1521] Input: User mode selection operation
[1522] Output: Avatar displayed in real or virtual space
[1523] (Application example 2)
[1524] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1525] Although systems that realistically recreate pet avatars in virtual spaces have existed, they have struggled to recognize the user's emotions and provide interactions that respond to the user's real-time mood. Furthermore, previous systems have limited the functionality that users can use to interact with their pet avatars, making the experience unrealistic and restrictive, especially in augmented reality (AR) and virtual reality (VR) environments. Furthermore, the pet avatars' conversational capabilities are limited, and systems that can provide more natural conversations based on the user's social networking service (SNS) posting data have yet to be fully realized.
[1526] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving image data, video data, and data posted to a social networking service provided by a user; image generation means for generating an animal avatar using the image data; multimodal means for extracting animal behavior and sound from the video data and reflecting them in the avatar; dialogue generation means for analyzing the user's social networking service post data and engaging in natural dialogue with the user; and emotion recognition means for analyzing the user's emotions in real time. This allows the avatar's behavior and dialogue content to be dynamically adjusted based on the user's emotional state, enabling more natural and realistic interactions. Furthermore, it is possible to enhance interaction with the avatar in augmented reality and virtual reality environments, improving the user experience.
[1527] "Image data" is digital data that contains visual information, such as photographs and illustrations.
[1528] "Moving image data" refers to digital data that includes visual information that changes over time, and includes moving images and animations.
[1529] "Social networking service posted data" refers to digital data including text, images, videos, and other content posted by users to social networking services.
[1530] An "animal avatar" is an animal character in a virtual space that is generated based on digital data provided by the user.
[1531] "Image generation means" means a technique or method for generating a visual representation, such as a 3D model, based on provided image data.
[1532] A "multimodal means" is a technique or method that integrates different types of data (e.g., images, audio, text) for analysis and processing.
[1533] The "dialogue generation means" is a technique or method for generating conversation content based on provided data in order to realize natural dialogue with the user.
[1534] "Emotion recognition means" refers to a technique or method for analyzing the user's voice and facial expressions in real time and recognizing the user's emotional state.
[1535] An "augmented reality environment" is a technology or method that overlays virtual visual information onto real visual information.
[1536] A "virtual reality environment" is a technology or method that provides a user with an immersive virtual space by generating completely virtual visual and audio information.
[1537] "Dynamic adjustment" means changing the system's operation or behavior in real time according to the situation or conditions.
[1538] The present invention is a system that generates realistic animal avatars in a virtual space using image data, video data, and data posted on social networking services provided by users, and further recognizes the user's emotions in real time and dynamically adjusts interactions with the avatars. Specific program processing and the hardware and software used are described below.
[1539] Server-side processing
[1540] 1. Receipt and storage of data
[1541] The server receives image data, video data, and data posted to social networking services that users have uploaded through the application, and stores this data in the appropriate folders. This process uses folders for images, video, and text data.
[1542] 2. Image Generation and Multimodal Analysis
[1543] The server analyzes the received image data and generates an animal avatar using image generation technology. It also extracts the animal's behavior and sounds from the video data and reflects them in the avatar. Specifically, the image generation technology uses 3D modeling software and AI-based image generation algorithms.
[1544] 3. Dialogue Generation Method
[1545] The server analyzes data posted on social networking services and generates conversations based on this data using large-scale language models (e.g., GPT-3 or GPT-4). This conversation generation model is adjusted based on the user's past posting data, enabling realistic and natural conversations.
[1546] 4. Emotion recognition means
[1547] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. The emotion engine uses Google's Emotion API and Microsoft's AI for Emotional Analysis. The avatar's behavior and dialogue content are dynamically adjusted according to the user's emotional state.
[1548] 5. Submission of Avatar Data
[1549] The server transmits the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the user's device.
[1550] Terminal side processing
[1551] 1. Displaying Avatars
[1552] The device displays the animal avatar data received from the server, and can be a smartphone with a high-performance graphics processor, smart glasses, or a head-mounted display.
[1553] 2. Providing training functions
[1554] The device provides a nurturing function that allows users to interact with the animal avatar, including the ability to virtually feed it and provide simple training.
[1555] 3. Collaboration with emotion engine
[1556] The device sends the user's voice and facial expression data to an emotion engine in real time, which adjusts the avatar's movements and speech based on the user's emotional state.
[1557] 4. Providing AR and VR functionality
[1558] The device switches to AR / VR mode based on the user's selection and displays an animal avatar in real space, allowing users to play with a virtual pet in a real space such as their living room.
[1559] Specific examples
[1560] As a concrete example, when a user launches a virtual pet shop app, a login screen is displayed. After logging in, the user enters the virtual shop, and the camera on their smartphone or smart glasses captures the surrounding environment to create a virtual pet shop interior. When the user approaches a pet avatar, detailed information about that pet is displayed. The detailed information reflects the actual characteristics of the pet using photos, videos, and data posted on social networking services uploaded by the user.
[1561] Prompt Sentence Examples
[1562] "I'm creating a virtual pet shop app. I want to create an app that allows users to interact with multiple virtual pets. I want the app to detect when the user is smiling or sad and change the avatar's reaction accordingly. I also want to display detailed information (photos, videos, social media posts) about the pet the user has selected. How can I do this?"
[1563] The above is a specific embodiment for carrying out the present invention, which allows users to enjoy dynamic interactions according to their emotional state.
[1564] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1565] Step 1:
[1566] Image data, video data, and data posted to social networking services provided by the user are uploaded to the server via the application. This involves the user selecting data from their smartphone or PC and pressing the send button. The input is the various data provided by the user, and the output is the state in which that data is received on the server side.
[1567] Step 2:
[1568] The server stores the received image data, video data, and data posted to social networking services in the appropriate folders. Specifically, the data is classified and stored in folders for images, video, and text data. The input is the uploaded raw data, and the output is the data saved in each folder.
[1569] Step 3:
[1570] The server analyzes the image data and generates animal avatars using image generation technology, utilizing AI-based image generation algorithms and 3D modeling software. The input is the image data, and the output is the generated 3D model of the animal avatar.
[1571] Step 4:
[1572] The server extracts the animal's behavior and sound from the video data and reflects it in the avatar. Specifically, it uses video analysis technology to analyze the pet's movements and cries, and incorporates these behaviors and sounds into the animal avatar. The input is video data, and the output is an animal avatar that reflects the behavior and sound.
[1573] Step 5:
[1574] The server analyzes data posted on social networking services and adjusts a dialogue generation model using a large-scale language model (such as GPT-3 or GPT-4). This model enables natural dialogue with users. The input is text data from the social networking service, and the output is the adjusted dialogue generation model.
[1575] Step 6:
[1576] The server analyzes the user's voice data and facial expression data and executes emotion recognition using tools such as Google's Emotion API and Microsoft's AI for Emotional Analysis. The input is voice data and facial expression data, and the output is the user's emotional state.
[1577] Step 7:
[1578] The server sends the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the device. The input is a series of avatar data generated on the server side, and the output is the avatar data received on the device side.
[1579] Step 8:
[1580] The device displays the received animal avatar data and allows the user to interact with the avatar. For example, the avatar appears on the screen and performs actions in response to user commands. The input is the avatar data received from the server, and the output is an interactive avatar displayed on the device screen.
[1581] Step 9:
[1582] The terminal sends the user's voice and facial expression data to the emotion recognition means in real time, and dynamically adjusts the avatar's behavior and dialogue based on the analysis results. The input is the user's emotion data collected in real time, and the output is the dynamically adjusted avatar's behavior and dialogue content.
[1583] Step 10:
[1584] The device switches to AR / VR mode according to the user's selection and displays an animal avatar in the real space, allowing the user to play with a virtual pet in a real room or environment. The input is the user's mode selection and camera data, and the output is the animal avatar displayed in the AR / VR environment.
[1585] The above are the specific processing steps for carrying out the present invention.
[1586] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1587] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1588] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1589] [Fourth embodiment]
[1590] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1591] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1592] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1593] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1594] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1595] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1596] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1597] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1598] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1599] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1600] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1601] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1602] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1603] This system uses photos, videos, and social media posts provided by users to realistically recreate pet avatars in a virtual space. This system consists of three major components: a server, a device, and users.
[1604] Server-side processing
[1605] 1. Receipt and storage of data
[1606] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[1607] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[1608] 2. Image Generation and Multimodal Analysis
[1609] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[1610] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[1611] 3. Adjusting the ChatGPT model
[1612] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[1613] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[1614] 4. Submission of Avatar Data
[1615] The server transmits the data of the generated pet avatar to the device, which enables the device to display the avatar.
[1616] Example: The server sends a 3D avatar, gesture data, sound data, and a conversation script to the device.
[1617] Terminal side processing
[1618] 1. Displaying Avatars
[1619] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[1620] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[1621] 2. Providing training functions
[1622] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[1623] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[1624] 3. Providing AR and VR functionality
[1625] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[1626] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[1627] User Behavior
[1628] 1. Upload your data
[1629] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[1630] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[1631] 2. Interacting with avatars
[1632] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[1633] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[1634] 3. Using AR / VR mode
[1635] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[1636] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[1637] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments.
[1638] The processing flow will be explained below.
[1639] Step 1:
[1640] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[1641] Step 2:
[1642] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[1643] Step 3:
[1644] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[1645] Step 4:
[1646] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[1647] Step 5:
[1648] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[1649] Step 6:
[1650] The server sends the generated pet avatar data (3D model, movement, voice, conversation script) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[1651] Step 7:
[1652] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[1653] Step 8:
[1654] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[1655] Step 9:
[1656] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[1657] Step 10:
[1658] When users talk to their pet avatars in the app, the avatars respond naturally using ChatGPT, which converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[1659] Step 11:
[1660] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[1661] Example 1
[1662] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1663] In recent years, the number of people who are unable to keep pets has increased. This has led to a growing demand for systems that allow users to interact with pets in virtual spaces. However, existing systems struggle to generate realistic 3D avatars using photos and videos of pets, to reflect the pet's movements and sounds, and to enable natural conversations with the user. Furthermore, there is a lack of systems that can effectively interact with real-world spaces in AR and VR environments. There is a need for a system that can address these technical challenges and provide a more realistic and intimate virtual pet experience.
[1664] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1665] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, and multimodal means for extracting the pet's movements and cries from the video data and reflecting them in the avatar. This realizes a system that generates a realistic 3D avatar of the pet based on the data provided by the user, enabling natural interaction with the user.
[1666] "Photo data" is still image data provided by the user.
[1667] "Video data" refers to data of moving images provided by the user.
[1668] "Image generation means" refers to the technology or method for generating a 3D avatar of a pet in a virtual space based on photographic data.
[1669] "Multimodal means" refers to technologies and methods for extracting pet gestures and sounds from video data and reflecting these characteristics in an avatar.
[1670] "Conversation generation means" refers to techniques and methods for analyzing users' SNS posting data and engaging in natural conversations with users.
[1671] "SNS posting data" refers to information such as text, images, and videos uploaded by users to social networking services.
[1672] A "generative AI model" is a model that is tuned to perform a specific task using AI techniques such as large-scale language models or image generation models.
[1673] A "terminal" is a hardware device that allows a user to run applications and receive and display data sent from a server.
[1674] An "augmented reality environment" is a technology or system that displays computer-generated information overlaid on the real world.
[1675] A "virtual reality environment" is a technology or system that allows users to have an immersive experience in a computer-generated virtual space.
[1676] A "3D avatar" is a three-dimensional model of a pet generated based on data provided by the user.
[1677] This invention is a system that realistically recreates a pet avatar in a virtual space based on digital data provided by the user. This system is composed of three major elements: a server, a terminal, and a user.
[1678] Server-side processing
[1679] The server receives photos and video data of pets uploaded by users through the app. This includes direct uploads from devices such as smartphones and tablets. The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data, while motion analysis software (e.g., OpenPose) is used to analyze the video data.
[1680] The server uses 3D modeling software (e.g., Blender) to generate a 3D avatar of your pet based on the photo data. This 3D avatar incorporates features from the photos and videos provided by the user. Additionally, the server extracts your pet's movements and sounds from the video data and integrates these attributes into the 3D avatar.
[1681] The server also analyzes users' social media posting data and adjusts generative AI models such as ChatGPT based on the information obtained. This allows the pet avatar to have natural conversations with the user. Finally, the generated pet avatar data is sent to the device. This includes the 3D avatar model data, movement data, voice data, and the adjusted conversation model.
[1682] Terminal side processing
[1683] The device uses a 3D rendering engine (e.g., Unity or Unreal Engine) to display the pet avatar based on the data received from the server. The user can view this avatar through the app. The device also provides an interface for controlling the pet's care (e.g., feeding, walking, etc.).
[1684] The device also has the ability to provide AR and VR environments. When a user selects AR mode, the device activates the camera and uses an AR kit (e.g., ARCore, ARKit) to overlay a pet avatar in real space. Additionally, by using a VR headset, users can interact with the pet avatar in a virtual reality environment.
[1685] User Behavior
[1686] Users begin by uploading photos, videos, and social media posts of their pet through the app. Specifically, they select photos and videos of their pet from their smartphone's gallery and upload the social media posts as well. This accumulates the basic data that the system uses to generate a pet avatar.
[1687] Once the avatar is created, users can talk to the pet avatar within the app and perform care operations (e.g., feeding, walking, etc.) Users can also enjoy realistic interactions with virtual pets by selecting AR or VR mode.
[1688] Specific examples
[1689] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server analyzes the received photos and generates a 3D avatar of their pet dog. It also extracts the dog's running movements and barks from the video data and reflects them in the 3D avatar. It then analyzes the social media posting data and adjusts ChatGPT based on the dog's frequently used phrases and characteristic behaviors. The model generated in this way allows the pet avatar to converse with the user in a familiar manner. The server then sends the 3D avatar, gesture data, bark data, and conversation script to the device.
[1690] The device uses a 3D rendering engine to display a 3D avatar of the dog based on the data received from the server. When the user clicks the food button in the app, the device displays the avatar eating food and showing a happy expression. When the user selects AR mode, the device activates the camera and overlays the dog's 3D avatar in real space.
[1691] Prompt Sentence Examples
[1692] "Generate a 3D avatar of your pet dog and adjust it so that it can have a natural conversation with the user. Also, provide a function to display it in real space in AR mode."
[1693] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1694] Step 1:
[1695] Users upload photos, videos, and social media posts of their pets through the app.
[1696] Specifically, users select photos or videos from their smartphone's gallery and tap the upload button. Users also link their social media accounts in the app's settings screen and allow access to the posted data.
[1697] Input: Pet photos, videos, and SNS posting data
[1698] Output: Photos, videos, and SNS posting data uploaded to the server
[1699] Step 2:
[1700] The server receives the photo and video data provided by the user and stores this data in dedicated storage.
[1701] Specifically, the server receives the uploaded data and stores the photos in a photo folder, the videos in a video folder, and the text data in a text folder.
[1702] Input: User-uploaded photos, videos, and social media posting data
[1703] Output: Data stored in the server storage
[1704] Step 3:
[1705] The server uses image analysis software (e.g., OpenCV, TensorFlow) to analyze the received photo data.
[1706] Specifically, the server analyzes the photo data and detects facial features, body contours, etc.
[1707] Input: Photo data stored in the server storage
[1708] Output: Facial features and body contour information
[1709] Step 4:
[1710] The server uses the analyzed information to generate a 3D avatar of the pet using 3D modeling software (e.g., Blender).
[1711] Specifically, the server uses Blender to create a 3D model of the pet's face and body.
[1712] Input: Facial features and body contour information
[1713] Output: 3D avatar of your pet
[1714] Step 5:
[1715] The server analyzes the video data using motion analysis software (e.g., OpenPose) to extract the pet's movements and sounds.
[1716] Specifically, the server extracts frames from the video and detects the pet's movements and sounds.
[1717] Input: Video data stored in the server storage
[1718] Output: Pet movement patterns and barking data
[1719] Step 6:
[1720] The server integrates the extracted movement patterns and sounds into a 3D avatar.
[1721] In terms of specific movements, the server reflects the movement data and sound data into the 3D model to create a more realistic avatar.
[1722] Input: Pet 3D avatar, movement patterns, and sound data
[1723] Output: 3D avatar of your pet, with movements and sounds reflected
[1724] Step 7:
[1725] The server analyzes the SNS post data and adjusts the generative AI model (e.g., ChatGPT) based on the information obtained.
[1726] Specifically, the server extracts language patterns and behavioral characteristics related to pets from social media data and retrains ChatGPT.
[1727] Input: Social media post data stored in server storage
[1728] Output: Pet-optimized ChatGPT model
[1729] Step 8:
[1730] The server sends the generated pet avatar data to the device, including a 3D model, movement data, voice data, and a conversation script.
[1731] As a specific operation, the server divides the generated data into packets and transmits them to the terminal.
[1732] Input: 3D pet avatar, movement data, voice data, conversation script
[1733] Output: Avatar data sent to the device
[1734] Step 9:
[1735] The device displays the avatar using a 3D rendering engine (e.g., Unity, Unreal Engine) based on the pet avatar data received from the server.
[1736] Specifically, the device loads the received data into a rendering engine and displays the 3D avatar.
[1737] Input: Avatar data received from the server
[1738] Output: 3D avatar of your pet displayed on your device
[1739] Step 10:
[1740] The user performs training operations within the app, and the device displays the corresponding movements on the 3D avatar.
[1741] Specifically, when the user presses the rice button, the device executes the action of the avatar eating rice.
[1742] Input: User interaction
[1743] Output: Avatar movement
[1744] Step 11:
[1745] The device switches to AR mode or VR mode depending on the user's selection.
[1746] Specifically, when a user selects AR mode, the device activates the camera and overlays an avatar onto the real world.
[1747] Input: User mode selection
[1748] Output: Avatar displayed in real space (AR) or avatar operating in virtual space (VR)
[1749] (Application example 1)
[1750] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1751] Today's pet owners want to consider pet products through direct interaction with physical products and services. However, with the spread of online shopping, physical contact and trying on products has become difficult. In particular, pet products require more detailed interaction to confirm size, fit, and usability. Furthermore, services that allow users to record and utilize their pet memories as digital data are not yet widely available. Therefore, there is a need to provide an environment where users can virtually recreate the characteristics and behavior of their pets and use that data to try out and select the best products.
[1752] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1753] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, and means for using the pet avatar to try out products in a virtual store. This allows the user to try out products in the virtual store through an avatar that reflects the characteristics of their pet and select the product that best suits them.
[1754] "Means for receiving photo and video data provided by users" refers to a function that allows users to upload photos and videos to a server via the Internet and receive that data on the server side.
[1755] The "image generation means for generating a pet avatar using photographic data" is a function for recreating the appearance of a pet as a digital 3D model based on photographic data provided by the user.
[1756] "Multimodal means for extracting pet movements and sounds from video data and reflecting them in the avatar" is a function for analyzing the pet's movements and sounds from video data provided by the user and reflecting them in the avatar.
[1757] "A conversation generation means for analyzing user SNS posting data and engaging in natural conversation with the user" is a function for analyzing text data posted by the user on the SNS and generating natural conversation with the user based on that information.
[1758] The "means for using the pet avatar in the virtual store to try out products" is a function that allows a user to try out products such as pet supplies using a pet avatar in a virtual space.
[1759] The "means for displaying an avatar" is a display function that allows the user to visually confirm the generated pet avatar.
[1760] The "fostering means for interaction" is a function that allows the user to interact with the pet avatar through various responses and actions.
[1761] "Means for interacting with the avatar in AR and VR environments" refers to functionality for interacting with a pet avatar in an augmented reality (AR) or virtual reality (VR) environment.
[1762] "Means for displaying a pet avatar in real space using a smartphone, smart glasses, or head-mounted display" refers to a function that uses various devices to display a pet avatar superimposed on real-world scenery, enabling interaction.
[1763] "A conversation generation means adjusted based on user SNS posting data using a large-scale language model" is a function that uses a large-scale natural language processing model to generate conversation content based on information obtained from user SNS posts.
[1764] This invention is a system that uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and uses the avatars to try out products in a virtual store. Below, we will explain an embodiment of this system.
[1765] Server-side processing
[1766] The server includes the following means for receiving data provided by the user and generating and adjusting a pet avatar based on the data:
[1767] Receiving and storing data
[1768] Users upload photos, videos, and social media posts of their pets to the server via their smartphones or other devices. The server receives this data and uses frameworks such as AWS S3 and Django to store it appropriately.
[1769] Image Generation and Multimodal Analysis
[1770] The server uses tools such as TensorFlow, PyTorch, and Blender to generate a 3D avatar of the pet based on the photo data provided by the user. It also extracts the pet's movements and sounds from the video data and applies multimodal analysis to the avatar.
[1771] Adjusting the ChatGPT model
[1772] The server analyzes users' social media posting data and adjusts the ChatGPT model based on the information obtained. This allows the pet avatar to have a natural conversation with the user. The tools used are PyTorch and Huggingface Transformers.
[1773] Product trial in a virtual store
[1774] The server then sends the generated avatar to the terminal as data for trying out products in a virtual store, allowing the user to try out products in the virtual space.
[1775] Terminal side processing
[1776] The terminal side includes the following means for allowing the user to interact with the pet avatar based on the data received from the server.
[1777] Display avatar
[1778] The device displays the pet avatar data received from the server. Users can view the avatar using devices such as smartphones, smart glasses, and head-mounted displays. Specific tools used include Unity and ARKit / ARCore.
[1779] Providing training functions
[1780] The device provides a nurturing element for users to interact with their pet avatar. For example, when a user clicks the food button in the app, the avatar will appear eating food.
[1781] Providing AR and VR functionality
[1782] The device switches to AR / VR mode based on the user's selection, displaying a pet avatar in real space, allowing users to play with their virtual pet in the living room.
[1783] User Behavior
[1784] Users use the system through the following actions:
[1785] Uploading data
[1786] Users use their smartphones or other devices to upload photos, videos, and social media posts of their pets to the server.
[1787] Interacting with avatars
[1788] Users can talk to the avatar and perform training operations to enjoy the avatar's reactions in real time.
[1789] Virtual product trials
[1790] Users can use their pet avatars in the virtual store to try out products, for example, by having their pet's 3D avatar try on a new collar.
[1791] Prompt Sentence Examples
[1792] Below are some examples of prompt sentences.
[1793] "I'll upload photos and videos of my dog and generate a 3D avatar for you in the virtual store."
[1794] "Try having this 3D avatar hold this ball."
[1795] In this way, by implementing the system based on the present invention, users can try out products in a virtual store through their pet avatar, providing a more realistic experience.
[1796] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1797] Step 1:
[1798] A server receives photo and video data from a user.
[1799] Input: Users upload photos and video data of their pets using their smartphones or other devices.
[1800] How it works: The server receives this data via HTTP requests and stores it in AWS S3 or a database.
[1801] Output: Saved photo and video data.
[1802] Step 2:
[1803] The server uses the received photo data to generate a 3D avatar of your pet.
[1804] Input: Saved photo data.
[1805] How it works: The server uses TensorFlow and PyTorch to process images and Blender to generate 3D avatars.
[1806] Output: A 3D avatar of the generated pet.
[1807] Step 3:
[1808] The server extracts the pet's movements and sounds from the video data it receives and reflects them on the avatar.
[1809] Input: Saved video data.
[1810] Motion: The server uses multimodal analysis techniques to analyze the audio and motion data from the video, which is then applied to an avatar in Blender or another 3D modeling tool.
[1811] Output: A 3D avatar of your pet, with movements and sounds.
[1812] Step 4:
[1813] The server analyzes the user's SNS posting data and adjusts the conversation generation model for the pet avatar.
[1814] Input: User's social media posting data.
[1815] How it works: The server uses Huggingface Transformers to train ChatGPT by feeding social media post data into a natural language processing model, allowing the avatar to have a natural conversation with the user.
[1816] Output: A tuned speech generation model.
[1817] Step 5:
[1818] The server sends the data of the generated pet avatar to the terminal.
[1819] Input: 3D pet avatar, gestures, sound data, conversation generation model.
[1820] Operation: The server sends this data to the device using a communication protocol such as HTTP or WebSocket.
[1821] Output: Avatar data sent to the device.
[1822] Step 6:
[1823] The device displays the pet avatar data received from the server.
[1824] Input: Avatar data sent from the server.
[1825] How it works: The device uses Unity and / or ARKit / ARCore to visually display the pet avatar to the user.
[1826] Output: A 3D avatar of the pet displayed on the user's screen.
[1827] Step 7:
[1828] The terminal provides a pet avatar-raising element in response to a user's interaction request.
[1829] Input: User input (e.g. clicking the rice button).
[1830] Action: The device receives the user's input and makes the pet avatar perform the specified action (e.g., eating food).
[1831] Output: Display of the pet avatar's behavior and pet breeding scene.
[1832] Step 8:
[1833] The device displays the pet avatar in an AR / VR environment according to the user's selection.
[1834] Input: The AR / VR mode selected by the user.
[1835] How it works: The device uses its camera and sensors to capture the real world and overlays a pet avatar on top of it.
[1836] Output: A pet avatar displayed in real space.
[1837] Step 9:
[1838] A user tries out products using a pet avatar in a virtual store.
[1839] Input: A user prompt (e.g., "Try letting him hold this ball").
[1840] Operation: The terminal receives instructions from the user and causes the pet avatar to try out the specified product.
[1841] Output: A pet avatar trying out products in a virtual store.
[1842] In this way, users can try out products in a virtual store through their pet avatar, providing a more realistic shopping experience.
[1843] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1844] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[1845] Server-side processing
[1846] 1. Receipt and storage of data
[1847] Users upload photos, videos, and social media posts of their pets to the server through the app, and the server receives and stores the data appropriately.
[1848] Example: A user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in a photo folder, video folder, and text folder, respectively.
[1849] 2. Image Generation and Multimodal Analysis
[1850] The server analyzes the received photo data and uses image generation technology to create a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the avatar.
[1851] Example: The server analyzes photos of a user's pet dog and generates a 3D avatar. It also extracts the dog's running movements and barking sounds from video data and reflects them in the 3D avatar.
[1852] 3. Adjusting the ChatGPT model
[1853] The server analyzes SNS posting data and adjusts ChatGPT based on the information obtained, allowing the pet avatar to have natural conversations with the user.
[1854] Example: The server analyzes the user's social media posting data and adjusts ChatGPT based on the phrases and characteristic behaviors of the pet dog. The generated model allows the pet avatar to converse with the user in a familiar way.
[1855] 4. Implementing the Emotion Engine
[1856] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions, and dynamically adjusts the avatar's behavior and conversation content based on the user's emotional state.
[1857] Example: The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. For example, if the user has a sad expression, the pet avatar will be adjusted to send a comforting message.
[1858] 5. Submission of Avatar Data
[1859] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[1860] Example: The server transmits a 3D avatar, gesture data, sound data, conversation scripts, and emotional response patterns to the terminal.
[1861] Terminal side processing
[1862] 1. Displaying Avatars
[1863] The device displays the pet avatar data received from the server, and the user can view the avatar through the app.
[1864] Example: When a user launches the app, the device displays a 3D avatar of their pet dog based on the received data.
[1865] 2. Providing training functions
[1866] The terminal provides elements for raising the pet avatar (for example, feeding it, taking it for walks, etc.) that allow the user to interact with the pet avatar.
[1867] Example: When a user clicks the meal button in the app, the device displays an avatar eating a meal and showing a happy expression.
[1868] 3. Collaboration with emotion engine
[1869] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[1870] Example: While a user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will provide playful behavior and fun topics.
[1871] 4. Providing AR and VR functionality
[1872] The device switches to AR / VR mode according to the user's selection and displays the pet avatar in real space.
[1873] Example: When a user selects AR mode, the device activates the camera and overlays a 3D avatar of their pet dog in the real world, allowing the user to play with the virtual pet in the living room.
[1874] User Behavior
[1875] 1. Upload your data
[1876] Through the app, users upload photos, videos, and social media posts of their pets, which provides the system with the basic data to generate a pet avatar.
[1877] Example: A user selects photos and videos of their pet dog and uploads them to a server through the app. Past social media posting data is also sent to the server.
[1878] 2. Interacting with avatars
[1879] Users can talk to their pet avatars and perform breeding operations within the app, which causes the avatars to respond in real time.
[1880] Example: A user talks to their dog using the in-app chat feature, and the dog avatar responds in a natural conversation using ChatGPT.
[1881] 3. Feeling a response through emotional connection
[1882] The pet avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[1883] Example: If the user shows signs of fatigue, the avatar will offer a kind message such as, "Take it easy today."
[1884] 4. Using AR / VR mode
[1885] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[1886] Example: A user activates AR mode and sees an avatar of their pet in a real room, making the experience feel as if the pet is actually there.
[1887] As described above, the system of the present invention allows users to enjoy a variety of interactions with pet avatars in virtual spaces and AR / VR environments. In addition, by combining it with an emotion engine, the avatar can respond and behave in accordance with the user's emotions, providing an even more realistic and moving experience.
[1888] The processing flow will be explained below.
[1889] Step 1:
[1890] The user opens the app and selects the "Pet Avatar Creation" menu. The user selects and uploads photos and videos of their pet from their photo gallery or camera. The user also allows the app to access their social media posting data and posting history.
[1891] Step 2:
[1892] The server receives uploaded photos, videos, and SNS post data. It analyzes the received data and sorts and saves it into image folders (image data), video folders (video data), and text folders (SNS data).
[1893] Step 3:
[1894] The server inputs the photo data into the image generation AI to generate a 3D avatar of the pet. Specifically, the server provides the photo data to the image generation AI model and obtains the 3D avatar data generated from the model.
[1895] Step 4:
[1896] The server inputs video data into the multimodal AI, extracts the pet's movements and sounds, and reflects them in the avatar. Specifically, it breaks down the video data into frames and analyzes movement patterns and sounds. The extracted data is then integrated into the avatar.
[1897] Step 5:
[1898] The server analyzes the social media posts, initializes the ChatGPT model based on the information obtained, and generates conversation data based on the pet's personality. Specifically, the social media data is analyzed to extract keywords and phrases, which are then used as training data for ChatGPT.
[1899] Step 6:
[1900] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. Specifically, it analyzes the tone and pitch of the voice and changes in facial expressions in real time to identify the user's emotional state.
[1901] Step 7:
[1902] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the device. Avatar data is sent to the device in real time using WebSocket or API.
[1903] Step 8:
[1904] The pet avatar data received by the device is displayed on the main screen of the app. Specifically, the 3D avatar is rendered and displayed on the screen. An event listener is set up to enable interaction.
[1905] Step 9:
[1906] The device provides an interface for the nurturing functions (feeding, walking, etc.) and updates the avatar's state according to the user's actions. Specifically, clicking the feed or walk button in the app executes logic that changes the avatar's behavior and reactions.
[1907] Step 10:
[1908] The device uses a camera and microphone to collect the user's voice and facial expression data and transmits it to the emotion engine, which then analyzes the collected voice and video data to determine the user's emotional state in real time.
[1909] Step 11:
[1910] Based on the results from the emotion engine, the device adjusts the behavior and responses of the pet avatar. For example, if the user is sad, the avatar will display comforting behavior and conversations.
[1911] Step 12:
[1912] The device switches to AR / VR mode upon user request and displays the pet avatar in real space. Specifically, it uses ARKit or ARCore to process camera input and overlay the avatar in real space.
[1913] Step 13:
[1914] When users talk to their pet avatars or perform pet-raising operations within the app, the avatars respond naturally using ChatGPT. Specifically, the app converts the user's voice input into text, sends it to ChatGPT, and plays back the generated response.
[1915] Step 14:
[1916] Users can take photos and videos while playing with their pet avatar in AR / VR mode, and can record interactions in AR / VR mode and save or share them within the app.
[1917] Example 2
[1918] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1919] Previous pet avatar systems lacked the realism of generated avatars due to insufficient analysis of user-provided photos and video data. Furthermore, emotion recognition was not performed during user interaction, making it impossible to realize behaviors and responses that correspond to the emotions of individual users. Furthermore, interaction in AR / VR environments was limited, making it difficult to seamlessly integrate reality and virtuality.
[1920] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1921] In this invention, the server includes means for receiving photo and video data provided by the user, image generation means for generating a pet avatar using the photo data, multimodal means for extracting the pet's gestures and sounds from the video data and reflecting them in the avatar, conversation generation means for analyzing the user's SNS posting data and engaging in natural conversation with the user, emotion recognition means for analyzing the user's voice and facial expression data and adjusting the avatar's behavior and conversation content based on the user's emotional state, and means for transmitting the generated pet avatar data to the terminal. This allows the server to analyze a variety of user data to generate a realistic and responsive pet avatar, enabling interaction based on the user's emotions.
[1922] The "means for receiving photo and video data provided by the user" refers to a method for uploading photo and video files to a server via the Internet from a terminal owned by the user.
[1923] "Image generation means" refers to the algorithms and technology that analyzes uploaded photo data and generates a 3D model based on it.
[1924] "Multimodal methods" are methods that integrate multiple pieces of information (gestures, sounds, etc.) extracted from video data and reflect them in a 3D avatar.
[1925] "Conversation generation means" is a technology that analyzes users' SNS posting data and generates natural conversations based on the information obtained.
[1926] "Emotion recognition means" refers to the algorithms and techniques utilized to analyze a user's voice and facial expression data and recognize the user's emotional state.
[1927] "Means for transmitting data of the generated pet avatar to the terminal" refers to a communication method for transmitting the 3D model and related data generated by the server to the user's terminal.
[1928] "Means for displaying an avatar" refers to a technique for displaying avatar data received from a server on a user's terminal.
[1929] "Raising means" refers to a method of providing various functions (e.g., feeding, walking, etc.) that allow users to interact with their pet avatar within the app.
[1930] "Means of interacting in AR and VR environments" refers to methods that use augmented reality and virtual reality technologies to combine the real world with the virtual world and allow users to interact with avatars.
[1931] "Means for recording user interactions" refers to technology that records all operations and conversations that a user has with an avatar as a log.
[1932] A "large-scale language model" is an artificial intelligence algorithm that learns from large amounts of text data and enables natural language generation.
[1933] This system uses photos, videos, and social media posting data provided by users to realistically recreate pet avatars in a virtual space, and furthermore, recognizes the user's emotions and adjusts interactions with the avatar. This system is composed of three major elements: a server, a terminal, and the user.
[1934] Server-side processing
[1935] Receiving and storing data
[1936] The server receives photos, videos, and social media posting data provided by users through the smartphone app. The received data is appropriately saved in the photo folder, video folder, and text folder, respectively.
[1937] Examples:
[1938] The user uploads photos and videos of their pet dog from their smartphone gallery and allows access to past social media posting data. The server receives these and stores them in various folders.
[1939] Image Generation and Multimodal Analysis
[1940] The server analyzes the received photo data using deep learning technology (e.g., TensorFlow, PyTorch) to generate a 3D avatar of your pet. It also extracts your pet's movements and sounds from the video data and reflects them in the 3D avatar.
[1941] Examples:
[1942] The server analyzes photos of pet dogs uploaded by users and generates a 3D avatar. It also extracts the dog's running movements and barks from video data and reflects them in the avatar.
[1943] Adjusting the ChatGPT model
[1944] The server analyzes users' social media posting data and adjusts a large-scale language model (e.g., OpenAI GPT-4), enabling the pet avatar to have natural conversations with the user.
[1945] Examples:
[1946] The server analyzes users' social media posting data, extracts their frequently used phrases and characteristic behaviors, and adjusts the ChatGPT model, allowing the pet avatar to converse with the user in a familiar way.
[1947] Implementing the Emotion Engine
[1948] The server implements an emotion engine that analyzes the user's emotions using voice recognition and facial expression recognition technologies (e.g., Microsoft Azure Face API), and adjusts the avatar's behavior and conversation content based on the user's emotional state.
[1949] Examples:
[1950] The server analyzes facial expression data from audio data and camera images collected while the user is using the app to recognize the user's emotions. If the user looks sad, the avatar will send a comforting message.
[1951] Sending avatar data
[1952] The server sends the generated pet avatar data (3D model, movement, voice, conversation script, emotional response pattern) to the terminal.
[1953] Examples:
[1954] The server transmits the 3D avatar, gesture data, sound data, conversation script, and emotional response patterns to the terminal.
[1955] Terminal side processing
[1956] Display avatar
[1957] The device displays the pet avatar data received from the server, and this avatar is shown to the user through the app.
[1958] Examples:
[1959] When a user launches the app, the device displays a 3D avatar of their beloved dog based on the received data.
[1960] Providing training functions
[1961] The terminal provides elements for raising the pet avatar (e.g., feeding it, taking it for a walk) that allow the user to interact with the pet avatar.
[1962] Examples:
[1963] When a user clicks the meal button in the app, the device displays an avatar eating the meal and showing a happy expression.
[1964] Collaboration with emotion engine
[1965] The device sends the user's voice and facial expression data to the emotion engine in real time, and adjusts the avatar's behavior and speech based on the user's emotional state.
[1966] Examples:
[1967] While the user is using the app, data collected through the camera and microphone is sent to the emotion engine. If the emotion engine recognizes the user's emotion as "fun," the avatar will display playful behavior and provide fun topics.
[1968] Providing AR and VR functionality
[1969] The device switches to AR (augmented reality) / VR (virtual reality) mode depending on the user's selection, and displays the pet avatar in real space.
[1970] Examples:
[1971] When a user selects AR mode, the device activates the camera and displays a 3D avatar of their beloved dog overlaid on the real world, allowing the user to play with the virtual pet in the living room.
[1972] User Behavior
[1973] Uploading data
[1974] Users can upload photos, videos, and social media posts of their beloved dogs through the app.
[1975] Examples:
[1976] Users select photos and videos of their pet dog and upload them to the server through the app. Past social media posting data is also sent to the server.
[1977] Interacting with avatars
[1978] Users can talk to their pet avatars and control their care within the app, and the avatars respond in real time.
[1979] Examples:
[1980] Users can talk to their pet dog using the chat function within the app, and the pet dog avatar will respond in natural conversation using ChatGPT.
[1981] Feeling a response through emotional connection
[1982] The avatar's movements and responses change based on the user's emotions, allowing for a more realistic experience.
[1983] Examples:
[1984] If the user shows signs of fatigue, the avatar will offer kind words such as, "Take it easy and rest today."
[1985] Using AR / VR mode
[1986] Users can select AR or VR mode and play with their pet avatar in real or virtual space.
[1987] Examples:
[1988] Users can activate the AR mode and have their pet's avatar displayed in the actual room, making it feel as if the pet is actually there.
[1989] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1990] Step 1:
[1991] Users upload photo and video data
[1992] Specific operation: The user selects photos and videos of their pet dog from the smartphone gallery and clicks the "Upload" button. This action sends the selected data to the server.
[1993] Input: User-selected photo and video data
[1994] Output: Photo and video data sent to the server
[1995] Step 2:
[1996] The server receives and stores the photo and video data.
[1997] Specific operation: The server classifies and saves the received data in the appropriate folder. Photos are saved in the "Photos folder" and videos in the "Videos folder."
[1998] Input: Photo and video data submitted by the user
[1999] Output: Photo and video data stored on the server
[2000] Step 3:
[2001] The server analyzes the photo data and generates a 3D avatar
[2002] Specific operation: The server uses deep learning techniques (e.g., TensorFlow, PyTorch) to analyze the photo data and generate a 3D model based on the pet's appearance.
[2003] Input: Photo data stored on the server
[2004] Output: Generated 3D avatar
[2005] Step 4:
[2006] The server analyzes the video data and extracts behaviors and sounds.
[2007] Specific movements: The server analyzes the video frame by frame to extract the pet's movement patterns and sounds, which then adds realistic movements and sounds to the avatar.
[2008] Input: Video data stored on the server
[2009] Output: Extracted movement patterns and call data
[2010] Step 5:
[2011] The server analyzes the user's SNS posting data and adjusts ChatGPT
[2012] How it works: The server analyzes social media posting data, extracts the user's linguistic expressions and preferences, and adjusts the ChatGPT model, allowing the avatar to engage in natural and personalized conversations.
[2013] Input: User's SNS post data
[2014] Output: Adjusted ChatGPT model
[2015] Step 6:
[2016] The server implements the emotion recognition engine.
[2017] Specific operation: The server uses voice recognition technology and facial expression recognition technology (e.g., Microsoft Azure Face API) to analyze the user's voice and facial expression data to recognize emotions. The avatar's behavior and conversation content are adjusted based on the emotions.
[2018] Input: User's voice and facial expression data
[2019] Output: Avatar response pattern based on the user's emotional state
[2020] Step 7:
[2021] The server sends the generated avatar data to the device.
[2022] Specific operation: The server sends the 3D avatar, movement data, sound data, conversation script, and emotional response patterns to the device. The user's device receives these and uses them within the app.
[2023] Input: Generated avatar data (3D model, movement data, sounds, conversation script, emotional response patterns)
[2024] Output: Avatar data sent to the device
[2025] Step 8:
[2026] The device displays an avatar
[2027] Specific operation: When a user launches the app, the device displays a 3D avatar based on the data received from the server.
[2028] Input: Avatar data sent from the server
[2029] Output: 3D avatar displayed on the device
[2030] Step 9:
[2031] The device provides training functions
[2032] Specific operation: The device provides functions for raising pets, such as feeding and walking, so that users can interact with their pet avatars.
[2033] Input: User interaction operations
[2034] Output: Real-time avatar response (e.g. eating behavior)
[2035] Step 10:
[2036] The device will work with the emotion engine
[2037] Specific operation: The device transmits the user's voice and facial expression data to the server's emotion recognition engine in real time, and adjusts the avatar's behavior and conversation.
[2038] Input: Real-time collected voice and facial expression data
[2039] Output: Adjusted avatar behavior and dialogue
[2040] Step 11:
[2041] The device offers AR and VR capabilities
[2042] Specific operation: When a user selects AR / VR mode within the app, the device activates the corresponding hardware (e.g., camera, VR goggles) and displays the pet avatar in real or virtual space.
[2043] Input: User mode selection operation
[2044] Output: Avatar displayed in real or virtual space
[2045] (Application example 2)
[2046] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2047] Although systems that realistically recreate pet avatars in virtual spaces have existed, they have struggled to recognize the user's emotions and provide interactions that respond to the user's real-time mood. Furthermore, previous systems have limited the functionality that users can use to interact with their pet avatars, making the experience unrealistic and restrictive, especially in augmented reality (AR) and virtual reality (VR) environments. Furthermore, the pet avatars' conversational capabilities are limited, and systems that can provide more natural conversations based on the user's social networking service (SNS) posting data have yet to be fully realized.
[2048] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving image data, video data, and data posted to a social networking service provided by a user; image generation means for generating an animal avatar using the image data; multimodal means for extracting animal behavior and sound from the video data and reflecting them in the avatar; dialogue generation means for analyzing the user's social networking service post data and engaging in natural dialogue with the user; and emotion recognition means for analyzing the user's emotions in real time. This allows the avatar's behavior and dialogue content to be dynamically adjusted based on the user's emotional state, enabling more natural and realistic interactions. Furthermore, it is possible to enhance interaction with the avatar in augmented reality and virtual reality environments, improving the user experience.
[2049] "Image data" is digital data that contains visual information, such as photographs and illustrations.
[2050] "Moving image data" refers to digital data that includes visual information that changes over time, and includes moving images and animations.
[2051] "Social networking service posted data" refers to digital data including text, images, videos, and other content posted by users to social networking services.
[2052] An "animal avatar" is an animal character in a virtual space that is generated based on digital data provided by the user.
[2053] "Image generation means" means a technique or method for generating a visual representation, such as a 3D model, based on provided image data.
[2054] A "multimodal means" is a technique or method that integrates different types of data (e.g., images, audio, text) for analysis and processing.
[2055] The "dialogue generation means" is a technique or method for generating conversation content based on provided data in order to realize natural dialogue with the user.
[2056] "Emotion recognition means" refers to a technique or method for analyzing the user's voice and facial expressions in real time and recognizing the user's emotional state.
[2057] An "augmented reality environment" is a technology or method that overlays virtual visual information onto real visual information.
[2058] A "virtual reality environment" is a technology or method that provides a user with an immersive virtual space by generating completely virtual visual and audio information.
[2059] "Dynamic adjustment" means changing the system's operation or behavior in real time according to the situation or conditions.
[2060] The present invention is a system that generates realistic animal avatars in a virtual space using image data, video data, and data posted on social networking services provided by users, and further recognizes the user's emotions in real time and dynamically adjusts interactions with the avatars. Specific program processing and the hardware and software used are described below.
[2061] Server-side processing
[2062] 1. Receipt and storage of data
[2063] The server receives image data, video data, and data posted to social networking services that users have uploaded through the application, and stores this data in the appropriate folders. This process uses folders for images, video, and text data.
[2064] 2. Image Generation and Multimodal Analysis
[2065] The server analyzes the received image data and generates an animal avatar using image generation technology. It also extracts the animal's behavior and sounds from the video data and reflects them in the avatar. Specifically, the image generation technology uses 3D modeling software and AI-based image generation algorithms.
[2066] 3. Dialogue Generation Method
[2067] The server analyzes data posted on social networking services and generates conversations based on this data using large-scale language models (e.g., GPT-3 or GPT-4). This conversation generation model is adjusted based on the user's past posting data, enabling realistic and natural conversations.
[2068] 4. Emotion recognition means
[2069] The server implements an emotion engine that analyzes the user's voice data and facial expression data to recognize emotions. The emotion engine uses Google's Emotion API and Microsoft's AI for Emotional Analysis. The avatar's behavior and dialogue content are dynamically adjusted according to the user's emotional state.
[2070] 5. Submission of Avatar Data
[2071] The server transmits the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the user's device.
[2072] Terminal side processing
[2073] 1. Displaying Avatars
[2074] The device displays the animal avatar data received from the server, and can be a smartphone with a high-performance graphics processor, smart glasses, or a head-mounted display.
[2075] 2. Providing training functions
[2076] The device provides a nurturing function that allows users to interact with the animal avatar, including the ability to virtually feed it and provide simple training.
[2077] 3. Collaboration with emotion engine
[2078] The device sends the user's voice and facial expression data to an emotion engine in real time, which adjusts the avatar's movements and speech based on the user's emotional state.
[2079] 4. Providing AR and VR functionality
[2080] The device switches to AR / VR mode based on the user's selection and displays an animal avatar in real space, allowing users to play with a virtual pet in a real space such as their living room.
[2081] Specific examples
[2082] As a concrete example, when a user launches a virtual pet shop app, a login screen is displayed. After logging in, the user enters the virtual shop, and the camera on their smartphone or smart glasses captures the surrounding environment to create a virtual pet shop interior. When the user approaches a pet avatar, detailed information about that pet is displayed. The detailed information reflects the actual characteristics of the pet using photos, videos, and data posted on social networking services uploaded by the user.
[2083] Prompt Sentence Examples
[2084] "I'm creating a virtual pet shop app. I want to create an app that allows users to interact with multiple virtual pets. I want the app to detect when the user is smiling or sad and change the avatar's reaction accordingly. I also want to display detailed information (photos, videos, social media posts) about the pet the user has selected. How can I do this?"
[2085] The above is a specific embodiment for carrying out the present invention, which allows users to enjoy dynamic interactions according to their emotional state.
[2086] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2087] Step 1:
[2088] Image data, video data, and data posted to social networking services provided by the user are uploaded to the server via the application. This involves the user selecting data from their smartphone or PC and pressing the send button. The input is the various data provided by the user, and the output is the state in which that data is received on the server side.
[2089] Step 2:
[2090] The server stores the received image data, video data, and data posted to social networking services in the appropriate folders. Specifically, the data is classified and stored in folders for images, video, and text data. The input is the uploaded raw data, and the output is the data saved in each folder.
[2091] Step 3:
[2092] The server analyzes the image data and generates animal avatars using image generation technology, utilizing AI-based image generation algorithms and 3D modeling software. The input is the image data, and the output is the generated 3D model of the animal avatar.
[2093] Step 4:
[2094] The server extracts the animal's behavior and sound from the video data and reflects it in the avatar. Specifically, it uses video analysis technology to analyze the pet's movements and cries, and incorporates these behaviors and sounds into the animal avatar. The input is video data, and the output is an animal avatar that reflects the behavior and sound.
[2095] Step 5:
[2096] The server analyzes data posted on social networking services and adjusts a dialogue generation model using a large-scale language model (such as GPT-3 or GPT-4). This model enables natural dialogue with users. The input is text data from the social networking service, and the output is the adjusted dialogue generation model.
[2097] Step 6:
[2098] The server analyzes the user's voice data and facial expression data and executes emotion recognition using tools such as Google's Emotion API and Microsoft's AI for Emotional Analysis. The input is voice data and facial expression data, and the output is the user's emotional state.
[2099] Step 7:
[2100] The server sends the generated animal avatar data (3D model, behavior, voice, conversation script, emotional response pattern) to the device. The input is a series of avatar data generated on the server side, and the output is the avatar data received on the device side.
[2101] Step 8:
[2102] The device displays the received animal avatar data and allows the user to interact with the avatar. For example, the avatar appears on the screen and performs actions in response to user commands. The input is the avatar data received from the server, and the output is an interactive avatar displayed on the device screen.
[2103] Step 9:
[2104] The terminal sends the user's voice and facial expression data to the emotion recognition means in real time, and dynamically adjusts the avatar's behavior and dialogue based on the analysis results. The input is the user's emotion data collected in real time, and the output is the dynamically adjusted avatar's behavior and dialogue content.
[2105] Step 10:
[2106] The device switches to AR / VR mode according to the user's selection and displays an animal avatar in the real space, allowing the user to play with a virtual pet in a real room or environment. The input is the user's mode selection and camera data, and the output is the animal avatar displayed in the AR / VR environment.
[2107] The above are the specific processing steps for carrying out the present invention.
[2108] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2109] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2110] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2111] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification m...
Claims
1. means for receiving user-provided photo and video data; an image generation means for generating an avatar of a pet using the photographic data; a multimodal means for extracting the pet's movements and cries from the video data and reflecting them on the avatar; A conversation generation means for analyzing user's SNS posting data and having natural conversation with the user; A system including:
2. means for displaying the avatar; a training means for allowing a user to interact with the avatar; means for interacting with said avatar in an AR and VR environment; The system of claim 1 further comprising:
3. The system according to claim 1 , wherein the conversation generation means is adjusted based on user SNS posting data using a large-scale language model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A