System
The system addresses the limitations of video calls by creating a 3D avatar and synchronizing user movements in a virtual space, providing a high-quality communication experience.
Patent Information
- Application Number
- JP2024123998
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Current video calls lack the quality of face-to-face communication, particularly in conveying physical movements and subtle facial expressions, making them inadequate for remote monitoring of children or elderly individuals, and are time-consuming.
A system that performs a 360-degree full-body scan, generates a 3D avatar using a generative AI model, captures user movements with a motion capture device, and synchronizes these in a virtual space for real-time communication.
Enables high-quality communication between distant individuals, allowing for accurate representation of physical movements and facial expressions, enhancing interaction experience.
Smart Images

Figure 2026022481000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, face-to-face communication with far-away family and friends via video calls is common, but this method suffers from the problem that the quality of communication is lower than when you meet in person. Another problem is that using video calls on a daily basis is time-consuming and difficult to continue. Furthermore, video calls have difficulty accurately conveying physical movements and subtle changes in facial expressions, making them particularly inadequate for remotely checking on a child's growth or the health of an elderly person. There is a need for a method that can solve these issues and enable high-quality communication with family and friends in distant locations. [Means for solving the problem]
[0005] The system includes a means for performing a 360-degree full-body scan of a user, a means for preprocessing the acquired image data and transmitting it to a server, a means for generating a 3D avatar based on the received image data using a generative AI model, a means for transmitting the generated 3D avatar to the user's device, and a means for using a motion capture device to acquire user movement data. The system also includes a means for transmitting the acquired movement data to a server in real time, a means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data, a means for synchronizing movement and position information between users in the virtual space, and a means for transmitting and receiving audio data for users to communicate in the virtual space. This system enables high-quality communication between users, even when they are physically separated, as if they were in the same space.
[0006] A "user" is an individual who undergoes a full-body scan and motion capture in order to use the system.
[0007] A "full body scan" is the process of capturing 360-degree images of a user's entire body using a camera or other image capture device to capture image data.
[0008] "Image data" refers to image information of a user's body obtained through a full-body scan.
[0009] "Preprocessing" refers to processing to convert acquired image data into a format that is easy to use, and specifically includes noise removal, resolution adjustment, image combination, and the like.
[0010] A "server" is a central computer system that processes and manages data for the entire system.
[0011] A "generative AI model" is an artificial intelligence model that generates 3D avatars for specific purposes based on input data.
[0012] A "3D avatar" is a three-dimensional virtual image of a person created based on the user's image data.
[0013] A "terminal" refers to a device such as a computer, smartphone, or VR device that a user actually uses.
[0014] A "motion capture device" is a device that captures a user's movements in real time and collects that data.
[0015] "Motion data" refers to data representing the user's body movements captured using a motion capture device.
[0016] A "virtual space" is a computer-generated virtual realm that users can visually experience through a 3D avatar.
[0017] "Real-time" is a time concept that refers to processing occurring immediately with almost no delay.
[0018] "Synchronization" refers to the process of adjusting multiple data and processes at the same time to ensure consistency.
[0019] "Voice data" refers to data that collects a user's voice in digital form and is used to exchange with other users. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data using motion capture, avatar synchronization in a virtual space, and real-time voice communication. Below, the program processing of the entire system and specific examples are explained in natural language.
[0042] Full body scan and image data transmission
[0043] The user starts a dedicated application and scans their entire body in 360 degrees using a smartphone or dedicated camera device, which captures multiple image data.
[0044] The device preprocesses the acquired image data, including noise removal, resolution adjustment, and image merging.
[0045] The terminal compresses the preprocessed image data and sends it to the server, thereby distributing the processing load.
[0046] Creating 3D avatars with generative AI
[0047] The server uses a generative AI model based on the received full-body image data to create a detailed 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[0048] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0049] Obtaining movement data using motion capture
[0050] The user wears a VR device (such as a head-mounted display or a motion capture suit), which captures the user's movements in real time.
[0051] The device acquires operational data from each sensor (e.g., data from the gyro sensor, acceleration sensor, and position sensor).
[0052] The device analyzes the acquired motion data, converts it into a suitable format, compresses it, and transmits it to the server in real time.
[0053] Projection into virtual space and avatar synchronization
[0054] The server processes the motion data received from the user and reproduces the avatar's motion in the virtual space. It also integrates motion data from other participants (family and friends) and synchronizes the motion and location information of all users.
[0055] Communication Features
[0056] Users communicate with other users in the virtual space visually and audibly, for example by talking or waving.
[0057] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0058] The server transfers the received audio data to the terminals of the other participants, and the audio is played back at each terminal.
[0059] Specific examples
[0060] For example, consider the case where grandparents who live far away use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear similar devices at their homes.
[0061] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server. At the same time, the grandparent's device also acquires the grandparent's motion data and sends it to the server.
[0062] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk with their grandchildren as if they were there in the virtual space.
[0063] The devices also exchange voice data between users in real time, providing a natural conversation experience. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[0064] The processing flow will be explained below.
[0065] Step 1:
[0066] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[0067] The user uses a camera to scan their entire body 360 degrees, obtaining multiple image data.
[0068] Step 2:
[0069] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[0070] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[0071] Step 3:
[0072] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[0073] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0074] Step 4:
[0075] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0076] Step 5:
[0077] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[0078] Step 6:
[0079] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[0080] Step 7:
[0081] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[0082] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[0083] Step 8:
[0084] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[0085] Step 9:
[0086] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[0087] Step 10:
[0088] The server receives the voice data of the other participants and transmits it to the user's terminal.
[0089] The audio data received by the device is played through speakers or a headset.
[0090] As described above, the system of the present invention realizes high-quality real-time communication that transcends physical distance.
[0091] Example 1
[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0093] In modern society, users need to communicate with each other in real time, regardless of physical distance. In particular, real-time data communication, both visual and audio, is essential for immersive communication with family and friends in distant locations. However, current systems have difficulty accurately capturing a user's full-body movements and voice in real time and reproducing them in a virtual space. This limits the communication experience and leads to problems such as connection delays and instability.
[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0095] In this invention, the server includes means for performing a 360-degree full-body scan, means for preprocessing the image data and transmitting it to the server, means for generating a 3D avatar using a generative AI model based on the received image data, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to capture user motion data, means for analyzing the captured motion data and converting it into a suitable format, means for transmitting the converted motion data to the server in real time, means for reproducing the motion of the 3D avatar in a virtual space based on the real-time motion data, means for synchronizing motion and position information between users in the virtual space, and means for transmitting and receiving audio data for users to communicate in the virtual space. This allows the user's full-body motions and voices to be reflected in the virtual space in real time, enabling a high-quality communication experience.
[0096] A "means for scanning the entire body 360 degrees" is a device or software that photographs the user's entire body from various angles and collects data from all directions.
[0097] "Means for preprocessing image data and sending it to the server" refers to devices or software that process the acquired image data by removing noise, adjusting resolution, combining images, etc., and then compressing and sending it to the server.
[0098] "Means for generating a 3D avatar using a generative AI model" refers to modules or software that use artificial intelligence technology to generate a three-dimensional avatar based on received image data.
[0099] "Means for transmitting the 3D avatar to the user's terminal" refers to a communication protocol or device for transmitting the generated 3D avatar as data to the end user's terminal.
[0100] "Means using a motion capture device" means a device or software for detecting and capturing a user's physical movements in real time.
[0101] The "means for analyzing the motion data and converting it into a suitable format" refers to a device or software for analyzing the acquired motion data and converting it into a suitable data format for reproduction in a virtual space.
[0102] The "means for transmitting the converted motion data to the server in real time" refers to a device or software for compressing the motion data converted into an appropriate format and transmitting it to the server in real time without any loss.
[0103] "Means for reproducing the movements of a 3D avatar in a virtual space" refers to an engine or software that faithfully reproduces the movements of a 3D avatar in a virtual space based on received movement data.
[0104] "Means for synchronizing the movements and location information of users in a virtual space" refers to a system or software that consistently reflects the movements and location information of multiple users in real time, enabling interaction within a virtual space.
[0105] "Means for transmitting and receiving voice data" refers to a communication system or software for capturing a user's voice in real time and transmitting and playing it back to other users.
[0106] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition and transmission of real-time motion data through motion capture, synchronization of the motion data in a virtual space, and voice communication. A specific embodiment of the entire system will be described below.
[0107] Full body scan and image data transmission
[0108] The user launches a dedicated application on a smartphone or dedicated camera device and scans their entire body in 360 degrees, obtaining multiple image data. Specifically, the user rotates the smartphone to capture their entire body, and the app instructs them on the direction and angle of the shot.
[0109] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. Noise removal is performed using edge detection algorithms, and resolution is optimized. The individual images are then stitched together to generate the entire image.
[0110] The pre-processed image data is compressed by the terminal and sent to the server via HTTP protocol, using the JPEG compression algorithm for efficient uploading to the server.
[0111] Creating 3D avatars with generative AI
[0112] The server uses a generative AI model (e.g., a Deep Learning-based 3D Reconstruction model) based on the received full-body image data to generate a 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[0113] The generated 3D avatar data (mesh data, texture data, etc.) is sent from the server to the user's device. The data is packaged in JSON format and sent via a RESTful API.
[0114] Obtaining movement data using motion capture
[0115] The user wears VR devices (such as a head-mounted display and a motion capture suit) and captures their movements in real time. The user wears VR goggles and sensors from the motion capture suit, which are attached to the user's body, and the sensors capture the user's body movements.
[0116] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received via Bluetooth or Wi-Fi.
[0117] The acquired motion data is analyzed by the device and converted into quaternion format, which is then compressed using a compression algorithm such as Zlib and sent to the server in real time using the WebSocket protocol.
[0118] Projection into virtual space and avatar synchronization
[0119] The server processes the motion data received from the user and reproduces the movements of a 3D avatar in the virtual space. All user motion data is integrated using a real-time motion engine. The data is then reflected in the virtual space using a game engine such as Unity.
[0120] Communication Features
[0121] Users can communicate with other users visually and audibly in the virtual space by talking to other avatars or waving to them. Facial tracking can also be used to reflect facial expressions on the avatars in real time.
[0122] The device captures the user's voice data from the microphone, encodes it using an audio codec such as Opus, and sends it to the server. The voice data is then transferred in real time to other devices and played back through the speakers.
[0123] Specific examples
[0124] For example, grandparents who live far away can use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear VR devices as well.
[0125] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server, while the grandparent's device simultaneously captures the grandparent's motion data and sends it to the server.
[0126] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk as if they were actually there with their grandchildren. The devices also exchange voice data between users in real time, providing a natural conversation experience.
[0127] Prompt Sentence Examples
[0128] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] Users launch a dedicated application on their smartphone or dedicated camera device and perform a 360-degree full-body scan. Users rotate their smartphone to capture their entire body, adjusting the angle and direction of the shot according to the application's instructions.
[0132] Input: Multiple image data captured by a smartphone or dedicated camera.
[0133] Output: 360-degree full-body scan image data.
[0134] Step 2:
[0135] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. After using an image edge detection algorithm to remove noise and optimize the resolution, the individual images are stitched together to generate a complete image.
[0136] Input: 360 degree scanned image data.
[0137] Output: Preprocessed whole-body image data.
[0138] Step 3:
[0139] The terminal compresses the pre-processed image data using the JPEG compression algorithm and sends it to the server via the HTTP protocol. Compressing the image data improves transfer speed and reduces the load on the network bandwidth.
[0140] Input: Preprocessed whole-body image data.
[0141] Output: Compressed image data.
[0142] Step 4:
[0143] The server uses a generative AI model (e.g., a deep learning-based 3D reconstruction model) based on the received image data to generate a detailed 3D avatar of the user, analyzes the image, and generates a 3D mesh and texture.
[0144] Input: Compressed whole-body image data.
[0145] Output: 3D avatar data (mesh data, texture data).
[0146] Step 5:
[0147] The server packages the generated 3D avatar data in JSON format and sends it to the user's device via a RESTful API, which ensures structured and efficient data transfer.
[0148] Input: 3D avatar data (mesh data, texture data).
[0149] Output: 3D avatar data packaged in JSON format.
[0150] Step 6:
[0151] Users wear VR devices (such as a head-mounted display or a motion capture suit) and their movements are captured in real time. Users wear VR goggles and the motion capture suit's sensors are attached to their bodies, allowing the user's movements to be captured in detail.
[0152] Input: User's physical movements.
[0153] Output: Real-time operating data.
[0154] Step 7:
[0155] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received in real time via Bluetooth or Wi-Fi.
[0156] Input: Operational data from each sensor.
[0157] Output: Operational data stored in a local database.
[0158] Step 8:
[0159] The device analyzes the acquired motion data, converts it into quaternion format, compresses it using a compression algorithm such as Zlib, and transmits it to the server in real time using the WebSocket protocol.
[0160] Input: Operational data stored in a local database.
[0161] Output: Compressed quaternion format motion data.
[0162] Step 9:
[0163] The server processes the motion data received from the user using a real-time motion engine and reproduces the movements of a 3D avatar in the virtual space. Motion data from other participants is also integrated and synchronized, and the motion data of all users is reflected in the virtual space.
[0164] Input: Compressed quaternion format motion data.
[0165] Output: 3D avatar movement data reproduced in virtual space.
[0166] Step 10:
[0167] The device captures the user's voice data in real time from the microphone, encodes it using a voice codec such as Opus, and sends it to the server. The voice data is divided into packets with timestamps.
[0168] Input: User's voice data.
[0169] Output: The encoded audio data.
[0170] Step 11:
[0171] The server transfers the user's encoded voice data to the other participants' terminals and plays it in real time. The other participants' voice data is similarly processed and played on the user's terminal.
[0172] Input: Encoded audio data.
[0173] Output: Audio data for playback.
[0174] Prompt Sentence Examples
[0175] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[0176] (Application example 1)
[0177] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0178] One drawback of modern online shopping is that users cannot actually try on products. This often results in purchased products that do not meet actual expectations, resulting in a poor user experience. There is also a demand for a richer shopping experience that does not rely on physical stores. Furthermore, there is a lack of a way for users in different locations to communicate with each other in real time while shopping.
[0179] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0180] In this invention, the server includes: means for scanning the user's entire body 360 degrees; means for preprocessing the acquired image data and transmitting it to the server; means for generating a 3D avatar based on the received image data using a generative AI model; means for transmitting the generated 3D avatar to the user's device; means for using a motion capture device to acquire user movement data; means for transmitting the acquired movement data to the server in real time; means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data; means for synchronizing movement and position information between users in the virtual space; means for transmitting and receiving audio data for users to communicate in the virtual space; and means for providing a system in a virtual store that allows users to try on products online using a 3D avatar. This allows users to enjoy shopping while trying on products in real time from the comfort of their own home and communicating with other users.
[0181] A "means for scanning the entire body 360 degrees" is a device or system that can photograph the user's entire body from each angle and obtain image data from all directions.
[0182] "Preprocessing" refers to processing the acquired image data, such as noise removal, resolution adjustment, and image combination.
[0183] A "generative AI model" is a model that uses an artificial intelligence algorithm to generate a 3D avatar based on acquired image data.
[0184] A "motion capture device" is a device that captures a user's movements in real time and acquires them as movement data.
[0185] A "virtual space" is a computer-generated interactive 3D space in which users can perform various actions through their avatars.
[0186] "Means for synchronizing movement and position information" refers to technologies and systems for matching the movements and positions of multiple users in a virtual space in real time.
[0187] "Means for transmitting and receiving voice data" refers to a device and system for capturing a user's voice and transmitting it to other users via the Internet.
[0188] "Means provided within a virtual store" refers to a system that provides an environment on an online shopping platform where users can virtually try on products using 3D avatars.
[0189] This system scans a user's entire body in 360 degrees, generates a 3D avatar using a generative AI model, and then recreates the avatar in a virtual space in conjunction with motion capture data. This system allows users to try on products in the virtual space and communicate with other users in real time.
[0190] First, the user scans their entire body using a smartphone or dedicated camera device. In this step, the user rotates the device to capture images from all directions (360 degrees). The device preprocesses the captured image data, removing noise and adjusting the resolution, before sending it to the server. The OpenCV library is used for preprocessing.
[0191] The server uses a generative AI model to create a 3D avatar based on the received image data. This AI model uses OpenAI's API to analyze the image data and generate a detailed 3D avatar of the user. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the user's device.
[0192] Next, the user wears a motion capture device (e.g., a VR device or motion capture suit) and their movement data is acquired in real time. The device acquires the movement data from various sensors (e.g., gyro sensor, accelerometer, position sensor), converts it into an appropriate format, compresses it, and sends it to the server. During this process, appropriate libraries (e.g., NumPy or the Python standard library) are used to format and compress the data.
[0193] The server then recreates the movements of the 3D avatar in the virtual space based on the received motion data, and synchronizes the motion and position information with the avatars of other users, allowing all users to experience the sensation of operating the avatar in real time in the same virtual space.
[0194] Furthermore, users can send and receive voice data to communicate within the virtual space. The device picks up the user's voice data from the microphone and sends it to the server. The server then forwards the received voice data to the devices of the other participants, and the voice is played on each device.
[0195] As an example, consider a scenario in which a user tries on products in a virtual store in a virtual space. The user uses a smartphone or VR device to scan their entire body and generate a 3D avatar, then logs into the virtual store. At this time, the user might input a prompt sentence like the following into the generative AI model:
[0196] "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3"
[0197] Users can then view and rotate the products using their own 3D avatar as they try them on in a virtual space, and can also interact with other users in real time to share their thoughts on the fit and design of the products.
[0198] As described above, the present invention allows users to transcend physical constraints and enjoy a richer shopping experience in a virtual space.
[0199] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0200] Step 1:
[0201] The user launches a dedicated application and scans their entire body from all directions using a smartphone or dedicated camera device, acquiring multiple image data. This provides a full-body image of the user as input.
[0202] Step 2:
[0203] The device preprocesses the acquired image data. This preprocessing includes noise removal, resolution adjustment, and image merging. The OpenCV library is used for noise removal, resolution adjustment, and image merging. The preprocessed image data is obtained as the output.
[0204] Step 3:
[0205] The device compresses the preprocessed image data and sends it to the server. The compression uses an appropriate library to efficiently transmit the data. The preprocessed and compressed image data is sent as input to the server.
[0206] Step 4:
[0207] The server generates a 3D avatar for the user using a generative AI model based on the received image data. The generative AI model uses OpenAI's API to analyze the image and generate a detailed 3D avatar. The generated 3D avatar is obtained as the output.
[0208] Step 5:
[0209] The server sends the generated 3D avatar data to the user's device using a standard data transfer protocol. The user's device receives the 3D avatar data.
[0210] Step 6:
[0211] The user wears a motion capture device (e.g., a VR device or a motion capture suit) and motion data is acquired. The motion capture device acquires motion data in real time using gyro sensors, accelerometers, and position sensors. The motion data is obtained as output.
[0212] Step 7:
[0213] The device analyzes the motion data acquired from each sensor and converts it into an appropriate format. This process involves shaping the data and converting it into a compatible format. The analyzed motion data is then output.
[0214] Step 8:
[0215] The terminal compresses the converted motion data and transmits it to the server in real time using a standard data compression algorithm. The compressed motion data is then transmitted as input to the server.
[0216] Step 9:
[0217] The server reproduces the 3D avatar's movements in the virtual space based on the received movement data. This is done using a simulation engine, which reproduces real-time movements in the virtual space. The reproduced movement data is obtained as output.
[0218] Step 10:
[0219] The server synchronizes the motion and location information with the avatars of other users, so that all users can see the synchronized motion in the same virtual space. The synchronized motion and location information is obtained as output.
[0220] Step 11:
[0221] Users send and receive voice data to communicate within the virtual space. The device captures the voice data from the microphone and sends it to the server. The server then transfers the received voice data to other users' devices and plays it back in real time. The sent and received voice data is obtained as output.
[0222] Step 12:
[0223] Users use their own 3D avatar to try on products in a virtual store. Specifically, they input the following prompt into the generative AI model: "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3." This allows users to try on products online and communicate with other users in real time.
[0224] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0225] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and the ability to recognize and reflect the user's emotions using an emotion engine. Below, the program processing of the entire system and specific examples are explained in natural language.
[0226] Full body scan and image data transmission
[0227] The user launches a dedicated application on their smartphone and switches to full-body scan mode, using the camera to scan the entire body 360 degrees and capture multiple image data.
[0228] The terminal preprocesses the acquired image data, for example, by removing noise, adjusting resolution, combining images, etc. The preprocessed image data is compressed and sent to the server.
[0229] Creating 3D avatars with generative AI
[0230] The server analyzes the received full-body image data and generates a detailed 3D avatar of the user using a generative AI model. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the device.
[0231] Obtaining movement data using motion capture
[0232] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0233] The device acquires real-time motion data from the VR device, specifically, data from the gyro sensor, accelerometer, and position sensor.
[0234] The device analyzes the acquired motion data, converts it into an appropriate format, compresses it, and transmits it to the server in real time.
[0235] Projection into virtual space and avatar synchronization
[0236] The server reproduces the user's 3D avatar's movements in the virtual space based on the movement data sent by the user. It also receives the movement data of other participants and synchronizes the movement and position information of all users in the virtual space.
[0237] Emotion recognition and reflection using emotion engine
[0238] The device collects the user's facial expression data and voice data and analyzes this data using an emotion engine.
[0239] The terminal compresses the analysis results and transmits them to the server in real time.
[0240] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space.
[0241] Communication Features
[0242] Users communicate with other users visually and audibly in the virtual space, for example, by talking, exchanging glances, and making gestures.
[0243] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0244] The server receives the audio data from other participants and sends it to the appropriate device, where it is played through speakers or a headset.
[0245] Specific examples
[0246] For example, consider a situation where parents and children living far apart use this system to communicate in a virtual space, with the parent wearing a VR device at home and the child wearing one as well.
[0247] The child's device captures the child's behavior and emotional data in real time and sends it to the server, while the parent's device simultaneously captures the parent's behavior and emotional data and sends it to the server.
[0248] The server integrates this data and recreates and synchronizes it in the virtual space, allowing parents to interact with their child's 3D avatar in real time visually and audibly, and even feel their child's emotions.
[0249] Furthermore, when a parent smiles with joy, that expression is reflected in the child's avatar in real time, enabling natural communication even from a distance. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[0250] The processing flow will be explained below.
[0251] Step 1:
[0252] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[0253] The user uses a camera to scan the entire body 360 degrees and acquires multiple image data.
[0254] Step 2:
[0255] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[0256] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[0257] Step 3:
[0258] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[0259] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0260] Step 4:
[0261] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0262] Step 5:
[0263] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[0264] Step 6:
[0265] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[0266] Step 7:
[0267] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[0268] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[0269] Step 8:
[0270] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[0271] Step 9:
[0272] The device uses an emotion engine to recognize the user's emotions and collects the user's facial expression data and voice data.
[0273] The device analyzes this data to identify the user's emotions.
[0274] Step 10:
[0275] The device compresses the analysis results and sends them to the server in real time.
[0276] Step 11:
[0277] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in the virtual space.
[0278] Step 12:
[0279] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[0280] Step 13:
[0281] The server receives the voice data of the other participants and transmits it to the user's terminal.
[0282] The audio data received by the device is played through speakers or a headset.
[0283] As described above, the system of the present invention realizes high-quality communication through the creation of a 3D avatar using generative AI from a full-body scan of the user, motion capture, emotion recognition using an emotion engine, and real-time reproduction of facial expressions and movements in a virtual space.
[0284] Example 2
[0285] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0286] Current communication technology makes it difficult for people living far apart to interact in a natural way. In particular, the real-time sharing of facial expressions, movements, and emotions has not been fully realized. This has led to a decline in the quality of communication in virtual reality spaces and a limited user experience.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0288] In this invention, the server includes a means for scanning the user's entire body in 360 degrees, a means for preprocessing the acquired image data and sending it to the server, and a means for generating a 3D avatar based on the received image data using a machine learning model, which allows the user to move in real time and communicate naturally with the user, reflecting their emotions.
[0289] "User" refers to a person who uses the system to perform full-body scans, acquire movement data, and communicate within the virtual space.
[0290] "Means for scanning the whole body 360 degrees" refers to a device that captures the user's whole body from all directions and a method of operating the device.
[0291] "Preprocessing" refers to the step of processing the acquired image data by noise removal, resolution adjustment, image combination, etc.
[0292] A "server" refers to a computer system that processes data sent from users, generates avatars in virtual space, and synchronizes movement data.
[0293] A "machine learning model" refers to an algorithm that analyzes input data, recognizes patterns, and generates a 3D avatar.
[0294] A "3D avatar" refers to a three-dimensional virtual character that reproduces the user's appearance and movements.
[0295] "Device" refers to a device required for a user to use the system, such as a smartphone or VR device.
[0296] "Motion detection device" refers to a sensor device or device for capturing a user's physical movements in real time.
[0297] "Real-time" refers to a short time range in which user actions and data are reflected almost instantaneously.
[0298] "Virtual space" refers to a three-dimensional digital environment generated by a computer.
[0299] "Voice Data" means data that records and transmits a user's voice in digital form.
[0300] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and voice data to recognize and reflect their emotional state.
[0301] The system of the present invention is designed to realize high-quality communication in a virtual space. Specifically, it includes functions such as full-body scanning of the user, 3D avatar generation using a machine learning model, real-time motion data acquisition using a motion detection device, avatar synchronization within the virtual space, communication using voice data, and emotion recognition and reflection using an emotion engine.
[0302] First, the user scans their entire body in 360 degrees using a dedicated smartphone application. The user's full-body image data is pre-processed on the device to remove noise, adjust resolution, and combine images, then compressed and sent to a cloud-based server. The server analyzes the received image data and generates a detailed 3D avatar of the user using stable diffusion and other generative AI models. The generated 3D avatar data is then sent back to the device and provided to the user.
[0303] Next, the user puts on a VR device such as a head-mounted display or motion capture suit and launches the corresponding dedicated application. At this time, the device collects data from the VR device in real time using the gyro sensor, accelerometer, and position sensor, converts it into an appropriate format, compresses it, and sends it to the server. This allows the server to instantly reproduce the movements of the user's 3D avatar in the virtual space.
[0304] Furthermore, the system incorporates an emotion engine that can recognize and reflect the user's emotional state by collecting and analyzing facial and voice data, allowing the user's 3D avatar to express emotions in a more natural way.
[0305] For example, when parents and children living far apart use this system to interact in a virtual space, they both wear VR devices at home and launch a dedicated application. The child's device captures movement and emotional data in real time and sends it to the server. At the same time, the parent's device also acquires the parent's movement and emotional data and sends it to the server. The server integrates this data and reproduces and synchronizes it in the virtual space, allowing the parent and child to interact visually and audibly in real time. Furthermore, when the parent smiles with joy, that expression is reflected in the child's avatar in real time.
[0306] In this way, this system provides an advanced communication experience that transcends physical distance.
[0307] Example prompt sentence:
[0308] "Use a generative AI model to generate a 3D avatar from the user's full-body image data. The specific process includes noise removal, resolution adjustment, and avatar generation using the generative AI model."
[0309] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0310] Program processing steps
[0311] Step 1: Full body scan and image data transmission
[0312] The user launches a dedicated application on their smartphone and switches to "full body scan" mode.
[0313] Specific operation: Use the front camera to scan your entire body 360 degrees and obtain multiple image data.
[0314] Input: User's full-body image data
[0315] Once the device receives the captured image data, it applies a noise reduction filter to remove unnecessary information, then adjusts the image resolution to an appropriate size, and finally merges the images to create a single full-body image.
[0316] Output: Preprocessed whole-body image data
[0317] The device compresses the preprocessed image data and sends it to a server on the cloud.
[0318] Specific operation: The terminal automatically compresses and transmits data in parallel.
[0319] Step 2: Creating a 3D avatar using generative AI
[0320] The server sends the received preprocessed image data to the AI processing module.
[0321] Input: Preprocessed whole-body image data
[0322] The server analyzes the image using a generative AI model such as Stable Diffusion to generate a 3D avatar, which includes mesh and texture data.
[0323] Output: Generated 3D avatar data (mesh data, texture data)
[0324] The server sends the generated 3D avatar data to the user's device.
[0325] How it works: The server analyzes the image data and generates and transmits a complete 3D avatar within seconds.
[0326] Step 3: Obtaining movement data using motion capture
[0327] The user puts on a head-mounted display or motion capture suit and launches the corresponding dedicated application.
[0328] Specific actions: The user puts on the VR device and starts moving in the virtual space.
[0329] Input: User behavior data
[0330] The terminal collects data from the gyro sensor, accelerometer, and position sensor in real time from the VR device and converts it into a suitable format.
[0331] Output: Motion data converted into a suitable format
[0332] The terminal compresses the converted motion data and transmits it to the server in real time.
[0333] Specific operation: The device collects and transmits data sequentially in accordance with the user's movements.
[0334] Step 4: Projection into virtual space and avatar synchronization
[0335] The server reproduces the movements of the user's 3D avatar in the virtual space based on the motion data sent by the user.
[0336] Input: Operation data
[0337] Output: Avatar movements reproduced in virtual space
[0338] Specific actions: A user's movements in the virtual space or hand movements are synchronized with other users in real time.
[0339] Step 5: Emotion recognition and reflection using the emotion engine
[0340] The terminal collects facial expression data and voice data of the user.
[0341] Input: Facial expression data, voice data
[0342] The terminal uses an emotion engine to analyze this data and recognize the user's emotional state.
[0343] Output: Emotion recognition result
[0344] The terminal compresses the analysis results and transmits them to the server in real time.
[0345] Specific operation: When a user smiles, the facial expression is immediately analyzed by the emotion engine and reflected on the server.
[0346] Step 6: Communication Functions
[0347] Users can communicate visually and audibly with other users in the virtual space.
[0348] Specific actions: The user speaks or gestures.
[0349] Input: Voice data, gesture data
[0350] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0351] Output: Voice communication data with other users
[0352] The server receives the voice data of the other participants and sends it to the appropriate terminal, where the voice data is played through speakers or headsets.
[0353] Specific operation: The user's voice is transmitted to other users in real time, enabling natural conversation.
[0354] (Application example 2)
[0355] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0356] Conventional communication systems in virtual spaces reproduce the movements and communication of an avatar based solely on the user's motion and voice data, and lack the ability to reflect the user's emotions in real time, making it difficult to achieve a natural communication experience. Furthermore, even in fitting experiences in virtual stores in the real world, there was an issue of not being able to grasp the user's emotions and provide optimal advice and suggestions.
[0357] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0358] In this invention, the server includes means for scanning the user's entire body 360 degrees, means for preprocessing the acquired image data and transmitting it to the server, means for generating a 3D avatar based on the received image data using a generative AI model, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to acquire user movement data, means for transmitting the acquired movement data to the server in real time, means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data, means for synchronizing movement and position information between users in the virtual space, means for transmitting and receiving voice data for users to communicate in the virtual space, means for capturing user facial expression data and analyzing emotions using an emotion engine, and means for reflecting the user's 3D avatar's facial expressions and movements in the virtual space in real time based on the analyzed emotion data. This enables users to communicate in the virtual space and try on clothes in a virtual store in a natural and high-quality manner.
[0359] A "means for scanning the entire body 360 degrees" is a device or method for capturing images of a user's entire body from various angles and acquiring a series of image data.
[0360] "Image data preprocessing means" refers to a method or device that performs processing such as noise removal and resolution adjustment on acquired image data to improve the quality of the data.
[0361] The "means for transmitting to the server" refers to a device or method for transmitting the pre-processed data to the server over a network.
[0362] A "generative AI model" is an artificial intelligence model that generates a detailed 3D avatar of a user based on received image data.
[0363] A "3D avatar" is a three-dimensional digital model that mimics the user's entire body.
[0364] "Motion capture device" is a general term for devices and sensors used to capture a user's movements in real time.
[0365] "Motion data" refers to data that includes information about a user's movements acquired by a motion capture device.
[0366] The "means for transmitting to the server in real time" refers to a device or method for instantly transmitting acquired motion data to the server.
[0367] A "virtual space" is a three-dimensional digital space generated using computer technology.
[0368] "Means for reproducing movements in a virtual space" refers to devices or methods that use acquired movement data to allow a 3D avatar to move in real time within a virtual space.
[0369] The "means for synchronizing motion and position information" refers to a device or method for consistently synchronizing the motion and position information of multiple users in a virtual space.
[0370] "Means for transmitting and receiving voice data" refers to a method or device for transmitting and receiving voice data that allows a user to communicate with other users by voice in a virtual space.
[0371] "Means for capturing facial expression data" refers to a device or method for acquiring a user's facial expression in real time.
[0372] An "emotion engine" is a program or device that analyzes captured facial expression data and voice data to recognize the user's emotional state.
[0373] The "means for analyzing emotions" refers to a device or method for analyzing data acquired using an emotion engine and identifying the user's emotions.
[0374] "Means for reflecting the facial expressions and movements of a 3D avatar in real time within a virtual space" refers to a device or method for instantly changing and reflecting the facial expressions and movements of a 3D avatar within a virtual space based on analyzed emotional data.
[0375] The system of the present invention includes functions such as full-body scanning of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and recognition and reflection of the user's emotions using an emotion engine.
[0376] First, the user uses the smartphone's full-body scanning function to scan their entire body in 360 degrees, acquiring multiple image data. The acquired image data undergoes preprocessing, such as noise removal and resolution adjustment, to improve the data quality before being sent to the server. This preprocessing is carried out using OpenCV, an open-source image processing library.
[0377] The server uses a generative AI model based on the received image data to generate a detailed 3D avatar of the user. The mesh and texture data of the generated 3D avatar are then sent to the user's device, where it can be displayed on a VR device or similar.
[0378] Next, the user puts on a VR device (such as a head-mounted display or motion capture suit) and movement data is acquired. This movement data is collected in real time using gyro sensors, acceleration sensors, position sensors, etc. The acquired movement data is sent to a server and reproduced in real time as the movement of a 3D avatar in the virtual space.
[0379] This system allows multiple users to participate simultaneously and synchronizes their movements and positions in the virtual space, allowing for natural interactions between users.
[0380] Furthermore, the system captures the user's facial expression and voice data and analyzes them using an emotion engine to identify the user's emotions. This analyzed emotion data is reflected in the facial expressions and movements of the user's 3D avatar in real time. The emotion engine uses an emotion recognition model using TensorFlow.
[0381] A concrete example of this system is a virtual store try-on experience. By having a 3D avatar instantly try on the clothes the user has selected, the experience becomes as if they are actually trying them on. In addition, by using an emotion engine, the system can understand the user's reactions and provide optimal fashion advice.
[0382] Example prompt sentence:
[0383] "I would like to demonstrate a real-time try-on experience using a 3D avatar. The 3D avatar will instantly wear the clothes the user has chosen, and an emotion engine will be used to analyze the user's reaction and display the next recommended item."
[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0385] Step 1:
[0386] Full body scan of the user
[0387] The user activates the full-body scan mode on their smartphone and uses the camera to scan their entire body 360 degrees. This acquires multiple image data. The input is the smartphone camera image, and the output is the acquired multiple image data. Specifically, the user rotates the camera to capture images of their entire body.
[0388] Step 2:
[0389] Image data preprocessing
[0390] The device preprocesses the acquired image data, performing noise removal, resolution adjustment, image merging, etc. The input is the acquired image data, and the output is the preprocessed image data. This processing uses the OpenCV library. Specifically, a noise removal filter is applied, and each image is merged to generate a high-resolution whole-body image.
[0391] Step 3:
[0392] Sending image data to the server
[0393] The terminal compresses the preprocessed image data and sends it to the server. The input is the preprocessed image data, and the output is the compressed data. The HTTP protocol is used to send the data. Specifically, the data is converted into a byte sequence and sent using an HTTP request.
[0394] Step 4:
[0395] 3D avatar generation
[0396] The server generates a 3D avatar using a generative AI model based on the received image data. The input is compressed image data, and the output is 3D avatar data (mesh data, texture data, etc.). Specifically, the generative AI model analyzes the image data and constructs a 3D model.
[0397] Step 5:
[0398] Sending a 3D avatar to your device
[0399] The server sends the generated 3D avatar data to the device. The input is the 3D avatar data, and the output is the avatar data sent to the device. Specifically, the data is sent in real time using WebSocket communication.
[0400] Step 6:
[0401] Obtaining movement data using motion capture
[0402] The user puts on a VR device (a head-mounted display or motion capture suit) and launches a dedicated application. The input is the user's movements and the device's sensor data, and the output is movement data. Specifically, data is collected from the gyro sensor, acceleration sensor, and position sensor.
[0403] Step 7:
[0404] Sending operation data to the server
[0405] The terminal analyzes the acquired motion data, converts it into an appropriate format, compresses it, and sends it to the server. The input is the acquired motion data, and the output is compressed motion data. Specifically, the data format is converted and the data is made smaller using a compression algorithm.
[0406] Step 8:
[0407] Reproducing movements in virtual space
[0408] The server reproduces the movements of the user's 3D avatar in the virtual space based on the movement data sent by the user. The input is compressed movement data, and the output is a 3D avatar that moves in the virtual space. Specifically, the server analyzes the data, calculates the position and posture in the virtual space, and reflects them in the 3D avatar.
[0409] Step 9:
[0410] Synchronization of movement and position information in virtual space
[0411] The server also receives other users' motion data and synchronizes the motion and location information of all users in the virtual space. The input is motion data from multiple users, and the output is a synchronized virtual space. Specifically, all user data is centrally managed and adjusted in real time.
[0412] Step 10:
[0413] Facial expression data capture and analysis
[0414] The device captures the user's facial expression data and analyzes the emotions using an emotion engine. The input is the user's facial expression data, and the output is the emotion data that is the analysis result. Specifically, the device analyzes the facial image captured by the camera and identifies the emotional state (joy, anger, sadness, etc.).
[0415] Step 11:
[0416] Sending emotion data to the server
[0417] The device compresses the analyzed emotion data and sends it to the server in real time. The input is the analyzed emotion data, and the output is the compressed emotion data. Specifically, the emotion data is converted into a byte sequence and sent using an HTTP request.
[0418] Step 12:
[0419] Reflecting facial expressions and movements in virtual space
[0420] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space. The input is emotion data, and the output is a 3D avatar that reflects the emotion. Specifically, the emotion data is converted into the avatar's facial motion and reflected in real time.
[0421] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0422] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0423] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0424] [Second embodiment]
[0425] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0426] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0427] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0428] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0429] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0430] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0431] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0432] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0433] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0434] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0435] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0436] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0437] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data using motion capture, avatar synchronization in a virtual space, and real-time voice communication. Below, the program processing of the entire system and specific examples are explained in natural language.
[0438] Full body scan and image data transmission
[0439] The user starts a dedicated application and scans their entire body in 360 degrees using a smartphone or dedicated camera device, which captures multiple image data.
[0440] The device preprocesses the acquired image data, including noise removal, resolution adjustment, and image merging.
[0441] The terminal compresses the preprocessed image data and sends it to the server, thereby distributing the processing load.
[0442] Creating 3D avatars with generative AI
[0443] The server uses a generative AI model based on the received full-body image data to create a detailed 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[0444] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0445] Obtaining movement data using motion capture
[0446] The user wears a VR device (such as a head-mounted display or a motion capture suit), which captures the user's movements in real time.
[0447] The device acquires operational data from each sensor (e.g., data from the gyro sensor, acceleration sensor, and position sensor).
[0448] The device analyzes the acquired motion data, converts it into a suitable format, compresses it, and transmits it to the server in real time.
[0449] Projection into virtual space and avatar synchronization
[0450] The server processes the motion data received from the user and reproduces the avatar's motion in the virtual space. It also integrates motion data from other participants (family and friends) and synchronizes the motion and location information of all users.
[0451] Communication Features
[0452] Users communicate with other users in the virtual space visually and audibly, for example by talking or waving.
[0453] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0454] The server transfers the received audio data to the terminals of the other participants, and the audio is played back at each terminal.
[0455] Specific examples
[0456] For example, consider the case where grandparents who live far away use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear similar devices at their homes.
[0457] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server. At the same time, the grandparent's device also acquires the grandparent's motion data and sends it to the server.
[0458] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk with their grandchildren as if they were there in the virtual space.
[0459] The devices also exchange voice data between users in real time, providing a natural conversation experience. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[0460] The processing flow will be explained below.
[0461] Step 1:
[0462] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[0463] The user uses a camera to scan their entire body 360 degrees, obtaining multiple image data.
[0464] Step 2:
[0465] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[0466] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[0467] Step 3:
[0468] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[0469] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0470] Step 4:
[0471] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0472] Step 5:
[0473] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[0474] Step 6:
[0475] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[0476] Step 7:
[0477] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[0478] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[0479] Step 8:
[0480] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[0481] Step 9:
[0482] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[0483] Step 10:
[0484] The server receives the voice data of the other participants and transmits it to the user's terminal.
[0485] The audio data received by the device is played through speakers or a headset.
[0486] As described above, the system of the present invention realizes high-quality real-time communication that transcends physical distance.
[0487] Example 1
[0488] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0489] In modern society, users need to communicate with each other in real time, regardless of physical distance. In particular, real-time data communication, both visual and audio, is essential for immersive communication with family and friends in distant locations. However, current systems have difficulty accurately capturing a user's full-body movements and voice in real time and reproducing them in a virtual space. This limits the communication experience and leads to problems such as connection delays and instability.
[0490] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0491] In this invention, the server includes means for performing a 360-degree full-body scan, means for preprocessing the image data and transmitting it to the server, means for generating a 3D avatar using a generative AI model based on the received image data, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to capture user motion data, means for analyzing the captured motion data and converting it into a suitable format, means for transmitting the converted motion data to the server in real time, means for reproducing the motion of the 3D avatar in a virtual space based on the real-time motion data, means for synchronizing motion and position information between users in the virtual space, and means for transmitting and receiving audio data for users to communicate in the virtual space. This allows the user's full-body motions and voices to be reflected in the virtual space in real time, enabling a high-quality communication experience.
[0492] A "means for scanning the entire body 360 degrees" is a device or software that photographs the user's entire body from various angles and collects data from all directions.
[0493] "Means for preprocessing image data and sending it to the server" refers to devices or software that process the acquired image data by removing noise, adjusting resolution, combining images, etc., and then compressing and sending it to the server.
[0494] "Means for generating a 3D avatar using a generative AI model" refers to modules or software that use artificial intelligence technology to generate a three-dimensional avatar based on received image data.
[0495] "Means for transmitting the 3D avatar to the user's terminal" refers to a communication protocol or device for transmitting the generated 3D avatar as data to the end user's terminal.
[0496] "Means using a motion capture device" means a device or software for detecting and capturing a user's physical movements in real time.
[0497] The "means for analyzing the motion data and converting it into a suitable format" refers to a device or software for analyzing the acquired motion data and converting it into a suitable data format for reproduction in a virtual space.
[0498] The "means for transmitting the converted motion data to the server in real time" refers to a device or software for compressing the motion data converted into an appropriate format and transmitting it to the server in real time without any loss.
[0499] "Means for reproducing the movements of a 3D avatar in a virtual space" refers to an engine or software that faithfully reproduces the movements of a 3D avatar in a virtual space based on received movement data.
[0500] "Means for synchronizing the movements and location information of users in a virtual space" refers to a system or software that consistently reflects the movements and location information of multiple users in real time, enabling interaction within a virtual space.
[0501] "Means for transmitting and receiving voice data" refers to a communication system or software for capturing a user's voice in real time and transmitting and playing it back to other users.
[0502] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition and transmission of real-time motion data through motion capture, synchronization of the motion data in a virtual space, and voice communication. A specific embodiment of the entire system will be described below.
[0503] Full body scan and image data transmission
[0504] The user launches a dedicated application on a smartphone or dedicated camera device and scans their entire body in 360 degrees, obtaining multiple image data. Specifically, the user rotates the smartphone to capture their entire body, and the app instructs them on the direction and angle of the shot.
[0505] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. Noise removal is performed using edge detection algorithms, and resolution is optimized. The individual images are then stitched together to generate the entire image.
[0506] The pre-processed image data is compressed by the terminal and sent to the server via HTTP protocol, using the JPEG compression algorithm for efficient uploading to the server.
[0507] Creating 3D avatars with generative AI
[0508] The server uses a generative AI model (e.g., a Deep Learning-based 3D Reconstruction model) based on the received full-body image data to generate a 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[0509] The generated 3D avatar data (mesh data, texture data, etc.) is sent from the server to the user's device. The data is packaged in JSON format and sent via a RESTful API.
[0510] Obtaining movement data using motion capture
[0511] The user wears VR devices (such as a head-mounted display and a motion capture suit) and captures their movements in real time. The user wears VR goggles and sensors from the motion capture suit, which are attached to the user's body, and the sensors capture the user's body movements.
[0512] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received via Bluetooth or Wi-Fi.
[0513] The acquired motion data is analyzed by the device and converted into quaternion format, which is then compressed using a compression algorithm such as Zlib and sent to the server in real time using the WebSocket protocol.
[0514] Projection into virtual space and avatar synchronization
[0515] The server processes the motion data received from the user and reproduces the movements of a 3D avatar in the virtual space. All user motion data is integrated using a real-time motion engine. The data is then reflected in the virtual space using a game engine such as Unity.
[0516] Communication Features
[0517] Users can communicate with other users visually and audibly in the virtual space by talking to other avatars or waving to them. Facial tracking can also be used to reflect facial expressions on the avatars in real time.
[0518] The device captures the user's voice data from the microphone, encodes it using an audio codec such as Opus, and sends it to the server. The voice data is then transferred in real time to other devices and played back through the speakers.
[0519] Specific examples
[0520] For example, grandparents who live far away can use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear VR devices as well.
[0521] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server, while the grandparent's device simultaneously captures the grandparent's motion data and sends it to the server.
[0522] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk as if they were actually there with their grandchildren. The devices also exchange voice data between users in real time, providing a natural conversation experience.
[0523] Prompt Sentence Examples
[0524] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[0525] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0526] Step 1:
[0527] Users launch a dedicated application on their smartphone or dedicated camera device and perform a 360-degree full-body scan. Users rotate their smartphone to capture their entire body, adjusting the angle and direction of the shot according to the application's instructions.
[0528] Input: Multiple image data captured by a smartphone or dedicated camera.
[0529] Output: 360-degree full-body scan image data.
[0530] Step 2:
[0531] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. After using an image edge detection algorithm to remove noise and optimize the resolution, the individual images are stitched together to generate a complete image.
[0532] Input: 360 degree scanned image data.
[0533] Output: Preprocessed whole-body image data.
[0534] Step 3:
[0535] The terminal compresses the pre-processed image data using the JPEG compression algorithm and sends it to the server via the HTTP protocol. Compressing the image data improves transfer speed and reduces the load on the network bandwidth.
[0536] Input: Preprocessed whole-body image data.
[0537] Output: Compressed image data.
[0538] Step 4:
[0539] The server uses a generative AI model (e.g., a deep learning-based 3D reconstruction model) based on the received image data to generate a detailed 3D avatar of the user, analyzes the image, and generates a 3D mesh and texture.
[0540] Input: Compressed whole-body image data.
[0541] Output: 3D avatar data (mesh data, texture data).
[0542] Step 5:
[0543] The server packages the generated 3D avatar data in JSON format and sends it to the user's device via a RESTful API, which ensures structured and efficient data transfer.
[0544] Input: 3D avatar data (mesh data, texture data).
[0545] Output: 3D avatar data packaged in JSON format.
[0546] Step 6:
[0547] Users wear VR devices (such as a head-mounted display or a motion capture suit) and their movements are captured in real time. Users wear VR goggles and the motion capture suit's sensors are attached to their bodies, allowing the user's movements to be captured in detail.
[0548] Input: User's physical movements.
[0549] Output: Real-time operating data.
[0550] Step 7:
[0551] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received in real time via Bluetooth or Wi-Fi.
[0552] Input: Operational data from each sensor.
[0553] Output: Operational data stored in a local database.
[0554] Step 8:
[0555] The device analyzes the acquired motion data, converts it into quaternion format, compresses it using a compression algorithm such as Zlib, and transmits it to the server in real time using the WebSocket protocol.
[0556] Input: Operational data stored in a local database.
[0557] Output: Compressed quaternion format motion data.
[0558] Step 9:
[0559] The server processes the motion data received from the user using a real-time motion engine and reproduces the movements of a 3D avatar in the virtual space. Motion data from other participants is also integrated and synchronized, and the motion data of all users is reflected in the virtual space.
[0560] Input: Compressed quaternion format motion data.
[0561] Output: 3D avatar movement data reproduced in virtual space.
[0562] Step 10:
[0563] The device captures the user's voice data in real time from the microphone, encodes it using a voice codec such as Opus, and sends it to the server. The voice data is divided into packets with timestamps.
[0564] Input: User's voice data.
[0565] Output: The encoded audio data.
[0566] Step 11:
[0567] The server transfers the user's encoded voice data to the other participants' terminals and plays it in real time. The other participants' voice data is similarly processed and played on the user's terminal.
[0568] Input: Encoded audio data.
[0569] Output: Audio data for playback.
[0570] Prompt Sentence Examples
[0571] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[0572] (Application example 1)
[0573] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0574] One drawback of modern online shopping is that users cannot actually try on products. This often results in purchased products that do not meet actual expectations, resulting in a poor user experience. There is also a demand for a richer shopping experience that does not rely on physical stores. Furthermore, there is a lack of a way for users in different locations to communicate with each other in real time while shopping.
[0575] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0576] In this invention, the server includes: means for scanning the user's entire body 360 degrees; means for preprocessing the acquired image data and transmitting it to the server; means for generating a 3D avatar based on the received image data using a generative AI model; means for transmitting the generated 3D avatar to the user's device; means for using a motion capture device to acquire user movement data; means for transmitting the acquired movement data to the server in real time; means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data; means for synchronizing movement and position information between users in the virtual space; means for transmitting and receiving audio data for users to communicate in the virtual space; and means for providing a system in a virtual store that allows users to try on products online using a 3D avatar. This allows users to enjoy shopping while trying on products in real time from the comfort of their own home and communicating with other users.
[0577] A "means for scanning the entire body 360 degrees" is a device or system that can photograph the user's entire body from each angle and obtain image data from all directions.
[0578] "Preprocessing" refers to processing the acquired image data, such as noise removal, resolution adjustment, and image combination.
[0579] A "generative AI model" is a model that uses an artificial intelligence algorithm to generate a 3D avatar based on acquired image data.
[0580] A "motion capture device" is a device that captures a user's movements in real time and acquires them as movement data.
[0581] A "virtual space" is a computer-generated interactive 3D space in which users can perform various actions through their avatars.
[0582] "Means for synchronizing movement and position information" refers to technologies and systems for matching the movements and positions of multiple users in a virtual space in real time.
[0583] "Means for transmitting and receiving voice data" refers to a device and system for capturing a user's voice and transmitting it to other users via the Internet.
[0584] "Means provided within a virtual store" refers to a system that provides an environment on an online shopping platform where users can virtually try on products using 3D avatars.
[0585] This system scans a user's entire body in 360 degrees, generates a 3D avatar using a generative AI model, and then recreates the avatar in a virtual space in conjunction with motion capture data. This system allows users to try on products in the virtual space and communicate with other users in real time.
[0586] First, the user scans their entire body using a smartphone or dedicated camera device. In this step, the user rotates the device to capture images from all directions (360 degrees). The device preprocesses the captured image data, removing noise and adjusting the resolution, before sending it to the server. The OpenCV library is used for preprocessing.
[0587] The server uses a generative AI model to create a 3D avatar based on the received image data. This AI model uses OpenAI's API to analyze the image data and generate a detailed 3D avatar of the user. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the user's device.
[0588] Next, the user wears a motion capture device (e.g., a VR device or motion capture suit) and their movement data is acquired in real time. The device acquires the movement data from various sensors (e.g., gyro sensor, accelerometer, position sensor), converts it into an appropriate format, compresses it, and sends it to the server. During this process, appropriate libraries (e.g., NumPy or the Python standard library) are used to format and compress the data.
[0589] The server then recreates the movements of the 3D avatar in the virtual space based on the received motion data, and synchronizes the motion and position information with the avatars of other users, allowing all users to experience the sensation of operating the avatar in real time in the same virtual space.
[0590] Furthermore, users can send and receive voice data to communicate within the virtual space. The device picks up the user's voice data from the microphone and sends it to the server. The server then forwards the received voice data to the devices of the other participants, and the voice is played on each device.
[0591] As an example, consider a scenario in which a user tries on products in a virtual store in a virtual space. The user uses a smartphone or VR device to scan their entire body and generate a 3D avatar, then logs into the virtual store. At this time, the user might input a prompt sentence like the following into the generative AI model:
[0592] "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3"
[0593] Users can then view and rotate the products using their own 3D avatar as they try them on in a virtual space, and can also interact with other users in real time to share their thoughts on the fit and design of the products.
[0594] As described above, the present invention allows users to transcend physical constraints and enjoy a richer shopping experience in a virtual space.
[0595] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0596] Step 1:
[0597] The user launches a dedicated application and scans their entire body from all directions using a smartphone or dedicated camera device, acquiring multiple image data. This provides a full-body image of the user as input.
[0598] Step 2:
[0599] The device preprocesses the acquired image data. This preprocessing includes noise removal, resolution adjustment, and image merging. The OpenCV library is used for noise removal, resolution adjustment, and image merging. The preprocessed image data is obtained as the output.
[0600] Step 3:
[0601] The device compresses the preprocessed image data and sends it to the server. The compression uses an appropriate library to efficiently transmit the data. The preprocessed and compressed image data is sent as input to the server.
[0602] Step 4:
[0603] The server generates a 3D avatar for the user using a generative AI model based on the received image data. The generative AI model uses OpenAI's API to analyze the image and generate a detailed 3D avatar. The generated 3D avatar is obtained as the output.
[0604] Step 5:
[0605] The server sends the generated 3D avatar data to the user's device using a standard data transfer protocol. The user's device receives the 3D avatar data.
[0606] Step 6:
[0607] The user wears a motion capture device (e.g., a VR device or a motion capture suit) and motion data is acquired. The motion capture device acquires motion data in real time using gyro sensors, accelerometers, and position sensors. The motion data is obtained as output.
[0608] Step 7:
[0609] The device analyzes the motion data acquired from each sensor and converts it into an appropriate format. This process involves shaping the data and converting it into a compatible format. The analyzed motion data is then output.
[0610] Step 8:
[0611] The terminal compresses the converted motion data and transmits it to the server in real time using a standard data compression algorithm. The compressed motion data is then transmitted as input to the server.
[0612] Step 9:
[0613] The server reproduces the 3D avatar's movements in the virtual space based on the received movement data. This is done using a simulation engine, which reproduces real-time movements in the virtual space. The reproduced movement data is obtained as output.
[0614] Step 10:
[0615] The server synchronizes the motion and location information with the avatars of other users, so that all users can see the synchronized motion in the same virtual space. The synchronized motion and location information is obtained as output.
[0616] Step 11:
[0617] Users send and receive voice data to communicate within the virtual space. The device captures the voice data from the microphone and sends it to the server. The server then transfers the received voice data to other users' devices and plays it back in real time. The sent and received voice data is obtained as output.
[0618] Step 12:
[0619] Users use their own 3D avatar to try on products in a virtual store. Specifically, they input the following prompt into the generative AI model: "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3." This allows users to try on products online and communicate with other users in real time.
[0620] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0621] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and the ability to recognize and reflect the user's emotions using an emotion engine. Below, the program processing of the entire system and specific examples are explained in natural language.
[0622] Full body scan and image data transmission
[0623] The user launches a dedicated application on their smartphone and switches to full-body scan mode, using the camera to scan the entire body 360 degrees and capture multiple image data.
[0624] The terminal preprocesses the acquired image data, for example, by removing noise, adjusting resolution, combining images, etc. The preprocessed image data is compressed and sent to the server.
[0625] Creating 3D avatars with generative AI
[0626] The server analyzes the received full-body image data and generates a detailed 3D avatar of the user using a generative AI model. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the device.
[0627] Obtaining movement data using motion capture
[0628] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0629] The device acquires real-time motion data from the VR device, specifically, data from the gyro sensor, accelerometer, and position sensor.
[0630] The device analyzes the acquired motion data, converts it into an appropriate format, compresses it, and transmits it to the server in real time.
[0631] Projection into virtual space and avatar synchronization
[0632] The server reproduces the user's 3D avatar's movements in the virtual space based on the movement data sent by the user. It also receives the movement data of other participants and synchronizes the movement and position information of all users in the virtual space.
[0633] Emotion recognition and reflection using emotion engine
[0634] The device collects the user's facial expression data and voice data and analyzes this data using an emotion engine.
[0635] The terminal compresses the analysis results and transmits them to the server in real time.
[0636] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space.
[0637] Communication Features
[0638] Users communicate with other users visually and audibly in the virtual space, for example, by talking, exchanging glances, and making gestures.
[0639] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0640] The server receives the audio data from other participants and sends it to the appropriate device, where it is played through speakers or a headset.
[0641] Specific examples
[0642] For example, consider a situation where parents and children living far apart use this system to communicate in a virtual space, with the parent wearing a VR device at home and the child wearing one as well.
[0643] The child's device captures the child's behavior and emotional data in real time and sends it to the server, while the parent's device simultaneously captures the parent's behavior and emotional data and sends it to the server.
[0644] The server integrates this data and recreates and synchronizes it in the virtual space, allowing parents to interact with their child's 3D avatar in real time visually and audibly, and even feel their child's emotions.
[0645] Furthermore, when a parent smiles with joy, that expression is reflected in the child's avatar in real time, enabling natural communication even from a distance. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[0646] The processing flow will be explained below.
[0647] Step 1:
[0648] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[0649] The user uses a camera to scan the entire body 360 degrees and acquires multiple image data.
[0650] Step 2:
[0651] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[0652] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[0653] Step 3:
[0654] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[0655] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0656] Step 4:
[0657] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0658] Step 5:
[0659] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[0660] Step 6:
[0661] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[0662] Step 7:
[0663] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[0664] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[0665] Step 8:
[0666] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[0667] Step 9:
[0668] The device uses an emotion engine to recognize the user's emotions and collects the user's facial expression data and voice data.
[0669] The device analyzes this data to identify the user's emotions.
[0670] Step 10:
[0671] The device compresses the analysis results and sends them to the server in real time.
[0672] Step 11:
[0673] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in the virtual space.
[0674] Step 12:
[0675] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[0676] Step 13:
[0677] The server receives the voice data of the other participants and transmits it to the user's terminal.
[0678] The audio data received by the device is played through speakers or a headset.
[0679] As described above, the system of the present invention realizes high-quality communication through the creation of a 3D avatar using generative AI from a full-body scan of the user, motion capture, emotion recognition using an emotion engine, and real-time reproduction of facial expressions and movements in a virtual space.
[0680] Example 2
[0681] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0682] Current communication technology makes it difficult for people living far apart to interact in a natural way. In particular, the real-time sharing of facial expressions, movements, and emotions has not been fully realized. This has led to a decline in the quality of communication in virtual reality spaces and a limited user experience.
[0683] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0684] In this invention, the server includes a means for scanning the user's entire body in 360 degrees, a means for preprocessing the acquired image data and sending it to the server, and a means for generating a 3D avatar based on the received image data using a machine learning model, which allows the user to move in real time and communicate naturally with the user, reflecting their emotions.
[0685] "User" refers to a person who uses the system to perform full-body scans, acquire movement data, and communicate within the virtual space.
[0686] "Means for scanning the whole body 360 degrees" refers to a device that captures the user's whole body from all directions and a method of operating the device.
[0687] "Preprocessing" refers to the step of processing the acquired image data by noise removal, resolution adjustment, image combination, etc.
[0688] A "server" refers to a computer system that processes data sent from users, generates avatars in virtual space, and synchronizes movement data.
[0689] A "machine learning model" refers to an algorithm that analyzes input data, recognizes patterns, and generates a 3D avatar.
[0690] A "3D avatar" refers to a three-dimensional virtual character that reproduces the user's appearance and movements.
[0691] "Device" refers to a device required for a user to use the system, such as a smartphone or VR device.
[0692] "Motion detection device" refers to a sensor device or device for capturing a user's physical movements in real time.
[0693] "Real-time" refers to a short time range in which user actions and data are reflected almost instantaneously.
[0694] "Virtual space" refers to a three-dimensional digital environment generated by a computer.
[0695] "Voice Data" means data that records and transmits a user's voice in digital form.
[0696] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and voice data to recognize and reflect their emotional state.
[0697] The system of the present invention is designed to realize high-quality communication in a virtual space. Specifically, it includes functions such as full-body scanning of the user, 3D avatar generation using a machine learning model, real-time motion data acquisition using a motion detection device, avatar synchronization within the virtual space, communication using voice data, and emotion recognition and reflection using an emotion engine.
[0698] First, the user scans their entire body in 360 degrees using a dedicated smartphone application. The user's full-body image data is pre-processed on the device to remove noise, adjust resolution, and combine images, then compressed and sent to a cloud-based server. The server analyzes the received image data and generates a detailed 3D avatar of the user using stable diffusion and other generative AI models. The generated 3D avatar data is then sent back to the device and provided to the user.
[0699] Next, the user puts on a VR device such as a head-mounted display or motion capture suit and launches the corresponding dedicated application. At this time, the device collects data from the VR device in real time using the gyro sensor, accelerometer, and position sensor, converts it into an appropriate format, compresses it, and sends it to the server. This allows the server to instantly reproduce the movements of the user's 3D avatar in the virtual space.
[0700] Furthermore, the system incorporates an emotion engine that can recognize and reflect the user's emotional state by collecting and analyzing facial and voice data, allowing the user's 3D avatar to express emotions in a more natural way.
[0701] For example, when parents and children living far apart use this system to interact in a virtual space, they both wear VR devices at home and launch a dedicated application. The child's device captures movement and emotional data in real time and sends it to the server. At the same time, the parent's device also acquires the parent's movement and emotional data and sends it to the server. The server integrates this data and reproduces and synchronizes it in the virtual space, allowing the parent and child to interact visually and audibly in real time. Furthermore, when the parent smiles with joy, that expression is reflected in the child's avatar in real time.
[0702] In this way, this system provides an advanced communication experience that transcends physical distance.
[0703] Example prompt sentence:
[0704] "Use a generative AI model to generate a 3D avatar from the user's full-body image data. The specific process includes noise removal, resolution adjustment, and avatar generation using the generative AI model."
[0705] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0706] Program processing steps
[0707] Step 1: Full body scan and image data transmission
[0708] The user launches a dedicated application on their smartphone and switches to "full body scan" mode.
[0709] Specific operation: Use the front camera to scan your entire body 360 degrees and obtain multiple image data.
[0710] Input: User's full-body image data
[0711] Once the device receives the captured image data, it applies a noise reduction filter to remove unnecessary information, then adjusts the image resolution to an appropriate size, and finally merges the images to create a single full-body image.
[0712] Output: Preprocessed whole-body image data
[0713] The device compresses the preprocessed image data and sends it to a server on the cloud.
[0714] Specific operation: The terminal automatically compresses and transmits data in parallel.
[0715] Step 2: Creating a 3D avatar using generative AI
[0716] The server sends the received preprocessed image data to the AI processing module.
[0717] Input: Preprocessed whole-body image data
[0718] The server analyzes the image using a generative AI model such as Stable Diffusion to generate a 3D avatar, which includes mesh and texture data.
[0719] Output: Generated 3D avatar data (mesh data, texture data)
[0720] The server sends the generated 3D avatar data to the user's device.
[0721] How it works: The server analyzes the image data and generates and transmits a complete 3D avatar within seconds.
[0722] Step 3: Obtaining movement data using motion capture
[0723] The user puts on a head-mounted display or motion capture suit and launches the corresponding dedicated application.
[0724] Specific actions: The user puts on the VR device and starts moving in the virtual space.
[0725] Input: User behavior data
[0726] The terminal collects data from the gyro sensor, accelerometer, and position sensor in real time from the VR device and converts it into a suitable format.
[0727] Output: Motion data converted into a suitable format
[0728] The terminal compresses the converted motion data and transmits it to the server in real time.
[0729] Specific operation: The device collects and transmits data sequentially in accordance with the user's movements.
[0730] Step 4: Projection into virtual space and avatar synchronization
[0731] The server reproduces the movements of the user's 3D avatar in the virtual space based on the motion data sent by the user.
[0732] Input: Operation data
[0733] Output: Avatar movements reproduced in virtual space
[0734] Specific actions: A user's movements in the virtual space or hand movements are synchronized with other users in real time.
[0735] Step 5: Emotion recognition and reflection using the emotion engine
[0736] The terminal collects facial expression data and voice data of the user.
[0737] Input: Facial expression data, voice data
[0738] The terminal uses an emotion engine to analyze this data and recognize the user's emotional state.
[0739] Output: Emotion recognition result
[0740] The terminal compresses the analysis results and transmits them to the server in real time.
[0741] Specific operation: When a user smiles, the facial expression is immediately analyzed by the emotion engine and reflected on the server.
[0742] Step 6: Communication Functions
[0743] Users can communicate visually and audibly with other users in the virtual space.
[0744] Specific actions: The user speaks or gestures.
[0745] Input: Voice data, gesture data
[0746] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0747] Output: Voice communication data with other users
[0748] The server receives the voice data of the other participants and sends it to the appropriate terminal, where the voice data is played through speakers or headsets.
[0749] Specific operation: The user's voice is transmitted to other users in real time, enabling natural conversation.
[0750] (Application example 2)
[0751] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0752] Conventional communication systems in virtual spaces reproduce the movements and communication of an avatar based solely on the user's motion and voice data, and lack the ability to reflect the user's emotions in real time, making it difficult to achieve a natural communication experience. Furthermore, even in fitting experiences in virtual stores in the real world, there was an issue of not being able to grasp the user's emotions and provide optimal advice and suggestions.
[0753] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0754] In this invention, the server includes means for scanning the user's entire body 360 degrees, means for preprocessing the acquired image data and transmitting it to the server, means for generating a 3D avatar based on the received image data using a generative AI model, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to acquire user movement data, means for transmitting the acquired movement data to the server in real time, means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data, means for synchronizing movement and position information between users in the virtual space, means for transmitting and receiving voice data for users to communicate in the virtual space, means for capturing user facial expression data and analyzing emotions using an emotion engine, and means for reflecting the user's 3D avatar's facial expressions and movements in the virtual space in real time based on the analyzed emotion data. This enables users to communicate in the virtual space and try on clothes in a virtual store in a natural and high-quality manner.
[0755] A "means for scanning the entire body 360 degrees" is a device or method for capturing images of a user's entire body from various angles and acquiring a series of image data.
[0756] "Image data preprocessing means" refers to a method or device that performs processing such as noise removal and resolution adjustment on acquired image data to improve the quality of the data.
[0757] The "means for transmitting to the server" refers to a device or method for transmitting the pre-processed data to the server over a network.
[0758] A "generative AI model" is an artificial intelligence model that generates a detailed 3D avatar of a user based on received image data.
[0759] A "3D avatar" is a three-dimensional digital model that mimics the user's entire body.
[0760] "Motion capture device" is a general term for devices and sensors used to capture a user's movements in real time.
[0761] "Motion data" refers to data that includes information about a user's movements acquired by a motion capture device.
[0762] The "means for transmitting to the server in real time" refers to a device or method for instantly transmitting acquired motion data to the server.
[0763] A "virtual space" is a three-dimensional digital space generated using computer technology.
[0764] "Means for reproducing movements in a virtual space" refers to devices or methods that use acquired movement data to allow a 3D avatar to move in real time within a virtual space.
[0765] The "means for synchronizing motion and position information" refers to a device or method for consistently synchronizing the motion and position information of multiple users in a virtual space.
[0766] "Means for transmitting and receiving voice data" refers to a method or device for transmitting and receiving voice data that allows a user to communicate with other users by voice in a virtual space.
[0767] "Means for capturing facial expression data" refers to a device or method for acquiring a user's facial expression in real time.
[0768] An "emotion engine" is a program or device that analyzes captured facial expression data and voice data to recognize the user's emotional state.
[0769] The "means for analyzing emotions" refers to a device or method for analyzing data acquired using an emotion engine and identifying the user's emotions.
[0770] "Means for reflecting the facial expressions and movements of a 3D avatar in real time within a virtual space" refers to a device or method for instantly changing and reflecting the facial expressions and movements of a 3D avatar within a virtual space based on analyzed emotional data.
[0771] The system of the present invention includes functions such as full-body scanning of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and recognition and reflection of the user's emotions using an emotion engine.
[0772] First, the user uses the smartphone's full-body scanning function to scan their entire body in 360 degrees, acquiring multiple image data. The acquired image data undergoes preprocessing, such as noise removal and resolution adjustment, to improve the data quality before being sent to the server. This preprocessing is carried out using OpenCV, an open-source image processing library.
[0773] The server uses a generative AI model based on the received image data to generate a detailed 3D avatar of the user. The mesh and texture data of the generated 3D avatar are then sent to the user's device, where it can be displayed on a VR device or similar.
[0774] Next, the user puts on a VR device (such as a head-mounted display or motion capture suit) and movement data is acquired. This movement data is collected in real time using gyro sensors, acceleration sensors, position sensors, etc. The acquired movement data is sent to a server and reproduced in real time as the movement of a 3D avatar in the virtual space.
[0775] This system allows multiple users to participate simultaneously and synchronizes their movements and positions in the virtual space, allowing for natural interactions between users.
[0776] Furthermore, the system captures the user's facial expression and voice data and analyzes them using an emotion engine to identify the user's emotions. This analyzed emotion data is reflected in the facial expressions and movements of the user's 3D avatar in real time. The emotion engine uses an emotion recognition model using TensorFlow.
[0777] A concrete example of this system is a virtual store try-on experience. By having a 3D avatar instantly try on the clothes the user has selected, the experience becomes as if they are actually trying them on. In addition, by using an emotion engine, the system can understand the user's reactions and provide optimal fashion advice.
[0778] Example prompt sentence:
[0779] "I would like to demonstrate a real-time try-on experience using a 3D avatar. The 3D avatar will instantly wear the clothes the user has chosen, and an emotion engine will be used to analyze the user's reaction and display the next recommended item."
[0780] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0781] Step 1:
[0782] Full body scan of the user
[0783] The user activates the full-body scan mode on their smartphone and uses the camera to scan their entire body 360 degrees. This acquires multiple image data. The input is the smartphone camera image, and the output is the acquired multiple image data. Specifically, the user rotates the camera to capture images of their entire body.
[0784] Step 2:
[0785] Image data preprocessing
[0786] The device preprocesses the acquired image data, performing noise removal, resolution adjustment, image merging, etc. The input is the acquired image data, and the output is the preprocessed image data. This processing uses the OpenCV library. Specifically, a noise removal filter is applied, and each image is merged to generate a high-resolution whole-body image.
[0787] Step 3:
[0788] Sending image data to the server
[0789] The terminal compresses the preprocessed image data and sends it to the server. The input is the preprocessed image data, and the output is the compressed data. The HTTP protocol is used to send the data. Specifically, the data is converted into a byte sequence and sent using an HTTP request.
[0790] Step 4:
[0791] 3D avatar generation
[0792] The server generates a 3D avatar using a generative AI model based on the received image data. The input is compressed image data, and the output is 3D avatar data (mesh data, texture data, etc.). Specifically, the generative AI model analyzes the image data and constructs a 3D model.
[0793] Step 5:
[0794] Sending a 3D avatar to your device
[0795] The server sends the generated 3D avatar data to the device. The input is the 3D avatar data, and the output is the avatar data sent to the device. Specifically, the data is sent in real time using WebSocket communication.
[0796] Step 6:
[0797] Obtaining movement data using motion capture
[0798] The user puts on a VR device (a head-mounted display or motion capture suit) and launches a dedicated application. The input is the user's movements and the device's sensor data, and the output is movement data. Specifically, data is collected from the gyro sensor, acceleration sensor, and position sensor.
[0799] Step 7:
[0800] Sending operation data to the server
[0801] The terminal analyzes the acquired motion data, converts it into an appropriate format, compresses it, and sends it to the server. The input is the acquired motion data, and the output is compressed motion data. Specifically, the data format is converted and the data is made smaller using a compression algorithm.
[0802] Step 8:
[0803] Reproducing movements in virtual space
[0804] The server reproduces the movements of the user's 3D avatar in the virtual space based on the movement data sent by the user. The input is compressed movement data, and the output is a 3D avatar that moves in the virtual space. Specifically, the server analyzes the data, calculates the position and posture in the virtual space, and reflects them in the 3D avatar.
[0805] Step 9:
[0806] Synchronization of movement and position information in virtual space
[0807] The server also receives other users' motion data and synchronizes the motion and location information of all users in the virtual space. The input is motion data from multiple users, and the output is a synchronized virtual space. Specifically, all user data is centrally managed and adjusted in real time.
[0808] Step 10:
[0809] Facial expression data capture and analysis
[0810] The device captures the user's facial expression data and analyzes the emotions using an emotion engine. The input is the user's facial expression data, and the output is the emotion data that is the analysis result. Specifically, the device analyzes the facial image captured by the camera and identifies the emotional state (joy, anger, sadness, etc.).
[0811] Step 11:
[0812] Sending emotion data to the server
[0813] The device compresses the analyzed emotion data and sends it to the server in real time. The input is the analyzed emotion data, and the output is the compressed emotion data. Specifically, the emotion data is converted into a byte sequence and sent using an HTTP request.
[0814] Step 12:
[0815] Reflecting facial expressions and movements in virtual space
[0816] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space. The input is emotion data, and the output is a 3D avatar that reflects the emotion. Specifically, the emotion data is converted into the avatar's facial motion and reflected in real time.
[0817] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0818] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0819] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0820] [Third embodiment]
[0821] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0822] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0823] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0824] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0825] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0826] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0827] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0828] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0829] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0830] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0831] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0832] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0833] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data using motion capture, avatar synchronization in a virtual space, and real-time voice communication. Below, the program processing of the entire system and specific examples are explained in natural language.
[0834] Full body scan and image data transmission
[0835] The user starts a dedicated application and scans their entire body in 360 degrees using a smartphone or dedicated camera device, which captures multiple image data.
[0836] The device preprocesses the acquired image data, including noise removal, resolution adjustment, and image merging.
[0837] The terminal compresses the preprocessed image data and sends it to the server, thereby distributing the processing load.
[0838] Creating 3D avatars with generative AI
[0839] The server uses a generative AI model based on the received full-body image data to create a detailed 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[0840] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0841] Obtaining movement data using motion capture
[0842] The user wears a VR device (such as a head-mounted display or a motion capture suit), which captures the user's movements in real time.
[0843] The device acquires operational data from each sensor (e.g., data from the gyro sensor, acceleration sensor, and position sensor).
[0844] The device analyzes the acquired motion data, converts it into a suitable format, compresses it, and transmits it to the server in real time.
[0845] Projection into virtual space and avatar synchronization
[0846] The server processes the motion data received from the user and reproduces the avatar's motion in the virtual space. It also integrates motion data from other participants (family and friends) and synchronizes the motion and location information of all users.
[0847] Communication Features
[0848] Users communicate with other users in the virtual space visually and audibly, for example by talking or waving.
[0849] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[0850] The server transfers the received audio data to the terminals of the other participants, and the audio is played back at each terminal.
[0851] Specific examples
[0852] For example, consider the case where grandparents who live far away use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear similar devices at their homes.
[0853] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server. At the same time, the grandparent's device also acquires the grandparent's motion data and sends it to the server.
[0854] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk with their grandchildren as if they were there in the virtual space.
[0855] The devices also exchange voice data between users in real time, providing a natural conversation experience. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[0856] The processing flow will be explained below.
[0857] Step 1:
[0858] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[0859] The user uses a camera to scan their entire body 360 degrees, obtaining multiple image data.
[0860] Step 2:
[0861] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[0862] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[0863] Step 3:
[0864] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[0865] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[0866] Step 4:
[0867] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[0868] Step 5:
[0869] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[0870] Step 6:
[0871] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[0872] Step 7:
[0873] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[0874] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[0875] Step 8:
[0876] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[0877] Step 9:
[0878] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[0879] Step 10:
[0880] The server receives the voice data of the other participants and transmits it to the user's terminal.
[0881] The audio data received by the device is played through speakers or a headset.
[0882] As described above, the system of the present invention realizes high-quality real-time communication that transcends physical distance.
[0883] Example 1
[0884] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0885] In modern society, users need to communicate with each other in real time, regardless of physical distance. In particular, real-time data communication, both visual and audio, is essential for immersive communication with family and friends in distant locations. However, current systems have difficulty accurately capturing a user's full-body movements and voice in real time and reproducing them in a virtual space. This limits the communication experience and leads to problems such as connection delays and instability.
[0886] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0887] In this invention, the server includes means for performing a 360-degree full-body scan, means for preprocessing the image data and transmitting it to the server, means for generating a 3D avatar using a generative AI model based on the received image data, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to capture user motion data, means for analyzing the captured motion data and converting it into a suitable format, means for transmitting the converted motion data to the server in real time, means for reproducing the motion of the 3D avatar in a virtual space based on the real-time motion data, means for synchronizing motion and position information between users in the virtual space, and means for transmitting and receiving audio data for users to communicate in the virtual space. This allows the user's full-body motions and voices to be reflected in the virtual space in real time, enabling a high-quality communication experience.
[0888] A "means for scanning the entire body 360 degrees" is a device or software that photographs the user's entire body from various angles and collects data from all directions.
[0889] "Means for preprocessing image data and sending it to the server" refers to devices or software that process the acquired image data by removing noise, adjusting resolution, combining images, etc., and then compressing and sending it to the server.
[0890] "Means for generating a 3D avatar using a generative AI model" refers to modules or software that use artificial intelligence technology to generate a three-dimensional avatar based on received image data.
[0891] "Means for transmitting the 3D avatar to the user's terminal" refers to a communication protocol or device for transmitting the generated 3D avatar as data to the end user's terminal.
[0892] "Means using a motion capture device" means a device or software for detecting and capturing a user's physical movements in real time.
[0893] The "means for analyzing the motion data and converting it into a suitable format" refers to a device or software for analyzing the acquired motion data and converting it into a suitable data format for reproduction in a virtual space.
[0894] The "means for transmitting the converted motion data to the server in real time" refers to a device or software for compressing the motion data converted into an appropriate format and transmitting it to the server in real time without any loss.
[0895] "Means for reproducing the movements of a 3D avatar in a virtual space" refers to an engine or software that faithfully reproduces the movements of a 3D avatar in a virtual space based on received movement data.
[0896] "Means for synchronizing the movements and location information of users in a virtual space" refers to a system or software that consistently reflects the movements and location information of multiple users in real time, enabling interaction within a virtual space.
[0897] "Means for transmitting and receiving voice data" refers to a communication system or software for capturing a user's voice in real time and transmitting and playing it back to other users.
[0898] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition and transmission of real-time motion data through motion capture, synchronization of the motion data in a virtual space, and voice communication. A specific embodiment of the entire system will be described below.
[0899] Full body scan and image data transmission
[0900] The user launches a dedicated application on a smartphone or dedicated camera device and scans their entire body in 360 degrees, obtaining multiple image data. Specifically, the user rotates the smartphone to capture their entire body, and the app instructs them on the direction and angle of the shot.
[0901] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. Noise removal is performed using edge detection algorithms, and resolution is optimized. The individual images are then stitched together to generate the entire image.
[0902] The pre-processed image data is compressed by the terminal and sent to the server via HTTP protocol, using the JPEG compression algorithm for efficient uploading to the server.
[0903] Creating 3D avatars with generative AI
[0904] The server uses a generative AI model (e.g., a Deep Learning-based 3D Reconstruction model) based on the received full-body image data to generate a 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[0905] The generated 3D avatar data (mesh data, texture data, etc.) is sent from the server to the user's device. The data is packaged in JSON format and sent via a RESTful API.
[0906] Obtaining movement data using motion capture
[0907] The user wears VR devices (such as a head-mounted display and a motion capture suit) and captures their movements in real time. The user wears VR goggles and sensors from the motion capture suit, which are attached to the user's body, and the sensors capture the user's body movements.
[0908] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received via Bluetooth or Wi-Fi.
[0909] The acquired motion data is analyzed by the device and converted into quaternion format, which is then compressed using a compression algorithm such as Zlib and sent to the server in real time using the WebSocket protocol.
[0910] Projection into virtual space and avatar synchronization
[0911] The server processes the motion data received from the user and reproduces the movements of a 3D avatar in the virtual space. All user motion data is integrated using a real-time motion engine. The data is then reflected in the virtual space using a game engine such as Unity.
[0912] Communication Features
[0913] Users can communicate with other users visually and audibly in the virtual space by talking to other avatars or waving to them. Facial tracking can also be used to reflect facial expressions on the avatars in real time.
[0914] The device captures the user's voice data from the microphone, encodes it using an audio codec such as Opus, and sends it to the server. The voice data is then transferred in real time to other devices and played back through the speakers.
[0915] Specific examples
[0916] For example, grandparents who live far away can use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear VR devices as well.
[0917] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server, while the grandparent's device simultaneously captures the grandparent's motion data and sends it to the server.
[0918] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk as if they were actually there with their grandchildren. The devices also exchange voice data between users in real time, providing a natural conversation experience.
[0919] Prompt Sentence Examples
[0920] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[0921] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0922] Step 1:
[0923] Users launch a dedicated application on their smartphone or dedicated camera device and perform a 360-degree full-body scan. Users rotate their smartphone to capture their entire body, adjusting the angle and direction of the shot according to the application's instructions.
[0924] Input: Multiple image data captured by a smartphone or dedicated camera.
[0925] Output: 360-degree full-body scan image data.
[0926] Step 2:
[0927] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. After using an image edge detection algorithm to remove noise and optimize the resolution, the individual images are stitched together to generate a complete image.
[0928] Input: 360 degree scanned image data.
[0929] Output: Preprocessed whole-body image data.
[0930] Step 3:
[0931] The terminal compresses the pre-processed image data using the JPEG compression algorithm and sends it to the server via the HTTP protocol. Compressing the image data improves transfer speed and reduces the load on the network bandwidth.
[0932] Input: Preprocessed whole-body image data.
[0933] Output: Compressed image data.
[0934] Step 4:
[0935] The server uses a generative AI model (e.g., a deep learning-based 3D reconstruction model) based on the received image data to generate a detailed 3D avatar of the user, analyzes the image, and generates a 3D mesh and texture.
[0936] Input: Compressed whole-body image data.
[0937] Output: 3D avatar data (mesh data, texture data).
[0938] Step 5:
[0939] The server packages the generated 3D avatar data in JSON format and sends it to the user's device via a RESTful API, which ensures structured and efficient data transfer.
[0940] Input: 3D avatar data (mesh data, texture data).
[0941] Output: 3D avatar data packaged in JSON format.
[0942] Step 6:
[0943] Users wear VR devices (such as a head-mounted display or a motion capture suit) and their movements are captured in real time. Users wear VR goggles and the motion capture suit's sensors are attached to their bodies, allowing the user's movements to be captured in detail.
[0944] Input: User's physical movements.
[0945] Output: Real-time operating data.
[0946] Step 7:
[0947] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received in real time via Bluetooth or Wi-Fi.
[0948] Input: Operational data from each sensor.
[0949] Output: Operational data stored in a local database.
[0950] Step 8:
[0951] The device analyzes the acquired motion data, converts it into quaternion format, compresses it using a compression algorithm such as Zlib, and transmits it to the server in real time using the WebSocket protocol.
[0952] Input: Operational data stored in a local database.
[0953] Output: Compressed quaternion format motion data.
[0954] Step 9:
[0955] The server processes the motion data received from the user using a real-time motion engine and reproduces the movements of a 3D avatar in the virtual space. Motion data from other participants is also integrated and synchronized, and the motion data of all users is reflected in the virtual space.
[0956] Input: Compressed quaternion format motion data.
[0957] Output: 3D avatar movement data reproduced in virtual space.
[0958] Step 10:
[0959] The device captures the user's voice data in real time from the microphone, encodes it using a voice codec such as Opus, and sends it to the server. The voice data is divided into packets with timestamps.
[0960] Input: User's voice data.
[0961] Output: The encoded audio data.
[0962] Step 11:
[0963] The server transfers the user's encoded voice data to the other participants' terminals and plays it in real time. The other participants' voice data is similarly processed and played on the user's terminal.
[0964] Input: Encoded audio data.
[0965] Output: Audio data for playback.
[0966] Prompt Sentence Examples
[0967] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[0968] (Application example 1)
[0969] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0970] One drawback of modern online shopping is that users cannot actually try on products. This often results in purchased products that do not meet actual expectations, resulting in a poor user experience. There is also a demand for a richer shopping experience that does not rely on physical stores. Furthermore, there is a lack of a way for users in different locations to communicate with each other in real time while shopping.
[0971] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0972] In this invention, the server includes: means for scanning the user's entire body 360 degrees; means for preprocessing the acquired image data and transmitting it to the server; means for generating a 3D avatar based on the received image data using a generative AI model; means for transmitting the generated 3D avatar to the user's device; means for using a motion capture device to acquire user movement data; means for transmitting the acquired movement data to the server in real time; means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data; means for synchronizing movement and position information between users in the virtual space; means for transmitting and receiving audio data for users to communicate in the virtual space; and means for providing a system in a virtual store that allows users to try on products online using a 3D avatar. This allows users to enjoy shopping while trying on products in real time from the comfort of their own home and communicating with other users.
[0973] A "means for scanning the entire body 360 degrees" is a device or system that can photograph the user's entire body from each angle and obtain image data from all directions.
[0974] "Preprocessing" refers to processing the acquired image data, such as noise removal, resolution adjustment, and image combination.
[0975] A "generative AI model" is a model that uses an artificial intelligence algorithm to generate a 3D avatar based on acquired image data.
[0976] A "motion capture device" is a device that captures a user's movements in real time and acquires them as movement data.
[0977] A "virtual space" is a computer-generated interactive 3D space in which users can perform various actions through their avatars.
[0978] "Means for synchronizing movement and position information" refers to technologies and systems for matching the movements and positions of multiple users in a virtual space in real time.
[0979] "Means for transmitting and receiving voice data" refers to a device and system for capturing a user's voice and transmitting it to other users via the Internet.
[0980] "Means provided within a virtual store" refers to a system that provides an environment on an online shopping platform where users can virtually try on products using 3D avatars.
[0981] This system scans a user's entire body in 360 degrees, generates a 3D avatar using a generative AI model, and then recreates the avatar in a virtual space in conjunction with motion capture data. This system allows users to try on products in the virtual space and communicate with other users in real time.
[0982] First, the user scans their entire body using a smartphone or dedicated camera device. In this step, the user rotates the device to capture images from all directions (360 degrees). The device preprocesses the captured image data, removing noise and adjusting the resolution, before sending it to the server. The OpenCV library is used for preprocessing.
[0983] The server uses a generative AI model to create a 3D avatar based on the received image data. This AI model uses OpenAI's API to analyze the image data and generate a detailed 3D avatar of the user. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the user's device.
[0984] Next, the user wears a motion capture device (e.g., a VR device or motion capture suit) and their movement data is acquired in real time. The device acquires the movement data from various sensors (e.g., gyro sensor, accelerometer, position sensor), converts it into an appropriate format, compresses it, and sends it to the server. During this process, appropriate libraries (e.g., NumPy or the Python standard library) are used to format and compress the data.
[0985] The server then recreates the movements of the 3D avatar in the virtual space based on the received motion data, and synchronizes the motion and position information with the avatars of other users, allowing all users to experience the sensation of operating the avatar in real time in the same virtual space.
[0986] Furthermore, users can send and receive voice data to communicate within the virtual space. The device picks up the user's voice data from the microphone and sends it to the server. The server then forwards the received voice data to the devices of the other participants, and the voice is played on each device.
[0987] As an example, consider a scenario in which a user tries on products in a virtual store in a virtual space. The user uses a smartphone or VR device to scan their entire body and generate a 3D avatar, then logs into the virtual store. At this time, the user might input a prompt sentence like the following into the generative AI model:
[0988] "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3"
[0989] Users can then view and rotate the products using their own 3D avatar as they try them on in a virtual space, and can also interact with other users in real time to share their thoughts on the fit and design of the products.
[0990] As described above, the present invention allows users to transcend physical constraints and enjoy a richer shopping experience in a virtual space.
[0991] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0992] Step 1:
[0993] The user launches a dedicated application and scans their entire body from all directions using a smartphone or dedicated camera device, acquiring multiple image data. This provides a full-body image of the user as input.
[0994] Step 2:
[0995] The device preprocesses the acquired image data. This preprocessing includes noise removal, resolution adjustment, and image merging. The OpenCV library is used for noise removal, resolution adjustment, and image merging. The preprocessed image data is obtained as the output.
[0996] Step 3:
[0997] The device compresses the preprocessed image data and sends it to the server. The compression uses an appropriate library to efficiently transmit the data. The preprocessed and compressed image data is sent as input to the server.
[0998] Step 4:
[0999] The server generates a 3D avatar for the user using a generative AI model based on the received image data. The generative AI model uses OpenAI's API to analyze the image and generate a detailed 3D avatar. The generated 3D avatar is obtained as the output.
[1000] Step 5:
[1001] The server sends the generated 3D avatar data to the user's device using a standard data transfer protocol. The user's device receives the 3D avatar data.
[1002] Step 6:
[1003] The user wears a motion capture device (e.g., a VR device or a motion capture suit) and motion data is acquired. The motion capture device acquires motion data in real time using gyro sensors, accelerometers, and position sensors. The motion data is obtained as output.
[1004] Step 7:
[1005] The device analyzes the motion data acquired from each sensor and converts it into an appropriate format. This process involves shaping the data and converting it into a compatible format. The analyzed motion data is then output.
[1006] Step 8:
[1007] The terminal compresses the converted motion data and transmits it to the server in real time using a standard data compression algorithm. The compressed motion data is then transmitted as input to the server.
[1008] Step 9:
[1009] The server reproduces the 3D avatar's movements in the virtual space based on the received movement data. This is done using a simulation engine, which reproduces real-time movements in the virtual space. The reproduced movement data is obtained as output.
[1010] Step 10:
[1011] The server synchronizes the motion and location information with the avatars of other users, so that all users can see the synchronized motion in the same virtual space. The synchronized motion and location information is obtained as output.
[1012] Step 11:
[1013] Users send and receive voice data to communicate within the virtual space. The device captures the voice data from the microphone and sends it to the server. The server then transfers the received voice data to other users' devices and plays it back in real time. The sent and received voice data is obtained as output.
[1014] Step 12:
[1015] Users use their own 3D avatar to try on products in a virtual store. Specifically, they input the following prompt into the generative AI model: "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3." This allows users to try on products online and communicate with other users in real time.
[1016] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1017] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and the ability to recognize and reflect the user's emotions using an emotion engine. Below, the program processing of the entire system and specific examples are explained in natural language.
[1018] Full body scan and image data transmission
[1019] The user launches a dedicated application on their smartphone and switches to full-body scan mode, using the camera to scan the entire body 360 degrees and capture multiple image data.
[1020] The terminal preprocesses the acquired image data, for example, by removing noise, adjusting resolution, combining images, etc. The preprocessed image data is compressed and sent to the server.
[1021] Creating 3D avatars with generative AI
[1022] The server analyzes the received full-body image data and generates a detailed 3D avatar of the user using a generative AI model. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the device.
[1023] Obtaining movement data using motion capture
[1024] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[1025] The device acquires real-time motion data from the VR device, specifically, data from the gyro sensor, accelerometer, and position sensor.
[1026] The device analyzes the acquired motion data, converts it into an appropriate format, compresses it, and transmits it to the server in real time.
[1027] Projection into virtual space and avatar synchronization
[1028] The server reproduces the user's 3D avatar's movements in the virtual space based on the movement data sent by the user. It also receives the movement data of other participants and synchronizes the movement and position information of all users in the virtual space.
[1029] Emotion recognition and reflection using emotion engine
[1030] The device collects the user's facial expression data and voice data and analyzes this data using an emotion engine.
[1031] The terminal compresses the analysis results and transmits them to the server in real time.
[1032] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space.
[1033] Communication Features
[1034] Users communicate with other users visually and audibly in the virtual space, for example, by talking, exchanging glances, and making gestures.
[1035] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[1036] The server receives the audio data from other participants and sends it to the appropriate device, where it is played through speakers or a headset.
[1037] Specific examples
[1038] For example, consider a situation where parents and children living far apart use this system to communicate in a virtual space, with the parent wearing a VR device at home and the child wearing one as well.
[1039] The child's device captures the child's behavior and emotional data in real time and sends it to the server, while the parent's device simultaneously captures the parent's behavior and emotional data and sends it to the server.
[1040] The server integrates this data and recreates and synchronizes it in the virtual space, allowing parents to interact with their child's 3D avatar in real time visually and audibly, and even feel their child's emotions.
[1041] Furthermore, when a parent smiles with joy, that expression is reflected in the child's avatar in real time, enabling natural communication even from a distance. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[1042] The processing flow will be explained below.
[1043] Step 1:
[1044] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[1045] The user uses a camera to scan the entire body 360 degrees and acquires multiple image data.
[1046] Step 2:
[1047] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[1048] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[1049] Step 3:
[1050] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[1051] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[1052] Step 4:
[1053] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[1054] Step 5:
[1055] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[1056] Step 6:
[1057] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[1058] Step 7:
[1059] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[1060] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[1061] Step 8:
[1062] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[1063] Step 9:
[1064] The device uses an emotion engine to recognize the user's emotions and collects the user's facial expression data and voice data.
[1065] The device analyzes this data to identify the user's emotions.
[1066] Step 10:
[1067] The device compresses the analysis results and sends them to the server in real time.
[1068] Step 11:
[1069] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in the virtual space.
[1070] Step 12:
[1071] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[1072] Step 13:
[1073] The server receives the voice data of the other participants and transmits it to the user's terminal.
[1074] The audio data received by the device is played through speakers or a headset.
[1075] As described above, the system of the present invention realizes high-quality communication through the creation of a 3D avatar using generative AI from a full-body scan of the user, motion capture, emotion recognition using an emotion engine, and real-time reproduction of facial expressions and movements in a virtual space.
[1076] Example 2
[1077] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1078] Current communication technology makes it difficult for people living far apart to interact in a natural way. In particular, the real-time sharing of facial expressions, movements, and emotions has not been fully realized. This has led to a decline in the quality of communication in virtual reality spaces and a limited user experience.
[1079] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1080] In this invention, the server includes a means for scanning the user's entire body in 360 degrees, a means for preprocessing the acquired image data and sending it to the server, and a means for generating a 3D avatar based on the received image data using a machine learning model, which allows the user to move in real time and communicate naturally with the user, reflecting their emotions.
[1081] "User" refers to a person who uses the system to perform full-body scans, acquire movement data, and communicate within the virtual space.
[1082] "Means for scanning the whole body 360 degrees" refers to a device that captures the user's whole body from all directions and a method of operating the device.
[1083] "Preprocessing" refers to the step of processing the acquired image data by noise removal, resolution adjustment, image combination, etc.
[1084] A "server" refers to a computer system that processes data sent from users, generates avatars in virtual space, and synchronizes movement data.
[1085] A "machine learning model" refers to an algorithm that analyzes input data, recognizes patterns, and generates a 3D avatar.
[1086] A "3D avatar" refers to a three-dimensional virtual character that reproduces the user's appearance and movements.
[1087] "Device" refers to a device required for a user to use the system, such as a smartphone or VR device.
[1088] "Motion detection device" refers to a sensor device or device for capturing a user's physical movements in real time.
[1089] "Real-time" refers to a short time range in which user actions and data are reflected almost instantaneously.
[1090] "Virtual space" refers to a three-dimensional digital environment generated by a computer.
[1091] "Voice Data" means data that records and transmits a user's voice in digital form.
[1092] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and voice data to recognize and reflect their emotional state.
[1093] The system of the present invention is designed to realize high-quality communication in a virtual space. Specifically, it includes functions such as full-body scanning of the user, 3D avatar generation using a machine learning model, real-time motion data acquisition using a motion detection device, avatar synchronization within the virtual space, communication using voice data, and emotion recognition and reflection using an emotion engine.
[1094] First, the user scans their entire body in 360 degrees using a dedicated smartphone application. The user's full-body image data is pre-processed on the device to remove noise, adjust resolution, and combine images, then compressed and sent to a cloud-based server. The server analyzes the received image data and generates a detailed 3D avatar of the user using stable diffusion and other generative AI models. The generated 3D avatar data is then sent back to the device and provided to the user.
[1095] Next, the user puts on a VR device such as a head-mounted display or motion capture suit and launches the corresponding dedicated application. At this time, the device collects data from the VR device in real time using the gyro sensor, accelerometer, and position sensor, converts it into an appropriate format, compresses it, and sends it to the server. This allows the server to instantly reproduce the movements of the user's 3D avatar in the virtual space.
[1096] Furthermore, the system incorporates an emotion engine that can recognize and reflect the user's emotional state by collecting and analyzing facial and voice data, allowing the user's 3D avatar to express emotions in a more natural way.
[1097] For example, when parents and children living far apart use this system to interact in a virtual space, they both wear VR devices at home and launch a dedicated application. The child's device captures movement and emotional data in real time and sends it to the server. At the same time, the parent's device also acquires the parent's movement and emotional data and sends it to the server. The server integrates this data and reproduces and synchronizes it in the virtual space, allowing the parent and child to interact visually and audibly in real time. Furthermore, when the parent smiles with joy, that expression is reflected in the child's avatar in real time.
[1098] In this way, this system provides an advanced communication experience that transcends physical distance.
[1099] Example prompt sentence:
[1100] "Use a generative AI model to generate a 3D avatar from the user's full-body image data. The specific process includes noise removal, resolution adjustment, and avatar generation using the generative AI model."
[1101] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1102] Program processing steps
[1103] Step 1: Full body scan and image data transmission
[1104] The user launches a dedicated application on their smartphone and switches to "full body scan" mode.
[1105] Specific operation: Use the front camera to scan your entire body 360 degrees and obtain multiple image data.
[1106] Input: User's full-body image data
[1107] Once the device receives the captured image data, it applies a noise reduction filter to remove unnecessary information, then adjusts the image resolution to an appropriate size, and finally merges the images to create a single full-body image.
[1108] Output: Preprocessed whole-body image data
[1109] The device compresses the preprocessed image data and sends it to a server on the cloud.
[1110] Specific operation: The terminal automatically compresses and transmits data in parallel.
[1111] Step 2: Creating a 3D avatar using generative AI
[1112] The server sends the received preprocessed image data to the AI processing module.
[1113] Input: Preprocessed whole-body image data
[1114] The server analyzes the image using a generative AI model such as Stable Diffusion to generate a 3D avatar, which includes mesh and texture data.
[1115] Output: Generated 3D avatar data (mesh data, texture data)
[1116] The server sends the generated 3D avatar data to the user's device.
[1117] How it works: The server analyzes the image data and generates and transmits a complete 3D avatar within seconds.
[1118] Step 3: Obtaining movement data using motion capture
[1119] The user puts on a head-mounted display or motion capture suit and launches the corresponding dedicated application.
[1120] Specific actions: The user puts on the VR device and starts moving in the virtual space.
[1121] Input: User behavior data
[1122] The terminal collects data from the gyro sensor, accelerometer, and position sensor in real time from the VR device and converts it into a suitable format.
[1123] Output: Motion data converted into a suitable format
[1124] The terminal compresses the converted motion data and transmits it to the server in real time.
[1125] Specific operation: The device collects and transmits data sequentially in accordance with the user's movements.
[1126] Step 4: Projection into virtual space and avatar synchronization
[1127] The server reproduces the movements of the user's 3D avatar in the virtual space based on the motion data sent by the user.
[1128] Input: Operation data
[1129] Output: Avatar movements reproduced in virtual space
[1130] Specific actions: A user's movements in the virtual space or hand movements are synchronized with other users in real time.
[1131] Step 5: Emotion recognition and reflection using the emotion engine
[1132] The terminal collects facial expression data and voice data of the user.
[1133] Input: Facial expression data, voice data
[1134] The terminal uses an emotion engine to analyze this data and recognize the user's emotional state.
[1135] Output: Emotion recognition result
[1136] The terminal compresses the analysis results and transmits them to the server in real time.
[1137] Specific operation: When a user smiles, the facial expression is immediately analyzed by the emotion engine and reflected on the server.
[1138] Step 6: Communication Functions
[1139] Users can communicate visually and audibly with other users in the virtual space.
[1140] Specific actions: The user speaks or gestures.
[1141] Input: Voice data, gesture data
[1142] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[1143] Output: Voice communication data with other users
[1144] The server receives the voice data of the other participants and sends it to the appropriate terminal, where the voice data is played through speakers or headsets.
[1145] Specific operation: The user's voice is transmitted to other users in real time, enabling natural conversation.
[1146] (Application example 2)
[1147] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1148] Conventional communication systems in virtual spaces reproduce the movements and communication of an avatar based solely on the user's motion and voice data, and lack the ability to reflect the user's emotions in real time, making it difficult to achieve a natural communication experience. Furthermore, even in fitting experiences in virtual stores in the real world, there was an issue of not being able to grasp the user's emotions and provide optimal advice and suggestions.
[1149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1150] In this invention, the server includes means for scanning the user's entire body 360 degrees, means for preprocessing the acquired image data and transmitting it to the server, means for generating a 3D avatar based on the received image data using a generative AI model, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to acquire user movement data, means for transmitting the acquired movement data to the server in real time, means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data, means for synchronizing movement and position information between users in the virtual space, means for transmitting and receiving voice data for users to communicate in the virtual space, means for capturing user facial expression data and analyzing emotions using an emotion engine, and means for reflecting the user's 3D avatar's facial expressions and movements in the virtual space in real time based on the analyzed emotion data. This enables users to communicate in the virtual space and try on clothes in a virtual store in a natural and high-quality manner.
[1151] A "means for scanning the entire body 360 degrees" is a device or method for capturing images of a user's entire body from various angles and acquiring a series of image data.
[1152] "Image data preprocessing means" refers to a method or device that performs processing such as noise removal and resolution adjustment on acquired image data to improve the quality of the data.
[1153] The "means for transmitting to the server" refers to a device or method for transmitting the pre-processed data to the server over a network.
[1154] A "generative AI model" is an artificial intelligence model that generates a detailed 3D avatar of a user based on received image data.
[1155] A "3D avatar" is a three-dimensional digital model that mimics the user's entire body.
[1156] "Motion capture device" is a general term for devices and sensors used to capture a user's movements in real time.
[1157] "Motion data" refers to data that includes information about a user's movements acquired by a motion capture device.
[1158] The "means for transmitting to the server in real time" refers to a device or method for instantly transmitting acquired motion data to the server.
[1159] A "virtual space" is a three-dimensional digital space generated using computer technology.
[1160] "Means for reproducing movements in a virtual space" refers to devices or methods that use acquired movement data to allow a 3D avatar to move in real time within a virtual space.
[1161] The "means for synchronizing motion and position information" refers to a device or method for consistently synchronizing the motion and position information of multiple users in a virtual space.
[1162] "Means for transmitting and receiving voice data" refers to a method or device for transmitting and receiving voice data that allows a user to communicate with other users by voice in a virtual space.
[1163] "Means for capturing facial expression data" refers to a device or method for acquiring a user's facial expression in real time.
[1164] An "emotion engine" is a program or device that analyzes captured facial expression data and voice data to recognize the user's emotional state.
[1165] The "means for analyzing emotions" refers to a device or method for analyzing data acquired using an emotion engine and identifying the user's emotions.
[1166] "Means for reflecting the facial expressions and movements of a 3D avatar in real time within a virtual space" refers to a device or method for instantly changing and reflecting the facial expressions and movements of a 3D avatar within a virtual space based on analyzed emotional data.
[1167] The system of the present invention includes functions such as full-body scanning of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and recognition and reflection of the user's emotions using an emotion engine.
[1168] First, the user uses the smartphone's full-body scanning function to scan their entire body in 360 degrees, acquiring multiple image data. The acquired image data undergoes preprocessing, such as noise removal and resolution adjustment, to improve the data quality before being sent to the server. This preprocessing is carried out using OpenCV, an open-source image processing library.
[1169] The server uses a generative AI model based on the received image data to generate a detailed 3D avatar of the user. The mesh and texture data of the generated 3D avatar are then sent to the user's device, where it can be displayed on a VR device or similar.
[1170] Next, the user puts on a VR device (such as a head-mounted display or motion capture suit) and movement data is acquired. This movement data is collected in real time using gyro sensors, acceleration sensors, position sensors, etc. The acquired movement data is sent to a server and reproduced in real time as the movement of a 3D avatar in the virtual space.
[1171] This system allows multiple users to participate simultaneously and synchronizes their movements and positions in the virtual space, allowing for natural interactions between users.
[1172] Furthermore, the system captures the user's facial expression and voice data and analyzes them using an emotion engine to identify the user's emotions. This analyzed emotion data is reflected in the facial expressions and movements of the user's 3D avatar in real time. The emotion engine uses an emotion recognition model using TensorFlow.
[1173] A concrete example of this system is a virtual store try-on experience. By having a 3D avatar instantly try on the clothes the user has selected, the experience becomes as if they are actually trying them on. In addition, by using an emotion engine, the system can understand the user's reactions and provide optimal fashion advice.
[1174] Example prompt sentence:
[1175] "I would like to demonstrate a real-time try-on experience using a 3D avatar. The 3D avatar will instantly wear the clothes the user has chosen, and an emotion engine will be used to analyze the user's reaction and display the next recommended item."
[1176] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1177] Step 1:
[1178] Full body scan of the user
[1179] The user activates the full-body scan mode on their smartphone and uses the camera to scan their entire body 360 degrees. This acquires multiple image data. The input is the smartphone camera image, and the output is the acquired multiple image data. Specifically, the user rotates the camera to capture images of their entire body.
[1180] Step 2:
[1181] Image data preprocessing
[1182] The device preprocesses the acquired image data, performing noise removal, resolution adjustment, image merging, etc. The input is the acquired image data, and the output is the preprocessed image data. This processing uses the OpenCV library. Specifically, a noise removal filter is applied, and each image is merged to generate a high-resolution whole-body image.
[1183] Step 3:
[1184] Sending image data to the server
[1185] The terminal compresses the preprocessed image data and sends it to the server. The input is the preprocessed image data, and the output is the compressed data. The HTTP protocol is used to send the data. Specifically, the data is converted into a byte sequence and sent using an HTTP request.
[1186] Step 4:
[1187] 3D avatar generation
[1188] The server generates a 3D avatar using a generative AI model based on the received image data. The input is compressed image data, and the output is 3D avatar data (mesh data, texture data, etc.). Specifically, the generative AI model analyzes the image data and constructs a 3D model.
[1189] Step 5:
[1190] Sending a 3D avatar to your device
[1191] The server sends the generated 3D avatar data to the device. The input is the 3D avatar data, and the output is the avatar data sent to the device. Specifically, the data is sent in real time using WebSocket communication.
[1192] Step 6:
[1193] Obtaining movement data using motion capture
[1194] The user puts on a VR device (a head-mounted display or motion capture suit) and launches a dedicated application. The input is the user's movements and the device's sensor data, and the output is movement data. Specifically, data is collected from the gyro sensor, acceleration sensor, and position sensor.
[1195] Step 7:
[1196] Sending operation data to the server
[1197] The terminal analyzes the acquired motion data, converts it into an appropriate format, compresses it, and sends it to the server. The input is the acquired motion data, and the output is compressed motion data. Specifically, the data format is converted and the data is made smaller using a compression algorithm.
[1198] Step 8:
[1199] Reproducing movements in virtual space
[1200] The server reproduces the movements of the user's 3D avatar in the virtual space based on the movement data sent by the user. The input is compressed movement data, and the output is a 3D avatar that moves in the virtual space. Specifically, the server analyzes the data, calculates the position and posture in the virtual space, and reflects them in the 3D avatar.
[1201] Step 9:
[1202] Synchronization of movement and position information in virtual space
[1203] The server also receives other users' motion data and synchronizes the motion and location information of all users in the virtual space. The input is motion data from multiple users, and the output is a synchronized virtual space. Specifically, all user data is centrally managed and adjusted in real time.
[1204] Step 10:
[1205] Facial expression data capture and analysis
[1206] The device captures the user's facial expression data and analyzes the emotions using an emotion engine. The input is the user's facial expression data, and the output is the emotion data that is the analysis result. Specifically, the device analyzes the facial image captured by the camera and identifies the emotional state (joy, anger, sadness, etc.).
[1207] Step 11:
[1208] Sending emotion data to the server
[1209] The device compresses the analyzed emotion data and sends it to the server in real time. The input is the analyzed emotion data, and the output is the compressed emotion data. Specifically, the emotion data is converted into a byte sequence and sent using an HTTP request.
[1210] Step 12:
[1211] Reflecting facial expressions and movements in virtual space
[1212] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space. The input is emotion data, and the output is a 3D avatar that reflects the emotion. Specifically, the emotion data is converted into the avatar's facial motion and reflected in real time.
[1213] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1214] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1215] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1216] [Fourth embodiment]
[1217] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1218] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1219] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1220] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1221] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1223] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1224] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1225] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1226] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1227] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1228] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1229] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1230] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data using motion capture, avatar synchronization in a virtual space, and real-time voice communication. Below, the program processing of the entire system and specific examples are explained in natural language.
[1231] Full body scan and image data transmission
[1232] The user starts a dedicated application and scans their entire body in 360 degrees using a smartphone or dedicated camera device, which captures multiple image data.
[1233] The device preprocesses the acquired image data, including noise removal, resolution adjustment, and image merging.
[1234] The terminal compresses the preprocessed image data and sends it to the server, thereby distributing the processing load.
[1235] Creating 3D avatars with generative AI
[1236] The server uses a generative AI model based on the received full-body image data to create a detailed 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[1237] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[1238] Obtaining movement data using motion capture
[1239] The user wears a VR device (such as a head-mounted display or a motion capture suit), which captures the user's movements in real time.
[1240] The device acquires operational data from each sensor (e.g., data from the gyro sensor, acceleration sensor, and position sensor).
[1241] The device analyzes the acquired motion data, converts it into a suitable format, compresses it, and transmits it to the server in real time.
[1242] Projection into virtual space and avatar synchronization
[1243] The server processes the motion data received from the user and reproduces the avatar's motion in the virtual space. It also integrates motion data from other participants (family and friends) and synchronizes the motion and location information of all users.
[1244] Communication Features
[1245] Users communicate with other users in the virtual space visually and audibly, for example by talking or waving.
[1246] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[1247] The server transfers the received audio data to the terminals of the other participants, and the audio is played back at each terminal.
[1248] Specific examples
[1249] For example, consider the case where grandparents who live far away use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear similar devices at their homes.
[1250] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server. At the same time, the grandparent's device also acquires the grandparent's motion data and sends it to the server.
[1251] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk with their grandchildren as if they were there in the virtual space.
[1252] The devices also exchange voice data between users in real time, providing a natural conversation experience. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[1253] The processing flow will be explained below.
[1254] Step 1:
[1255] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[1256] The user uses a camera to scan their entire body 360 degrees, obtaining multiple image data.
[1257] Step 2:
[1258] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[1259] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[1260] Step 3:
[1261] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[1262] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[1263] Step 4:
[1264] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[1265] Step 5:
[1266] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[1267] Step 6:
[1268] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[1269] Step 7:
[1270] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[1271] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[1272] Step 8:
[1273] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[1274] Step 9:
[1275] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[1276] Step 10:
[1277] The server receives the voice data of the other participants and transmits it to the user's terminal.
[1278] The audio data received by the device is played through speakers or a headset.
[1279] As described above, the system of the present invention realizes high-quality real-time communication that transcends physical distance.
[1280] Example 1
[1281] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1282] In modern society, users need to communicate with each other in real time, regardless of physical distance. In particular, real-time data communication, both visual and audio, is essential for immersive communication with family and friends in distant locations. However, current systems have difficulty accurately capturing a user's full-body movements and voice in real time and reproducing them in a virtual space. This limits the communication experience and leads to problems such as connection delays and instability.
[1283] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1284] In this invention, the server includes means for performing a 360-degree full-body scan, means for preprocessing the image data and transmitting it to the server, means for generating a 3D avatar using a generative AI model based on the received image data, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to capture user motion data, means for analyzing the captured motion data and converting it into a suitable format, means for transmitting the converted motion data to the server in real time, means for reproducing the motion of the 3D avatar in a virtual space based on the real-time motion data, means for synchronizing motion and position information between users in the virtual space, and means for transmitting and receiving audio data for users to communicate in the virtual space. This allows the user's full-body motions and voices to be reflected in the virtual space in real time, enabling a high-quality communication experience.
[1285] A "means for scanning the entire body 360 degrees" is a device or software that photographs the user's entire body from various angles and collects data from all directions.
[1286] "Means for preprocessing image data and sending it to the server" refers to devices or software that process the acquired image data by removing noise, adjusting resolution, combining images, etc., and then compressing and sending it to the server.
[1287] "Means for generating a 3D avatar using a generative AI model" refers to modules or software that use artificial intelligence technology to generate a three-dimensional avatar based on received image data.
[1288] "Means for transmitting the 3D avatar to the user's terminal" refers to a communication protocol or device for transmitting the generated 3D avatar as data to the end user's terminal.
[1289] "Means using a motion capture device" means a device or software for detecting and capturing a user's physical movements in real time.
[1290] The "means for analyzing the motion data and converting it into a suitable format" refers to a device or software for analyzing the acquired motion data and converting it into a suitable data format for reproduction in a virtual space.
[1291] The "means for transmitting the converted motion data to the server in real time" refers to a device or software for compressing the motion data converted into an appropriate format and transmitting it to the server in real time without any loss.
[1292] "Means for reproducing the movements of a 3D avatar in a virtual space" refers to an engine or software that faithfully reproduces the movements of a 3D avatar in a virtual space based on received movement data.
[1293] "Means for synchronizing the movements and location information of users in a virtual space" refers to a system or software that consistently reflects the movements and location information of multiple users in real time, enabling interaction within a virtual space.
[1294] "Means for transmitting and receiving voice data" refers to a communication system or software for capturing a user's voice in real time and transmitting and playing it back to other users.
[1295] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition and transmission of real-time motion data through motion capture, synchronization of the motion data in a virtual space, and voice communication. A specific embodiment of the entire system will be described below.
[1296] Full body scan and image data transmission
[1297] The user launches a dedicated application on a smartphone or dedicated camera device and scans their entire body in 360 degrees, obtaining multiple image data. Specifically, the user rotates the smartphone to capture their entire body, and the app instructs them on the direction and angle of the shot.
[1298] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. Noise removal is performed using edge detection algorithms, and resolution is optimized. The individual images are then stitched together to generate the entire image.
[1299] The pre-processed image data is compressed by the terminal and sent to the server via HTTP protocol, using the JPEG compression algorithm for efficient uploading to the server.
[1300] Creating 3D avatars with generative AI
[1301] The server uses a generative AI model (e.g., a Deep Learning-based 3D Reconstruction model) based on the received full-body image data to generate a 3D avatar of the user. The generation process involves image analysis and 3D model generation.
[1302] The generated 3D avatar data (mesh data, texture data, etc.) is sent from the server to the user's device. The data is packaged in JSON format and sent via a RESTful API.
[1303] Obtaining movement data using motion capture
[1304] The user wears VR devices (such as a head-mounted display and a motion capture suit) and captures their movements in real time. The user wears VR goggles and sensors from the motion capture suit, which are attached to the user's body, and the sensors capture the user's body movements.
[1305] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received via Bluetooth or Wi-Fi.
[1306] The acquired motion data is analyzed by the device and converted into quaternion format, which is then compressed using a compression algorithm such as Zlib and sent to the server in real time using the WebSocket protocol.
[1307] Projection into virtual space and avatar synchronization
[1308] The server processes the motion data received from the user and reproduces the movements of a 3D avatar in the virtual space. All user motion data is integrated using a real-time motion engine. The data is then reflected in the virtual space using a game engine such as Unity.
[1309] Communication Features
[1310] Users can communicate with other users visually and audibly in the virtual space by talking to other avatars or waving to them. Facial tracking can also be used to reflect facial expressions on the avatars in real time.
[1311] The device captures the user's voice data from the microphone, encodes it using an audio codec such as Opus, and sends it to the server. The voice data is then transferred in real time to other devices and played back through the speakers.
[1312] Specific examples
[1313] For example, grandparents who live far away can use this system to check on their grandchildren's growth. The grandparents wear VR devices at home, and the grandchildren wear VR devices as well.
[1314] The grandchild's device captures the grandchild's movements in real time and sends the motion data to the server, while the grandparent's device simultaneously captures the grandparent's motion data and sends it to the server.
[1315] The server integrates this motion data and reflects it in the virtual space, allowing grandparents to wave and talk as if they were actually there with their grandchildren. The devices also exchange voice data between users in real time, providing a natural conversation experience.
[1316] Prompt Sentence Examples
[1317] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[1318] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1319] Step 1:
[1320] Users launch a dedicated application on their smartphone or dedicated camera device and perform a 360-degree full-body scan. Users rotate their smartphone to capture their entire body, adjusting the angle and direction of the shot according to the application's instructions.
[1321] Input: Multiple image data captured by a smartphone or dedicated camera.
[1322] Output: 360-degree full-body scan image data.
[1323] Step 2:
[1324] The device performs pre-processing on the acquired image data, such as noise removal, resolution adjustment, and image merging. After using an image edge detection algorithm to remove noise and optimize the resolution, the individual images are stitched together to generate a complete image.
[1325] Input: 360 degree scanned image data.
[1326] Output: Preprocessed whole-body image data.
[1327] Step 3:
[1328] The terminal compresses the pre-processed image data using the JPEG compression algorithm and sends it to the server via the HTTP protocol. Compressing the image data improves transfer speed and reduces the load on the network bandwidth.
[1329] Input: Preprocessed whole-body image data.
[1330] Output: Compressed image data.
[1331] Step 4:
[1332] The server uses a generative AI model (e.g., a deep learning-based 3D reconstruction model) based on the received image data to generate a detailed 3D avatar of the user, analyzes the image, and generates a 3D mesh and texture.
[1333] Input: Compressed whole-body image data.
[1334] Output: 3D avatar data (mesh data, texture data).
[1335] Step 5:
[1336] The server packages the generated 3D avatar data in JSON format and sends it to the user's device via a RESTful API, which ensures structured and efficient data transfer.
[1337] Input: 3D avatar data (mesh data, texture data).
[1338] Output: 3D avatar data packaged in JSON format.
[1339] Step 6:
[1340] Users wear VR devices (such as a head-mounted display or a motion capture suit) and their movements are captured in real time. Users wear VR goggles and the motion capture suit's sensors are attached to their bodies, allowing the user's movements to be captured in detail.
[1341] Input: User's physical movements.
[1342] Output: Real-time operating data.
[1343] Step 7:
[1344] The device receives motion data from each sensor (e.g., gyro sensor, accelerometer, and location sensor) and stores it in a local database. The data is received in real time via Bluetooth or Wi-Fi.
[1345] Input: Operational data from each sensor.
[1346] Output: Operational data stored in a local database.
[1347] Step 8:
[1348] The device analyzes the acquired motion data, converts it into quaternion format, compresses it using a compression algorithm such as Zlib, and transmits it to the server in real time using the WebSocket protocol.
[1349] Input: Operational data stored in a local database.
[1350] Output: Compressed quaternion format motion data.
[1351] Step 9:
[1352] The server processes the motion data received from the user using a real-time motion engine and reproduces the movements of a 3D avatar in the virtual space. Motion data from other participants is also integrated and synchronized, and the motion data of all users is reflected in the virtual space.
[1353] Input: Compressed quaternion format motion data.
[1354] Output: 3D avatar movement data reproduced in virtual space.
[1355] Step 10:
[1356] The device captures the user's voice data in real time from the microphone, encodes it using a voice codec such as Opus, and sends it to the server. The voice data is divided into packets with timestamps.
[1357] Input: User's voice data.
[1358] Output: The encoded audio data.
[1359] Step 11:
[1360] The server transfers the user's encoded voice data to the other participants' terminals and plays it in real time. The other participants' voice data is similarly processed and played on the user's terminal.
[1361] Input: Encoded audio data.
[1362] Output: Audio data for playback.
[1363] Prompt Sentence Examples
[1364] "Generate a 3D avatar based on the user's full-body image data and project it into the virtual space in real time. Also, synchronize movement data and voice data in real time."
[1365] (Application example 1)
[1366] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1367] One drawback of modern online shopping is that users cannot actually try on products. This often results in purchased products that do not meet actual expectations, resulting in a poor user experience. There is also a demand for a richer shopping experience that does not rely on physical stores. Furthermore, there is a lack of a way for users in different locations to communicate with each other in real time while shopping.
[1368] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1369] In this invention, the server includes: means for scanning the user's entire body 360 degrees; means for preprocessing the acquired image data and transmitting it to the server; means for generating a 3D avatar based on the received image data using a generative AI model; means for transmitting the generated 3D avatar to the user's device; means for using a motion capture device to acquire user movement data; means for transmitting the acquired movement data to the server in real time; means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data; means for synchronizing movement and position information between users in the virtual space; means for transmitting and receiving audio data for users to communicate in the virtual space; and means for providing a system in a virtual store that allows users to try on products online using a 3D avatar. This allows users to enjoy shopping while trying on products in real time from the comfort of their own home and communicating with other users.
[1370] A "means for scanning the entire body 360 degrees" is a device or system that can photograph the user's entire body from each angle and obtain image data from all directions.
[1371] "Preprocessing" refers to processing the acquired image data, such as noise removal, resolution adjustment, and image combination.
[1372] A "generative AI model" is a model that uses an artificial intelligence algorithm to generate a 3D avatar based on acquired image data.
[1373] A "motion capture device" is a device that captures a user's movements in real time and acquires them as movement data.
[1374] A "virtual space" is a computer-generated interactive 3D space in which users can perform various actions through their avatars.
[1375] "Means for synchronizing movement and position information" refers to technologies and systems for matching the movements and positions of multiple users in a virtual space in real time.
[1376] "Means for transmitting and receiving voice data" refers to a device and system for capturing a user's voice and transmitting it to other users via the Internet.
[1377] "Means provided within a virtual store" refers to a system that provides an environment on an online shopping platform where users can virtually try on products using 3D avatars.
[1378] This system scans a user's entire body in 360 degrees, generates a 3D avatar using a generative AI model, and then recreates the avatar in a virtual space in conjunction with motion capture data. This system allows users to try on products in the virtual space and communicate with other users in real time.
[1379] First, the user scans their entire body using a smartphone or dedicated camera device. In this step, the user rotates the device to capture images from all directions (360 degrees). The device preprocesses the captured image data, removing noise and adjusting the resolution, before sending it to the server. The OpenCV library is used for preprocessing.
[1380] The server uses a generative AI model to create a 3D avatar based on the received image data. This AI model uses OpenAI's API to analyze the image data and generate a detailed 3D avatar of the user. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the user's device.
[1381] Next, the user wears a motion capture device (e.g., a VR device or motion capture suit) and their movement data is acquired in real time. The device acquires the movement data from various sensors (e.g., gyro sensor, accelerometer, position sensor), converts it into an appropriate format, compresses it, and sends it to the server. During this process, appropriate libraries (e.g., NumPy or the Python standard library) are used to format and compress the data.
[1382] The server then recreates the movements of the 3D avatar in the virtual space based on the received motion data, and synchronizes the motion and position information with the avatars of other users, allowing all users to experience the sensation of operating the avatar in real time in the same virtual space.
[1383] Furthermore, users can send and receive voice data to communicate within the virtual space. The device picks up the user's voice data from the microphone and sends it to the server. The server then forwards the received voice data to the devices of the other participants, and the voice is played on each device.
[1384] As an example, consider a scenario in which a user tries on products in a virtual store in a virtual space. The user uses a smartphone or VR device to scan their entire body and generate a 3D avatar, then logs into the virtual store. At this time, the user might input a prompt sentence like the following into the generative AI model:
[1385] "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3"
[1386] Users can then view and rotate the products using their own 3D avatar as they try them on in a virtual space, and can also interact with other users in real time to share their thoughts on the fit and design of the products.
[1387] As described above, the present invention allows users to transcend physical constraints and enjoy a richer shopping experience in a virtual space.
[1388] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1389] Step 1:
[1390] The user launches a dedicated application and scans their entire body from all directions using a smartphone or dedicated camera device, acquiring multiple image data. This provides a full-body image of the user as input.
[1391] Step 2:
[1392] The device preprocesses the acquired image data. This preprocessing includes noise removal, resolution adjustment, and image merging. The OpenCV library is used for noise removal, resolution adjustment, and image merging. The preprocessed image data is obtained as the output.
[1393] Step 3:
[1394] The device compresses the preprocessed image data and sends it to the server. The compression uses an appropriate library to efficiently transmit the data. The preprocessed and compressed image data is sent as input to the server.
[1395] Step 4:
[1396] The server generates a 3D avatar for the user using a generative AI model based on the received image data. The generative AI model uses OpenAI's API to analyze the image and generate a detailed 3D avatar. The generated 3D avatar is obtained as the output.
[1397] Step 5:
[1398] The server sends the generated 3D avatar data to the user's device using a standard data transfer protocol. The user's device receives the 3D avatar data.
[1399] Step 6:
[1400] The user wears a motion capture device (e.g., a VR device or a motion capture suit) and motion data is acquired. The motion capture device acquires motion data in real time using gyro sensors, accelerometers, and position sensors. The motion data is obtained as output.
[1401] Step 7:
[1402] The device analyzes the motion data acquired from each sensor and converts it into an appropriate format. This process involves shaping the data and converting it into a compatible format. The analyzed motion data is then output.
[1403] Step 8:
[1404] The terminal compresses the converted motion data and transmits it to the server in real time using a standard data compression algorithm. The compressed motion data is then transmitted as input to the server.
[1405] Step 9:
[1406] The server reproduces the 3D avatar's movements in the virtual space based on the received movement data. This is done using a simulation engine, which reproduces real-time movements in the virtual space. The reproduced movement data is obtained as output.
[1407] Step 10:
[1408] The server synchronizes the motion and location information with the avatars of other users, so that all users can see the synchronized motion in the same virtual space. The synchronized motion and location information is obtained as output.
[1409] Step 11:
[1410] Users send and receive voice data to communicate within the virtual space. The device captures the voice data from the microphone and sends it to the server. The server then transfers the received voice data to other users' devices and plays it back in real time. The sent and received voice data is obtained as output.
[1411] Step 12:
[1412] Users use their own 3D avatar to try on products in a virtual store. Specifically, they input the following prompt into the generative AI model: "Generate a 3D avatar of the user based on the following images: Image 1, Image 2, Image 3." This allows users to try on products online and communicate with other users in real time.
[1413] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1414] The system of the present invention includes a full-body scan of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and the ability to recognize and reflect the user's emotions using an emotion engine. Below, the program processing of the entire system and specific examples are explained in natural language.
[1415] Full body scan and image data transmission
[1416] The user launches a dedicated application on their smartphone and switches to full-body scan mode, using the camera to scan the entire body 360 degrees and capture multiple image data.
[1417] The terminal preprocesses the acquired image data, for example, by removing noise, adjusting resolution, combining images, etc. The preprocessed image data is compressed and sent to the server.
[1418] Creating 3D avatars with generative AI
[1419] The server analyzes the received full-body image data and generates a detailed 3D avatar of the user using a generative AI model. The generated 3D avatar data (mesh data, texture data, etc.) is then sent to the device.
[1420] Obtaining movement data using motion capture
[1421] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[1422] The device acquires real-time motion data from the VR device, specifically, data from the gyro sensor, accelerometer, and position sensor.
[1423] The device analyzes the acquired motion data, converts it into an appropriate format, compresses it, and transmits it to the server in real time.
[1424] Projection into virtual space and avatar synchronization
[1425] The server reproduces the user's 3D avatar's movements in the virtual space based on the movement data sent by the user. It also receives the movement data of other participants and synchronizes the movement and position information of all users in the virtual space.
[1426] Emotion recognition and reflection using emotion engine
[1427] The device collects the user's facial expression data and voice data and analyzes this data using an emotion engine.
[1428] The terminal compresses the analysis results and transmits them to the server in real time.
[1429] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space.
[1430] Communication Features
[1431] Users communicate with other users visually and audibly in the virtual space, for example, by talking, exchanging glances, and making gestures.
[1432] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[1433] The server receives the audio data from other participants and sends it to the appropriate device, where it is played through speakers or a headset.
[1434] Specific examples
[1435] For example, consider a situation where parents and children living far apart use this system to communicate in a virtual space, with the parent wearing a VR device at home and the child wearing one as well.
[1436] The child's device captures the child's behavior and emotional data in real time and sends it to the server, while the parent's device simultaneously captures the parent's behavior and emotional data and sends it to the server.
[1437] The server integrates this data and recreates and synchronizes it in the virtual space, allowing parents to interact with their child's 3D avatar in real time visually and audibly, and even feel their child's emotions.
[1438] Furthermore, when a parent smiles with joy, that expression is reflected in the child's avatar in real time, enabling natural communication even from a distance. In this way, this system realizes a high-quality communication experience that transcends physical distance.
[1439] The processing flow will be explained below.
[1440] Step 1:
[1441] The user launches a dedicated application on their smartphone and switches to full-body scan mode.
[1442] The user uses a camera to scan the entire body 360 degrees and acquires multiple image data.
[1443] Step 2:
[1444] Preprocessing of multiple image data acquired by the device, specifically noise removal, resolution adjustment, image merging, etc.
[1445] The terminal compresses the preprocessed whole-body image data and transmits it to the server.
[1446] Step 3:
[1447] The server analyzes the received full-body image data and generates a 3D avatar of the user using a generative AI model.
[1448] The server sends the generated 3D avatar data (mesh data, texture data, etc.) to the user's device.
[1449] Step 4:
[1450] The user puts on a VR device (such as a head-mounted display or motion capture suit) and launches a dedicated application.
[1451] Step 5:
[1452] The device acquires motion data from the VR device in real time, specifically from sensors such as the gyro sensor, accelerometer, and position sensor.
[1453] Step 6:
[1454] The device analyzes the collected motion data, converts it into an appropriate format, compresses the converted motion data, and transmits it to the server in real time.
[1455] Step 7:
[1456] Based on the movement data received by the server in real time, the movements of the user's 3D avatar are reproduced in the virtual space.
[1457] The server also receives motion data from other participants and synchronizes the motion and position information of all users within the virtual space.
[1458] Step 8:
[1459] Users communicate visually and audibly with other users in a virtual space, for example by speaking or using gestures.
[1460] Step 9:
[1461] The device uses an emotion engine to recognize the user's emotions and collects the user's facial expression data and voice data.
[1462] The device analyzes this data to identify the user's emotions.
[1463] Step 10:
[1464] The device compresses the analysis results and sends them to the server in real time.
[1465] Step 11:
[1466] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in the virtual space.
[1467] Step 12:
[1468] The device captures the user's voice data from the microphone and transmits it to the server in real time.
[1469] Step 13:
[1470] The server receives the voice data of the other participants and transmits it to the user's terminal.
[1471] The audio data received by the device is played through speakers or a headset.
[1472] As described above, the system of the present invention realizes high-quality communication through the creation of a 3D avatar using generative AI from a full-body scan of the user, motion capture, emotion recognition using an emotion engine, and real-time reproduction of facial expressions and movements in a virtual space.
[1473] Example 2
[1474] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1475] Current communication technology makes it difficult for people living far apart to interact in a natural way. In particular, the real-time sharing of facial expressions, movements, and emotions has not been fully realized. This has led to a decline in the quality of communication in virtual reality spaces and a limited user experience.
[1476] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1477] In this invention, the server includes a means for scanning the user's entire body in 360 degrees, a means for preprocessing the acquired image data and sending it to the server, and a means for generating a 3D avatar based on the received image data using a machine learning model, which allows the user to move in real time and communicate naturally with the user, reflecting their emotions.
[1478] "User" refers to a person who uses the system to perform full-body scans, acquire movement data, and communicate within the virtual space.
[1479] "Means for scanning the whole body 360 degrees" refers to a device that captures the user's whole body from all directions and a method of operating the device.
[1480] "Preprocessing" refers to the step of processing the acquired image data by noise removal, resolution adjustment, image combination, etc.
[1481] A "server" refers to a computer system that processes data sent from users, generates avatars in virtual space, and synchronizes movement data.
[1482] A "machine learning model" refers to an algorithm that analyzes input data, recognizes patterns, and generates a 3D avatar.
[1483] A "3D avatar" refers to a three-dimensional virtual character that reproduces the user's appearance and movements.
[1484] "Device" refers to a device required for a user to use the system, such as a smartphone or VR device.
[1485] "Motion detection device" refers to a sensor device or device for capturing a user's physical movements in real time.
[1486] "Real-time" refers to a short time range in which user actions and data are reflected almost instantaneously.
[1487] "Virtual space" refers to a three-dimensional digital environment generated by a computer.
[1488] "Voice Data" means data that records and transmits a user's voice in digital form.
[1489] An "emotion engine" refers to an algorithm or system that analyzes a user's facial expressions and voice data to recognize and reflect their emotional state.
[1490] The system of the present invention is designed to realize high-quality communication in a virtual space. Specifically, it includes functions such as full-body scanning of the user, 3D avatar generation using a machine learning model, real-time motion data acquisition using a motion detection device, avatar synchronization within the virtual space, communication using voice data, and emotion recognition and reflection using an emotion engine.
[1491] First, the user scans their entire body in 360 degrees using a dedicated smartphone application. The user's full-body image data is pre-processed on the device to remove noise, adjust resolution, and combine images, then compressed and sent to a cloud-based server. The server analyzes the received image data and generates a detailed 3D avatar of the user using stable diffusion and other generative AI models. The generated 3D avatar data is then sent back to the device and provided to the user.
[1492] Next, the user puts on a VR device such as a head-mounted display or motion capture suit and launches the corresponding dedicated application. At this time, the device collects data from the VR device in real time using the gyro sensor, accelerometer, and position sensor, converts it into an appropriate format, compresses it, and sends it to the server. This allows the server to instantly reproduce the movements of the user's 3D avatar in the virtual space.
[1493] Furthermore, the system incorporates an emotion engine that can recognize and reflect the user's emotional state by collecting and analyzing facial and voice data, allowing the user's 3D avatar to express emotions in a more natural way.
[1494] For example, when parents and children living far apart use this system to interact in a virtual space, they both wear VR devices at home and launch a dedicated application. The child's device captures movement and emotional data in real time and sends it to the server. At the same time, the parent's device also acquires the parent's movement and emotional data and sends it to the server. The server integrates this data and reproduces and synchronizes it in the virtual space, allowing the parent and child to interact visually and audibly in real time. Furthermore, when the parent smiles with joy, that expression is reflected in the child's avatar in real time.
[1495] In this way, this system provides an advanced communication experience that transcends physical distance.
[1496] Example prompt sentence:
[1497] "Use a generative AI model to generate a 3D avatar from the user's full-body image data. The specific process includes noise removal, resolution adjustment, and avatar generation using the generative AI model."
[1498] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1499] Program processing steps
[1500] Step 1: Full body scan and image data transmission
[1501] The user launches a dedicated application on their smartphone and switches to "full body scan" mode.
[1502] Specific operation: Use the front camera to scan your entire body 360 degrees and obtain multiple image data.
[1503] Input: User's full-body image data
[1504] Once the device receives the captured image data, it applies a noise reduction filter to remove unnecessary information, then adjusts the image resolution to an appropriate size, and finally merges the images to create a single full-body image.
[1505] Output: Preprocessed whole-body image data
[1506] The device compresses the preprocessed image data and sends it to a server on the cloud.
[1507] Specific operation: The terminal automatically compresses and transmits data in parallel.
[1508] Step 2: Creating a 3D avatar using generative AI
[1509] The server sends the received preprocessed image data to the AI processing module.
[1510] Input: Preprocessed whole-body image data
[1511] The server analyzes the image using a generative AI model such as Stable Diffusion to generate a 3D avatar, which includes mesh and texture data.
[1512] Output: Generated 3D avatar data (mesh data, texture data)
[1513] The server sends the generated 3D avatar data to the user's device.
[1514] How it works: The server analyzes the image data and generates and transmits a complete 3D avatar within seconds.
[1515] Step 3: Obtaining movement data using motion capture
[1516] The user puts on a head-mounted display or motion capture suit and launches the corresponding dedicated application.
[1517] Specific actions: The user puts on the VR device and starts moving in the virtual space.
[1518] Input: User behavior data
[1519] The terminal collects data from the gyro sensor, accelerometer, and position sensor in real time from the VR device and converts it into a suitable format.
[1520] Output: Motion data converted into a suitable format
[1521] The terminal compresses the converted motion data and transmits it to the server in real time.
[1522] Specific operation: The device collects and transmits data sequentially in accordance with the user's movements.
[1523] Step 4: Projection into virtual space and avatar synchronization
[1524] The server reproduces the movements of the user's 3D avatar in the virtual space based on the motion data sent by the user.
[1525] Input: Operation data
[1526] Output: Avatar movements reproduced in virtual space
[1527] Specific actions: A user's movements in the virtual space or hand movements are synchronized with other users in real time.
[1528] Step 5: Emotion recognition and reflection using the emotion engine
[1529] The terminal collects facial expression data and voice data of the user.
[1530] Input: Facial expression data, voice data
[1531] The terminal uses an emotion engine to analyze this data and recognize the user's emotional state.
[1532] Output: Emotion recognition result
[1533] The terminal compresses the analysis results and transmits them to the server in real time.
[1534] Specific operation: When a user smiles, the facial expression is immediately analyzed by the emotion engine and reflected on the server.
[1535] Step 6: Communication Functions
[1536] Users can communicate visually and audibly with other users in the virtual space.
[1537] Specific actions: The user speaks or gestures.
[1538] Input: Voice data, gesture data
[1539] The terminal acquires the user's voice data from a microphone and transmits it to the server in real time.
[1540] Output: Voice communication data with other users
[1541] The server receives the voice data of the other participants and sends it to the appropriate terminal, where the voice data is played through speakers or headsets.
[1542] Specific operation: The user's voice is transmitted to other users in real time, enabling natural conversation.
[1543] (Application example 2)
[1544] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1545] Conventional communication systems in virtual spaces reproduce the movements and communication of an avatar based solely on the user's motion and voice data, and lack the ability to reflect the user's emotions in real time, making it difficult to achieve a natural communication experience. Furthermore, even in fitting experiences in virtual stores in the real world, there was an issue of not being able to grasp the user's emotions and provide optimal advice and suggestions.
[1546] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1547] In this invention, the server includes means for scanning the user's entire body 360 degrees, means for preprocessing the acquired image data and transmitting it to the server, means for generating a 3D avatar based on the received image data using a generative AI model, means for transmitting the generated 3D avatar to the user's device, means for using a motion capture device to acquire user movement data, means for transmitting the acquired movement data to the server in real time, means for reproducing the 3D avatar's movements in a virtual space based on the real-time movement data, means for synchronizing movement and position information between users in the virtual space, means for transmitting and receiving voice data for users to communicate in the virtual space, means for capturing user facial expression data and analyzing emotions using an emotion engine, and means for reflecting the user's 3D avatar's facial expressions and movements in the virtual space in real time based on the analyzed emotion data. This enables users to communicate in the virtual space and try on clothes in a virtual store in a natural and high-quality manner.
[1548] A "means for scanning the entire body 360 degrees" is a device or method for capturing images of a user's entire body from various angles and acquiring a series of image data.
[1549] "Image data preprocessing means" refers to a method or device that performs processing such as noise removal and resolution adjustment on acquired image data to improve the quality of the data.
[1550] The "means for transmitting to the server" refers to a device or method for transmitting the pre-processed data to the server over a network.
[1551] A "generative AI model" is an artificial intelligence model that generates a detailed 3D avatar of a user based on received image data.
[1552] A "3D avatar" is a three-dimensional digital model that mimics the user's entire body.
[1553] "Motion capture device" is a general term for devices and sensors used to capture a user's movements in real time.
[1554] "Motion data" refers to data that includes information about a user's movements acquired by a motion capture device.
[1555] The "means for transmitting to the server in real time" refers to a device or method for instantly transmitting acquired motion data to the server.
[1556] A "virtual space" is a three-dimensional digital space generated using computer technology.
[1557] "Means for reproducing movements in a virtual space" refers to devices or methods that use acquired movement data to allow a 3D avatar to move in real time within a virtual space.
[1558] The "means for synchronizing motion and position information" refers to a device or method for consistently synchronizing the motion and position information of multiple users in a virtual space.
[1559] "Means for transmitting and receiving voice data" refers to a method or device for transmitting and receiving voice data that allows a user to communicate with other users by voice in a virtual space.
[1560] "Means for capturing facial expression data" refers to a device or method for acquiring a user's facial expression in real time.
[1561] An "emotion engine" is a program or device that analyzes captured facial expression data and voice data to recognize the user's emotional state.
[1562] The "means for analyzing emotions" refers to a device or method for analyzing data acquired using an emotion engine and identifying the user's emotions.
[1563] "Means for reflecting the facial expressions and movements of a 3D avatar in real time within a virtual space" refers to a device or method for instantly changing and reflecting the facial expressions and movements of a 3D avatar within a virtual space based on analyzed emotional data.
[1564] The system of the present invention includes functions such as full-body scanning of the user, generation of a 3D avatar using a generative AI model, acquisition of movement data through motion capture, synchronization between avatars in the virtual space, real-time voice communication, and recognition and reflection of the user's emotions using an emotion engine.
[1565] First, the user uses the smartphone's full-body scanning function to scan their entire body in 360 degrees, acquiring multiple image data. The acquired image data undergoes preprocessing, such as noise removal and resolution adjustment, to improve the data quality before being sent to the server. This preprocessing is carried out using OpenCV, an open-source image processing library.
[1566] The server uses a generative AI model based on the received image data to generate a detailed 3D avatar of the user. The mesh and texture data of the generated 3D avatar are then sent to the user's device, where it can be displayed on a VR device or similar.
[1567] Next, the user puts on a VR device (such as a head-mounted display or motion capture suit) and movement data is acquired. This movement data is collected in real time using gyro sensors, acceleration sensors, position sensors, etc. The acquired movement data is sent to a server and reproduced in real time as the movement of a 3D avatar in the virtual space.
[1568] This system allows multiple users to participate simultaneously and synchronizes their movements and positions in the virtual space, allowing for natural interactions between users.
[1569] Furthermore, the system captures the user's facial expression and voice data and analyzes them using an emotion engine to identify the user's emotions. This analyzed emotion data is reflected in the facial expressions and movements of the user's 3D avatar in real time. The emotion engine uses an emotion recognition model using TensorFlow.
[1570] A concrete example of this system is a virtual store try-on experience. By having a 3D avatar instantly try on the clothes the user has selected, the experience becomes as if they are actually trying them on. In addition, by using an emotion engine, the system can understand the user's reactions and provide optimal fashion advice.
[1571] Example prompt sentence:
[1572] "I would like to demonstrate a real-time try-on experience using a 3D avatar. The 3D avatar will instantly wear the clothes the user has chosen, and an emotion engine will be used to analyze the user's reaction and display the next recommended item."
[1573] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1574] Step 1:
[1575] Full body scan of the user
[1576] The user activates the full-body scan mode on their smartphone and uses the camera to scan their entire body 360 degrees. This acquires multiple image data. The input is the smartphone camera image, and the output is the acquired multiple image data. Specifically, the user rotates the camera to capture images of their entire body.
[1577] Step 2:
[1578] Image data preprocessing
[1579] The device preprocesses the acquired image data, performing noise removal, resolution adjustment, image merging, etc. The input is the acquired image data, and the output is the preprocessed image data. This processing uses the OpenCV library. Specifically, a noise removal filter is applied, and each image is merged to generate a high-resolution whole-body image.
[1580] Step 3:
[1581] Sending image data to the server
[1582] The terminal compresses the preprocessed image data and sends it to the server. The input is the preprocessed image data, and the output is the compressed data. The HTTP protocol is used to send the data. Specifically, the data is converted into a byte sequence and sent using an HTTP request.
[1583] Step 4:
[1584] 3D avatar generation
[1585] The server generates a 3D avatar using a generative AI model based on the received image data. The input is compressed image data, and the output is 3D avatar data (mesh data, texture data, etc.). Specifically, the generative AI model analyzes the image data and constructs a 3D model.
[1586] Step 5:
[1587] Sending a 3D avatar to your device
[1588] The server sends the generated 3D avatar data to the device. The input is the 3D avatar data, and the output is the avatar data sent to the device. Specifically, the data is sent in real time using WebSocket communication.
[1589] Step 6:
[1590] Obtaining movement data using motion capture
[1591] The user puts on a VR device (a head-mounted display or motion capture suit) and launches a dedicated application. The input is the user's movements and the device's sensor data, and the output is movement data. Specifically, data is collected from the gyro sensor, acceleration sensor, and position sensor.
[1592] Step 7:
[1593] Sending operation data to the server
[1594] The terminal analyzes the acquired motion data, converts it into an appropriate format, compresses it, and sends it to the server. The input is the acquired motion data, and the output is compressed motion data. Specifically, the data format is converted and the data is made smaller using a compression algorithm.
[1595] Step 8:
[1596] Reproducing movements in virtual space
[1597] The server reproduces the movements of the user's 3D avatar in the virtual space based on the movement data sent by the user. The input is compressed movement data, and the output is a 3D avatar that moves in the virtual space. Specifically, the server analyzes the data, calculates the position and posture in the virtual space, and reflects them in the 3D avatar.
[1598] Step 9:
[1599] Synchronization of movement and position information in virtual space
[1600] The server also receives other users' motion data and synchronizes the motion and location information of all users in the virtual space. The input is motion data from multiple users, and the output is a synchronized virtual space. Specifically, all user data is centrally managed and adjusted in real time.
[1601] Step 10:
[1602] Facial expression data capture and analysis
[1603] The device captures the user's facial expression data and analyzes the emotions using an emotion engine. The input is the user's facial expression data, and the output is the emotion data that is the analysis result. Specifically, the device analyzes the facial image captured by the camera and identifies the emotional state (joy, anger, sadness, etc.).
[1604] Step 11:
[1605] Sending emotion data to the server
[1606] The device compresses the analyzed emotion data and sends it to the server in real time. The input is the analyzed emotion data, and the output is the compressed emotion data. Specifically, the emotion data is converted into a byte sequence and sent using an HTTP request.
[1607] Step 12:
[1608] Reflecting facial expressions and movements in virtual space
[1609] Based on the analysis results of the emotion engine, the server reflects the facial expressions and movements of the user's 3D avatar in real time within the virtual space. The input is emotion data, and the output is a 3D avatar that reflects the emotion. Specifically, the emotion data is converted into the avatar's facial motion and reflected in real time.
[1610] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1611] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1612] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1613] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1614] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1615] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1616] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1617] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1618] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1619] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1620] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1621] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1622] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1623] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1624] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1625] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1626] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1627] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1628] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1629] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1630] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1631] The following is further disclosed regarding the above embodiment.
[1632] (Claim 1)
[1633] A means of scanning the user's entire body in 360 degrees;
[1634] means for preprocessing the acquired image data and transmitting the same to a server;
[1635] A means for generating a 3D avatar using a generative AI model based on the received image data;
[1636] A means for transmitting the generated 3D avatar to a user's device;
[1637] a means for using a motion capture device to acquire motion data from a user;
[1638] means for transmitting the acquired motion data to a server in real time;
[1639] A method for reproducing the movements of a 3D avatar in a virtual space based on real-time movement data, and
[1640] means for synchronizing motion and position information between users in a virtual space;
[1641] means for transmitting and receiving voice data for users to communicate in the virtual space;
[1642] A system including:
[1643] (Claim 2)
[1644] 10. The system of claim 1, further comprising means for analyzing the user's motion data and converting it into a suitable format within the virtual space.
[1645] (Claim 3)
[1646] 10. The system of claim 1, further comprising means for compressing the acquired motion data and transmitting the data to a server in real time.
[1647] "Example 1"
[1648] (Claim 1)
[1649] A means of scanning the user's entire body in 360 degrees;
[1650] means for preprocessing the acquired image data and transmitting the same to a server;
[1651] A means for generating a 3D avatar using a generative AI model based on the received image data;
[1652] A means for transmitting the generated 3D avatar to a user's device;
[1653] a means for a user to use a motion capture device to capture motion data;
[1654] means for analyzing and converting the acquired motion data into a suitable format;
[1655] means for transmitting the converted motion data to a server in real time;
[1656] A method for reproducing the movements of a 3D avatar in a virtual space based on real-time movement data, and
[1657] means for synchronizing motion and position information between users in a virtual space;
[1658] means for transmitting and receiving voice data for users to communicate in the virtual space;
[1659] A system including:
[1660] (Claim 2)
[1661] 10. The system of claim 1, further comprising: means for compressing the user's motion data and transmitting the data to the server in real time.
[1662] (Claim 3)
[1663] 2. The system according to claim 1, further comprising means for transmitting the user's voice data to the server in real time and transferring it to another terminal.
[1664] "Application Example 1"
[1665] (Claim 1)
[1666] A means of scanning the user's entire body in 360 degrees;
[1667] means for preprocessing the acquired image data and transmitting the same to a server;
[1668] A means for generating a 3D avatar using a generative AI model based on the received image data;
[1669] A means for transmitting the generated 3D avatar to a user's device;
[1670] a means for using a motion capture device to acquire motion data from a user;
[1671] means for transmitting the acquired motion data to a server in real time;
[1672] A method for reproducing the movements of a 3D avatar in a virtual space based on real-time movement data, and
[1673] means for synchronizing motion and position information between users in a virtual space;
[1674] means for transmitting and receiving voice data for users to communicate in the virtual space;
[1675] a means for providing a system in a virtual store that allows users to try on products online using a 3D avatar;
[1676] A system including:
[1677] (Claim 2)
[1678] 10. The system of claim 1, further comprising means for analyzing the user's motion data and converting it into a suitable format within the virtual space.
[1679] (Claim 3)
[1680] 10. The system of claim 1, further comprising means for compressing the acquired motion data and transmitting the data to a server in real time.
[1681] "Example 2: Combining Emotion Engines"
[1682] (Claim 1)
[1683] A means of scanning the user's entire body in 360 degrees;
[1684] means for preprocessing the acquired image data and transmitting the same to a server;
[1685] a means for generating a 3D avatar based on the received image data using a machine learning model;
[1686] means for transmitting the generated 3D avatar to a user device;
[1687] a means for using a motion detection device to obtain motion data by a user;
[1688] means for transmitting the acquired motion data to a server in real time;
[1689] A method for reproducing the movements of a 3D avatar in a virtual space based on real-time movement data, and
[1690] means for synchronizing motion and position information between users in a virtual space;
[1691] means for transmitting and receiving voice data for users to communicate in the virtual space;
[1692] A means for analyzing the user's facial expression data and voice data to recognize and reflect emotions;
[1693] A system including:
[1694] (Claim 2)
[1695] 10. The system of claim 1, further comprising means for analyzing the user's motion data and converting it into a suitable format within the virtual space.
[1696] (Claim 3)
[1697] 10. The system of claim 1, further comprising means for compressing the acquired motion data and facial expression data and transmitting the compressed data to a server in real time.
[1698] "Application example 2 when combining emotion engines"
[1699] (Claim 1)
[1700] A means of scanning the user's entire body in 360 degrees;
[1701] means for preprocessing the acquired image data and transmitting the same to a server;
[1702] A means for generating a 3D avatar using a generative AI model based on the received image data;
[1703] A means for transmitting the generated 3D avatar to a user's device;
[1704] a means for using a motion capture device to acquire motion data from a user;
[1705] means for transmitting the acquired motion data to a server in real time;
[1706] A method for reproducing the movements of a 3D avatar in a virtual space based on real-time movement data, and
[1707] means for synchronizing motion and position information between users in a virtual space;
[1708] means for transmitting and receiving voice data for users to communicate in the virtual space;
[1709] means for capturing facial expression data of a user and analyzing emotions using an emotion engine;
[1710] Based on the analyzed emotional data, a method is provided to reflect the facial expressions and movements of the user's 3D avatar in real time within the virtual space.
[1711] A system including:
[1712] (Claim 2)
[1713] 10. The system of claim 1, further comprising means for analyzing the user's motion data and facial expression data and converting them into a suitable format in the virtual space.
[1714] (Claim 3)
[1715] 10. The system of claim 1, further comprising means for compressing the acquired motion data and facial expression data and transmitting the compressed data to a server in real time. [Explanation of symbols]
[1716] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of scanning the user's entire body in 360 degrees; means for preprocessing the acquired image data and transmitting the same to a server; A means for generating a 3D avatar using a generative AI model based on the received image data; A means for transmitting the generated 3D avatar to a user's device; a means for using a motion capture device to acquire motion data from a user; means for transmitting the acquired motion data to a server in real time; A method for reproducing the movements of a 3D avatar in a virtual space based on real-time movement data, and means for synchronizing motion and position information between users in a virtual space; means for transmitting and receiving voice data for users to communicate in the virtual space; A system including:
2. 2. The system according to claim 1, further comprising means for analyzing the user's motion data and converting it into a suitable format in the virtual space.
3. 10. The system of claim 1, further comprising means for compressing the acquired motion data and transmitting the data to a server in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A