Method for processing real-time three-dimensional data on mobile device and apparatus therefor

The method addresses latency issues in mobile 3D telepresence by transmitting and reconstructing 3D data using RGB-depth data and deep learning, achieving efficient and accurate 3D data processing on mobile devices.

WO2025211936A1PCT designated stage Publication Date: 2025-10-09SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099590
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-03-06
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing 3D telepresence systems face limitations in mobile environments due to reliance on low-quality mobile sensors, leading to increased latency in real-time compression, transmission, and rendering of 3D data, and require stable wired communication and high-performance server computing.

Method used

A method for processing real-time 3D data by transmitting RGB-depth data from a transmitter to a receiver, applying deep learning models to identify facial feature points and head rotation information, and reconstructing 3D data using RGB-depth data, which includes encoding, decoding, and rendering operations on mobile devices.

Benefits of technology

Enables low-latency processing of real-time 3D data on mobile devices, achieving high accuracy in 3D data reconstruction with reduced latency and improved rendering speed compared to existing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099590_09102025_PF_FP_ABST
    Figure KR2025099590_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for processing real-time 3D data in a 3D telepresence system according to an embodiment of the present invention may comprise the operations of: acquiring, by a transmitter, multiple pieces of RGB-depth data associated with respective 2D images of a 2D video obtained by capturing a user's face; encoding, by the transmitter, the multiple pieces of RGB-depth data by using a video codec; applying, by the transmitter, a deep learning model to the multiple pieces of RGB-depth data to extract 3D modeling-based facial feature points and head orientation information regarding the user's face; transmitting, by the transmitter, the encoded RGB-depth data, the extracted facial feature points, and the head orientation information to a receiver via a wireless network; receiving, by the receiver, the encoded RGB-depth data from the transmitter; and decoding, by the receiver, the encoded RGB-depth data to generate 3D data regarding the user's face. Various other embodiments may be possible.
Need to check novelty before this filing date? Find Prior Art

Description

Method for processing real-time 3D data on a mobile device and device therefor

[0001] The present invention relates to a method for generating and transmitting 3D data in real time to a mobile device in a 3D telepresence system, and to a transmitter and receiver device therefor.

[0002] This application claims priority to Korean Patent Application No. 10-2024-0044168, filed on April 1, 2024, the entire contents of which are disclosed in the specification and drawings of the said application are incorporated herein by reference.

[0003] Meanwhile, the present invention was supported by the following national research and development project.

[0004] Assignment ID: 1711179518

[0005] Assignment Number: 2022R1A2C3008495

[0006] Ministry of Science and ICT

[0007] Project Management Agency Name: National Research Foundation of Korea

[0008] Research Project Name: Individual Basic Research (Ministry of Science and ICT)

[0009] Research Project Name: Hyper-Realistic Permanent Hybrid Telepresence Platform

[0010] Project Performing Institution Name: Seoul National University

[0011] Research period: January 1, 2023 - February 29, 2024

[0012] Assignment ID: 1711193554

[0013] Assignment Number: 2022-0-00531-002

[0014] Ministry of Science and ICT

[0015] Project Management Agency Name: Information and Communications Technology Planning and Evaluation Institute

[0016] Research Project Name: Broadcasting and Communications Industry Technology Development

[0017] Research Project Name: Development of Core Technologies for AI Networks Based on In-Network Computing

[0018] Project implementing organization name: Korea Advanced Institute of Science and Technology

[0019] Research period: January 1, 2023 - December 31, 2023

[0020] Assignment ID: 1711197751

[0021] Assignment Number: 00218601

[0022] Ministry of Science and ICT

[0023] Project Management Agency Name: National Research Foundation of Korea

[0024] Research Project Name: Group Research Support

[0025] Research Project Title: Foundation Model for Understanding 3D Human-Space Interaction

[0026] Project implementation organization name: Seoul National University Industry-Academic Cooperation Foundation

[0027] Research period: June 1, 2023 - February 29, 2024

[0028] 3D telepresence is a next-generation communication platform that allows users in physically separated spaces to feel as if they are in the same space. Typically, 3D telepresence systems reconstruct users and spaces as 3D data, share the reconstructed content bidirectionally, and provide immersive interactions within the content.

[0029] Meanwhile, since stable wired communication and high-performance server computing devices are essential for existing 3D telepresence platforms, there are problems in that the wired communication environment is limited and the possibility of confirmation is limited in order to build a 3D telepresence platform.

[0030] Furthermore, 3D telepresence platforms can be delivered via mobile devices, necessitating the provision of a communications platform designed specifically for mobile end systems. Current end systems rely on low-quality mobile sensors to process real-time 3D data, resulting in increased latency compared to wired communication environments in real-time compression, transmission, and rendering of 3D data.

[0031] The technical problem to be solved by the present invention is to provide a method for processing real-time 3D data by transmitting RGB-depth data of a 2D image captured in real time of a user's face to a receiver and composing 3D data using the RGB-depth data transmitted from the transmitter, and a device therefor.

[0032] Another technical problem to be solved by the present invention is to provide a method for processing real-time 3D data and a device therefor, which applies RGB-depth data to a deep learning model to identify facial feature points and head rotation information for a user's face and construct 3D data based on the identified information.

[0033] The technical problems of the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0034] A method for processing real-time 3D data in a 3D telepresence system according to an embodiment of the present invention for achieving the above technical task may include an operation of acquiring, by a transmitter, a plurality of RGB-depth data related to each 2D image of a 2D video that captures a user's face, an operation of encoding, by the transmitter, the plurality of RGB-depth data using a video codec, an operation of transmitting, by the transmitter, the encoded RGB-depth data and the extracted facial feature points and head direction information to a receiver through a wireless network, an operation of receiving, by the receiver, the encoded RGB-depth data from the transmitter, an operation of decoding, by the receiver, the encoded RGB-depth data to generate 3D data for the user's face, and an operation of rendering, by the receiver, the generated 3D data.

[0035] A method according to one embodiment of the present invention can be sequentially performed by the transmitter and the receiver for each of a plurality of 2D images included in the 2D video.

[0036] A method in a transmitter according to one embodiment of the present invention may further include an operation of acquiring a plurality of RGB-depth data by mapping RGB data acquired from an image sensor of the sensor module and a plurality of depth data acquired from a plurality of depth sensors of the sensor module installed at different locations at a time when a specific 2D image is acquired by the transmitter.

[0037] A method in a transmitter according to one embodiment of the present invention may further include an operation of extracting facial feature points and head direction information based on 3D modeling of the user's face by applying a deep learning model to the plurality of RGB-depth data by the transmitter, and an operation of transmitting the extracted facial feature points and head direction information together with the encoded RGB-depth data to the receiver.

[0038] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may include an operation of decoding the encoded RGB-depth data by the receiver to confirm the plurality of RGB-depth data and an operation of rendering the generated 3D data through a display.

[0039] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of generating 2D facial feature points for the face of the user by performing 3D modeling on the plurality of RGB-depth data, and an operation of generating 3D facial feature points for the plurality of RGB-depth data using the generated 2D facial feature points and the plurality of depth data.

[0040] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of aligning the generated 2D and 3D facial feature points and an operation of aligning the plurality of RGB-depth data by applying an inverse transformation of the received head rotation information to the aligned RGB-depth data.

[0041] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of merging the aligned plurality of RGB-depth data and converting them into the 3D data.

[0042] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of merging a plurality of RGB-depth data acquired at different locations using a volumetric fusion technique and an operation of converting the merged RGB-depth data into the 3D data using a marching cube technique.

[0043] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of generating mask data of a corresponding 2D image using the plurality of RGB-depth data.

[0044] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of identifying a first face part that does not change in real time and a second face part that changes for a user's face from the plurality of RGB-depth data, and an operation of generating an area connecting the first face part and the second face part as the mask data.

[0045] A method for processing real-time 3D data in a receiver according to one embodiment of the present invention may further include an operation of generating the 3D data using the second face portion among the plurality of RGB-depth data.

[0046] A receiver for processing real-time 3D data according to one embodiment of the present invention may include a receiving module for receiving, from a transmitter, RGB-depth data encoded from a 2D video that captures a user's face, a decoder for decoding the encoded RGB-depth data, a 3D configuration unit for identifying a plurality of RGB-depth data for a specific 2D image of the 2D video from the decoded RGB-depth data, and generating 3D data for the user's face using the plurality of RGB-depth data, a rendering unit for rendering the generated 3D data, and a display for outputting the rendered 3D data.

[0047] According to the present invention as described above, the method for processing real-time 3D data and the device therefor have the effect of processing real-time 3D data using a mobile device in a 3D telepresence system.

[0048] In addition, the present invention has the effect of processing 3D data with low latency by applying RGB-depth data to a deep learning model to identify facial feature points and head rotation information for a user's face and constructing 3D data based on the identified information.

[0049] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0050] FIG. 1 is a diagram illustrating a 3D telepresence system according to one embodiment of the present invention.

[0051] FIG. 2 is a block diagram illustrating a transmitter configuration according to one embodiment of the present invention.

[0052] FIG. 3 is a block diagram illustrating a receiver configuration according to one embodiment of the present invention.

[0053] FIG. 4 is a flowchart illustrating an operation of transmitting a 2D image to a receiver in real time from a transmitter according to one embodiment of the present invention.

[0054] FIG. 5 is a flowchart illustrating an operation of processing a 2D image received in real time into 3D data in a receiver according to one embodiment of the present invention.

[0055] FIG. 6 is a diagram illustrating an operation of rendering input data into 3D data in a receiver according to one embodiment of the present invention.

[0056] FIG. 7 is a diagram illustrating an operation of performing 3D modeling using a 2D image in a receiver according to one embodiment of the present invention.

[0057] FIG. 8 is a diagram illustrating a configuration of RGB-depth data combined in a receiver according to one embodiment of the present invention.

[0058] FIG. 9 is a diagram illustrating an operation of performing data rendering using RGB-depth data in a receiver according to one embodiment of the present invention.

[0059] FIG. 10 is a graph showing the results of performance evaluation according to 3D data processing in a transmitter, receiver, and wireless network according to one embodiment of the present invention.

[0060] Figure 11 is a graph showing the results of a performance evaluation of a 3D data processing method according to one embodiment of the present invention compared to existing techniques.

[0061] FIG. 12 is a flowchart illustrating an operation of processing real-time 3D data in a 3D telepresence system according to one embodiment of the present invention.

[0062] The purpose, technical configuration, and resulting operational effects of the present invention will be more clearly understood through the following detailed description based on the drawings attached to the specification of the present invention. Reference will now be made to the accompanying drawings, which will further describe embodiments of the present invention.

[0063] The embodiments disclosed herein should not be construed or used to limit the scope of the present invention. Those skilled in the art will readily appreciate that the descriptions of the embodiments herein, including the embodiments, have a wide range of applications. Therefore, any embodiments described in the detailed description of the present invention are intended to serve as illustrative examples to better illustrate the present invention and are not intended to limit the scope of the present invention to the embodiments.

[0064] The functional blocks depicted in the drawings and described below are merely examples of possible implementations. Other implementations may utilize other functional blocks without departing from the spirit and scope of the detailed description. Furthermore, while one or more functional blocks of the present invention are depicted as individual blocks, one or more of the functional blocks of the present invention may be a combination of various hardware and software configurations that perform the same function.

[0065] Additionally, the expression "including certain components" is an "open" expression, simply indicating the presence of those components, and should not be construed as excluding additional components.

[0066] Furthermore, when it is said that a component is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.

[0067] Hereinafter, detailed embodiments of the present invention will be described with reference to the drawings.

[0068] FIG. 1 is a diagram illustrating a 3D telepresence system according to one embodiment of the present invention.

[0069] Referring to FIG. 1, a 3D telepresence system includes a transmitter (101) and a receiver (102), and the transmitter (101) and the receiver (102) can perform communication connection via a wireless network (103). For example, the transmitter (101) and the receiver (102) can be mobile devices such as a smartphone or tablet PC to establish a 3D telepresence in a mobile environment.

[0070] According to one embodiment of the present invention, the transmitter (101) can acquire a 2D video including a plurality of 2D images captured in real time of a user's face. In addition, the transmitter (101) can acquire RGB data and first to third depth data for each 2D image when acquiring each 2D image.

[0071] According to one embodiment of the present invention, the transmitter (101) can perform encoding to compress each of three depth data and RGB data for a specific 2D image using a video codec. Accordingly, the transmitter (101) can obtain first RGB-depth data, second RGB-depth data, and third RGB-depth data for one 2D image.

[0072] According to one embodiment of the present invention, the transmitter (101) can apply a deep learning model to each of the RGB-depth data to extract 2D / 3D facial landmarks and head direction information indicating the head rotation direction for the user.

[0073] According to one embodiment of the present invention, the transmitter (101) can perform 2D streaming for the 2D video. For example, when RGB-depth data for a first 2D image among the 2D video is encoded in real time, the transmitter (101) transmits each encoded RGB-depth data to a receiver (102) in real time via a wireless network (103), and performs real-time encoding and transmission for a second 2D image, which is the next image of the first 2D image, to perform 2D streaming.

[0074] According to one embodiment of the present invention, the receiver (102) can check a 2D video including a 2D image received in real time from a transmitter (101) via a wireless network (103). For example, each 2D image can be encoded and received by the transmitter (101).

[0075] According to one embodiment of the present invention, the receiver (102) decodes the encoded 2D video to check first to third RGB-depth data for each 2D image and to check the user's facial feature points and head direction information.

[0076] According to one embodiment of the present invention, the receiver (102) can reconstruct a 3D image by combining the first to third RGB-depth data decoded from the 2D video. For example, the receiver (102) can use the identified facial feature points to generate an interpolation area, which is mask data representing a first face portion that does not change in real time, a second face portion that changes, and an area connecting the first face portion and the second face portion, for the user's face in the 3D video. Thereafter, the receiver (102) can combine data corresponding to the second face portion from the first to third RGB-depth data to reconstruct the 3D image.

[0077] According to one embodiment of the present invention, the receiver (102) can perform 3D modeling by applying the received 2D image to a pre-designated reference coordinate system to output a 3D model according to the input of the 2D image. For example, the receiver (102) can output a 3D model according to the input by transforming the input data into a 3D model and transforming the transformed 3D model into a 3DMM template, which is the reference coordinate system.

[0078] According to one embodiment of the present invention, the receiver (102) can perform 3D modeling by applying a deep learning model to the received 2D image. As a result of performing the 3D modeling, the receiver (102) can generate 2D 2D facial feature points using the received first to third depth data, and can generate 3D 3D facial feature points further based on the generated 2D facial feature points on the first to third depth data. Thereafter, the receiver (102) can align the first to third RGB-depth data by applying an inverse transformation of the head rotation information received from the transmitter (101) to the 2D / 3D facial feature points.

[0079] According to one embodiment of the present invention, the receiver (102) can generate 3D facial feature points for a plurality of RGB-depth data (or, first to third RGB-depth data) by using 2D facial feature points and a plurality of depth data generated by performing 3D modeling. Thereafter, the receiver (102) can align the generated 2D facial feature points and the 3D facial feature points, and then apply an inverse transformation of the received head rotation information to the aligned RGB-depth data to align the plurality of depth data.

[0080] According to one embodiment of the present invention, the receiver (102) can merge the aligned first to third RGB-depth data to generate 3D mesh data. For example, the receiver (102) can merge the first to third RGB-depth data acquired at different locations using a volumetric fusion technique, and convert the merged RGB-depth data into 3D mesh data using a marching cube technique.

[0081] According to one embodiment of the present invention, the receiver (102) can perform rendering of the 3D mesh data into a 3D image and control the rendered 3D image to be output.

[0082] FIG. 2 is a block diagram illustrating a transmitter configuration according to one embodiment of the present invention.

[0083] Referring to FIG. 2, the transmitter (101) may include a sensor module (111), an encoder (112), and a transmission module (113). In addition, the transmitter (101) includes a processor, and the overall 3D data processing operation in the transmitter (101) may be performed by the processor.

[0084] According to one embodiment of the present invention, the sensor module (111) may include an image sensor that acquires a plurality of 2D images captured in real time of a user's face, and first to third depth sensors provided at different locations of the transmitter (101). For example, the sensor module (111) may check a 2D video including the plurality of 2D images.

[0085] According to one embodiment of the present invention, when each 2D image is acquired, the sensor module (111) can acquire RGB data for the corresponding 2D image through the image sensor and first to third depth data through the first to third depth sensors.

[0086] According to one embodiment of the present invention, the sensor module (111) can transmit RGB data and first to third depth data acquired from the image sensor and a plurality of depth sensors to the encoder (112). In addition, the sensor module (111) may be provided in multiple numbers and may include an image sensor module and first to third depth sensor modules corresponding to the image sensor and the plurality of depth sensors, respectively. At this time, each of the image sensor module and the first to third depth sensor modules can acquire RGB data or depth data and transmit the acquired data to the encoder (112).

[0087] According to one embodiment of the present invention, the encoder (112) can perform encoding to compress each of three depth data and RGB data for a specific 2D image using a video codec. Accordingly, the encoder (112) can obtain first RGB-depth data, second RGB-depth data, and third RGB-depth data for one 2D image.

[0088] According to one embodiment of the present invention, the encoder (112) can apply a deep learning model to each of the RGB-depth data to extract 2D and 3D facial feature points and head direction information indicating the head rotation direction for the user.

[0089] According to one embodiment of the present invention, the transmission module (113) can perform 2D streaming for the 2D video. For example, when RGB-depth data for a first 2D image among the 2D video is encoded in real time, the transmission module (113) transmits each encoded RGB-depth data to a receiver (102) in real time via a wireless network (103), and performs real-time encoding and transmission for a second 2D image, which is the next image of the first 2D image, to perform 2D streaming.

[0090] FIG. 3 is a block diagram illustrating a receiver configuration according to one embodiment of the present invention.

[0091] Referring to FIG. 3, the receiver (102) may include a receiving module (121), a decoder (122), a 3D configuration unit (123), and a rendering unit (124). In addition, the receiver (102) includes a processor, and the overall 3D data processing operation in the receiver (102) may be performed by the processor.

[0092] According to one embodiment of the present invention, the receiving module (121) can check a 2D video including a 2D image received in real time from the transmitter (101) via a wireless network (103). For example, each 2D image can be encoded and received by the transmitter (101).

[0093] According to one embodiment of the present invention, the decoder (122) decodes the encoded 2D video to check first to third RGB-depth data for each 2D image and to check the user's facial feature points and head direction information.

[0094] According to one embodiment of the present invention, the 3D configuration unit (123) can combine the first to third RGB-depth data decoded from the 2D video to configure a 3D image.

[0095] According to one embodiment of the present invention, the 3D configuration unit (123) can use the identified facial feature points to generate an interpolation area, which is mask data representing a first face portion that does not change in real time, a second face portion that changes, and an area connecting the first face portion and the second face portion, for the user's face in the 3D video. Thereafter, the 3D configuration unit (123) can combine data corresponding to the second face portion from the first to third RGB-depth data to configure the 3D image.

[0096] According to one embodiment of the present invention, the 3D configuration unit (123) can perform 3D modeling on the plurality of RGB-depth data to generate 2D and 3D facial feature points for the user's face. For example, the receiver (102) can apply an inverse transformation of the head rotation information to the generated facial feature points to align the plurality of RGB-depth data, merge the aligned plurality of RGB-depth data, and convert the merged plurality of RGB-depth data into the 3D data.

[0097] According to one embodiment of the present invention, the 3D configuration unit (123) can generate 3D facial feature points for a plurality of RGB-depth data (or, first to third RGB-depth data) by using 2D facial feature points and a plurality of depth data generated by performing 3D modeling. Thereafter, the 3D configuration unit (123) can align the generated 2D facial feature points and the 3D facial feature points, and then apply an inverse transformation of the received head rotation information to the aligned RGB-depth data to align the plurality of depth data.

[0098] According to one embodiment of the present invention, the 3D configuration unit (123) can acquire 3D data by aligning 3D facial feature points acquired through RGB-depth data combined with a rendered image. At this time, the 3D configuration unit (123) can generate the 3D image using only the second facial portion among the plurality of RGB-depth data.

[0099] According to one embodiment of the present invention, the 3D configuration unit (123) can merge the aligned RGB-depth data to generate 3D mesh data. For example, the 3D configuration unit (123) can merge the first to third RGB-depth data acquired from different locations using a volumetric fusion technique, and convert the merged RGB-depth data into 3D mesh data using a marching cube technique.

[0100] According to one embodiment of the present invention, the rendering unit (124) can control the 3D mesh data output through the 3D configuration unit (123) to be rendered into a 3D image and output through the display (125).

[0101] According to one embodiment of the present invention, the rendering unit (124) can apply the combined RGB-depth data to a suitable model (e.g., DNN) to render it into a 3D image. For example, the DNN model can utilize coefficients related to a 3DMM, head rotation information, or camera parameters (e.g., RGB data, multiple depth data) for image rendering.

[0102] FIG. 4 is a flowchart illustrating an operation of transmitting a 2D image to a receiver in real time from a transmitter according to one embodiment of the present invention.

[0103] According to one embodiment of the present invention, the transmitter (101) can obtain a 2D video including a 2D image captured in real time through the sensor module (111).

[0104] Referring to FIG. 4, in operation S110, the transmitter (101) can acquire a plurality of RGB-depth data corresponding to each 2D image through the image sensor and the first to third depth sensors of the sensor module (111). For example, the plurality of RGB-depth data may include RGB data when the corresponding 2D image is acquired and first to third RGB-depth data to which first to third depth data corresponding to each of the RGB data are mapped.

[0105] In operation S120, the transmitter (101) can encode the plurality of RGB-depth data using a video codec. By performing the encoding, the transmitter (101) can confirm RGB-depth data in which the plurality of RGB-depth data are combined.

[0106] In operation S130, the transmitter (101) can extract 2D / 3D facial feature points and head direction information from the combined RGB-depth data based on 3D modeling.

[0107] In operation S140, the transmitter (101) can transmit the combined RGB-depth data and the extracted facial feature points and head direction information to the receiver (102).

[0108] According to one embodiment of the present invention, the transmitter (101) can sequentially perform the above-described operations S210 to S240 for each of a plurality of 2D images included in the 2D video.

[0109] FIG. 5 is a flowchart illustrating an operation of processing a 2D image received in real time into 3D data in a receiver according to one embodiment of the present invention.

[0110] Referring to FIG. 5, in operation S210, the receiver (102) can receive combined RGB-depth data from the transmitter (101).

[0111] In operation S220, the receiver (102) can decode the combined RGB-depth data. For example, by performing the decoding, the receiver (102) can identify a plurality of RGB-depth data for a specific 2D image compressed in the combined RGB-depth data, 2D / 3D facial feature points for the user's face captured in the 2D image, and head rotation information.

[0112] In operation S230, the receiver (102) can generate mask data of the 2D image using the plurality of RGB-depth data and facial feature points. For example, the receiver (102) can identify a first face part that does not change in real time and a second face part that changes for the user's face from the plurality of RGB-depth data, and generate an area connecting the first face part and the second face part as the mask data.

[0113] According to one embodiment of the present invention, the receiver (102) can combine a plurality of RGB-depth data, and identify a facial part that does not change in the combined RGB-depth data based on the identified facial feature points as the first facial part, and identify a facial part that changes as the second facial part. Thereafter, the receiver (102) can construct a 3D image using only the data corresponding to the second facial part in the combined RGB-depth data.

[0114] In operation S240, the receiver (102) can align the plurality of RGB-depth data to a preset reference coordinate system (e.g., a 3D morphable model (3DMM) template). For example, the receiver (102) can confirm RGB-depth data aligned to the reference coordinate system by applying the plurality of RGB-depth data to the reference coordinate system.

[0115] According to one embodiment of the present invention, the receiver (102) can generate 3D facial feature points for a plurality of RGB-depth data by using 2D facial feature points and a plurality of depth data generated by performing 3D modeling. Thereafter, the receiver (102) can align the generated 2D facial feature points and the 3D facial feature points, and then apply an inverse transformation of the received head rotation information to the aligned RGB-depth data to align the plurality of depth data.

[0116] In operation S250, the receiver (102) can generate 3D data using the second face portion from the aligned RGB-depth data.

[0117] According to one embodiment of the present invention, the receiver (102) can perform 3D modeling on the plurality of RGB-depth data to generate 2D and 3D facial feature points for the user's face. For example, the receiver (102) can apply an inverse transformation of the head rotation information to the generated facial feature points to align the plurality of RGB-depth data, merge the aligned plurality of RGB-depth data, and convert the merged plurality of RGB-depth data into the 3D data.

[0118] According to one embodiment of the present invention, the receiver (102) can apply the combined RGB-depth data to a suitable model (e.g., DNN) to render it into a 3D image. For example, the DNN model can utilize coefficients related to a 3DMM, head rotation information, or camera parameters (e.g., RGB data, multiple depth data) for image rendering.

[0119] According to one embodiment of the present invention, the receiver (102) can acquire 3D data by aligning 3D facial feature points acquired through RGB-depth data combined with a rendered image (720). At this time, the receiver (102) can generate the 3D image using only the second facial portion among the plurality of RGB-depth data.

[0120] In operation S260, the receiver (102) can render the generated 3D data. Thereafter, the receiver (102) can output the rendered 3D image through the display (125).

[0121] According to one embodiment of the present invention, the receiver (102) renders a 3D image using 3D facial feature points acquired through combined RGB-depth data, thereby enabling alignment with higher accuracy than when performing image rendering using 3D facial feature points acquired through performing 3D modeling.

[0122] FIG. 6 is a diagram illustrating an operation of rendering input data into 3D data in a receiver according to one embodiment of the present invention.

[0123] Referring to FIG. 6, the receiver (102) can receive a 2D image (410) and first to third RGB-depth data of the 2D image (410) from the transmitter (101). In addition, the receiver (102) can receive facial feature points and head direction information of the 2D image (510) extracted by the transmitter (101).

[0124] According to one embodiment of the present invention, the receiver (102) can perform 3D modeling by applying a deep learning model to the received 2D image (510). As a result of performing the 3D modeling, the receiver (102) can generate a 2D 2D facial feature point (421) using the received first to third depth data, and can generate a 3D 3D facial feature point (422) further based on the generated 2D facial feature point (421) to the first to third depth data. Thereafter, the receiver (102) can align the first to third RGB-depth data by applying an inverse transformation of the head rotation information received from the transmitter (101) to the 2D / 3D facial feature point (421, 422).

[0125] According to one embodiment of the present invention, the receiver (102) can generate 3D facial feature points for a plurality of RGB-depth data (or, first to third RGB-depth data) by using 2D facial feature points and a plurality of depth data generated by performing 3D modeling. Thereafter, the receiver (102) can align the generated 2D facial feature points and the 3D facial feature points, and then apply an inverse transformation of the received head rotation information to the aligned RGB-depth data to align the plurality of depth data.

[0126] According to one embodiment of the present invention, the receiver (102) can merge the aligned first to third RGB-depth data to generate 3D mesh data. For example, the receiver (102) can merge the first to third RGB-depth data acquired from different locations using a volumetric fusion technique, and convert the merged RGB-depth data into 3D mesh data using a marching cube technique. Thereafter, the receiver (102) can render the converted 3D mesh data on the display (125).

[0127] FIG. 7 is a diagram illustrating an operation of performing 3D modeling using a 2D image in a receiver according to one embodiment of the present invention.

[0128] Referring to FIG. 7, when the receiver (102) receives a plurality of 2D images (511, 512) and facial expressions and head rotation information related to the images from the transmitter (101), the receiver (102) can confirm the plurality of 2D images (511, 512) as input data for configuring 3D data.

[0129] According to one embodiment of the present invention, the receiver (102) can use a pre-designated 3DMM template (501) as a reference coordinate system and output a 3D model (3DMM fitted to input) (521, 522) according to a plurality of 2D images (511, 512) input by applying the input data to the reference coordinate system. For example, the receiver (102) can output a 3D model (521, 522) according to the input by transforming the input data into a 3D model and applying transformation to the transformed 3D model with the 3DMM template (501) as the reference coordinate system.

[0130] According to one embodiment of the present invention, the receiver (102) can align 3D data with high mathematical accuracy (topologically consistent) with input data while reducing the amount of 3D data by performing 3D modeling in real time using 3DMM.

[0131] FIG. 8 is a diagram illustrating a configuration of RGB-depth data combined in a receiver according to one embodiment of the present invention.

[0132] According to one embodiment of the present invention, the receiver (102) can receive first to third RGB-depth data for 3D video from the transmitter (101). Thereafter, the receiver (102) can combine the first to third RGB-depth data.

[0133] Referring to FIG. 8, the receiver (102) can identify a first face portion (610) that does not change in real time for the user's face, a second face portion (620) that changes, and an interpolation area (620), which is mask data representing an area connecting the first face portion and the second face portion, from the combined RGB-depth data (600). Thereafter, the receiver (102) can use the data corresponding to the second face portion (620) from the combined RGB-depth data (600) to construct a 3D image.

[0134] FIG. 9 is a diagram illustrating an operation of performing data rendering using RGB-depth data in a receiver according to one embodiment of the present invention.

[0135] Referring to FIG. 9, the receiver (102) can identify combined RGB-depth data (710) using a plurality of RGB-depth data, and apply the combined RGB-depth data (710) to a suitable model (e.g., DNN) to render it into a 3D image (720). For example, the DNN model can use coefficients related to 3DMM, head rotation information, or camera parameters (e.g., RGB data, depth data) for image rendering.

[0136] According to one embodiment of the present invention, the receiver (102) can obtain a 3D image (720) by aligning 3D facial feature points (722) obtained through RGB-depth data (710) combined with a rendered image (720).

[0137] According to one embodiment of the present invention, the receiver (102) renders a 3D image (720) using 3D facial feature points (722) acquired through combined RGB-depth data (710), thereby performing alignment with higher accuracy than image rendering using 3D facial feature points (721) obtained through 3D modeling.

[0138] FIG. 10 is a graph showing the results of performance evaluation according to 3D data processing in a transmitter, receiver, and wireless network according to one embodiment of the present invention.

[0139] Referring to FIG. 10, a graph (1000) represents latency in a wireless network (103) of a sender (101), a receiver (102).

[0140] According to one embodiment of the present invention, a transmitter (101) simultaneously experiences delays in face detection (face det.), landmark extraction (landmark det.), and encoding (enc.) operations for a captured 2D video, and it can be confirmed that a delay time of 20 ms occurs in landmark extraction, which exhibits the longest delay time.

[0141] According to one embodiment of the present invention, a receiver (102) simultaneously experiences delays in operations such as decoding (dec.) the received 2D video, combining RGB-depth data (fusion), meshing with 3D mesh data, and rendering with 3D images. It can be confirmed that a delay time of 30 ms occurs in meshing with 3D mesh data, which exhibits the longest delay time.

[0142] In a wireless network (103) according to one embodiment of the present invention, it can be confirmed that a 2D streaming operation is performed and the total delay time for this is 30 ms.

[0143] Accordingly, the total delay time (Total) is 80 m, which is the sum of the delay time of 20 ms in the transmitter (101), the delay time of 30 ms in the wireless network (103), and the delay time of 30 ms in the receiver (102), and it can be confirmed that the terminal delay time of 100 ms or less is satisfied.

[0144] Additionally, the delay time for all actions is less than 30ms, satisfying the processing speed of 30 frames per second, and the delay time for the action of rendering 3D images is 3~4ms, satisfying the rendering speed of less than 16ms.

[0145] Figure 11 is a graph showing the results of a performance evaluation of a 3D data processing method according to one embodiment of the present invention compared to existing techniques.

[0146] Referring to FIG. 11, a graph (1100) represents the delay time for each operation of each technique when a fusion operation and a meshing operation are performed in the 3D data processing method of the present invention (Farfetch fusion) and an existing technique (e.g., Multiview fusion).

[0147] The total delay time of the combining and meshing operation according to the 3D data processing method according to one embodiment of the present invention may be measured as 50 ms or less, and the total delay time of the combining and meshing operation according to the existing technique may be measured as 200 ms or more.

[0148] A 3D data processing method according to one embodiment of the present invention can process 3D data in real time quickly with a small delay time, as it exhibits a delay time of 25% or less compared to using existing techniques.

[0149] FIG. 12 is a flowchart illustrating an operation of processing real-time 3D data in a 3D telepresence system according to one embodiment of the present invention.

[0150] Referring to FIG. 12, in operation S310, the transmitter (101) can obtain multiple RGB-depth data related to each 2D image of the 2D video that captures the user's face.

[0151] According to one embodiment of the present invention, the transmitter (101) can acquire the plurality of RGB-depth data by mapping the RGB data acquired from the image sensor of the transmitter (101) and the plurality of depth data acquired from the plurality of depth sensors installed at different locations at the time when a specific 2D image is acquired.

[0152] In the S320 operation, the transmitter (101) can encode the plurality of RGB-depth data using a video codec.

[0153] In operation S330, the transmitter (101) can transmit RGB-depth data to the receiver (102). For example, the transmitter (101) can apply a deep learning model to the plurality of RGB-depth data to extract facial feature points and head direction information based on 3D modeling of the user's face, and transmit the extracted facial feature points and head direction information together with the encoded RGB-depth data to the receiver (102).

[0154] In operation S340, the receiver (102) can receive RGB-depth data from the transmitter (101).

[0155] In the S350 operation, the receiver (102) can generate 3D data using RGB-depth data.

[0156] According to one embodiment of the present invention, the receiver (102) can perform 3D modeling on the plurality of RGB-depth data to generate 2D and 3D facial feature points for the user's face, and align the plurality of RGB-depth data by applying an inverse transformation of the received head rotation information to the generated facial feature points.

[0157] According to one embodiment of the present invention, the receiver (102) can merge the plurality of aligned RGB-depth data and convert them into the 3D data. For example, the receiver (102) can merge the plurality of RGB-depth data acquired at different locations using a volumetric fusion technique, and convert the merged RGB-depth data into the 3D data using a marching cube technique.

[0158] According to one embodiment of the present invention, the receiver (102) can generate mask data of the corresponding 2D image using the plurality of RGB-depth data. For example, the receiver (102) can identify a first face part that does not change in real time and a second face part that changes for the user's face from the plurality of RGB-depth data, and can generate an area connecting the first face part and the second face part as the mask data by the receiver. Thereafter, the receiver (102) can generate the 3D data using the second face part among the plurality of RGB-depth data.

[0159] According to one embodiment of the present invention, the receiver (102) can decode the encoded RGB-depth data, confirm the plurality of RGB-depth data, and render the generated 3D data through a display.

[0160] According to one embodiment of the present invention, the above-described operations S310 to S350 can be sequentially performed for each of a plurality of 2D images included in the 2D video by the transmitter (101) and the receiver (102).

[0161] Although embodiments of the present invention have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.

Claims

1. A method for processing real-time 3D data in a 3D telepresence system, An operation of acquiring, by a transmitter, multiple RGB-depth data associated with each 2D image of a 2D video capturing a user's face; An operation of encoding the plurality of RGB-depth data using a video codec by the transmitter; An operation of transmitting the encoded RGB-depth data to a receiver via a wireless network by the transmitter; An operation of receiving, by the receiver, the encoded RGB-depth data from the transmitter; and A method comprising: generating 3D data for the user's face using the encoded RGB-depth data by the receiver; 2. In paragraph 1, The above method, A method characterized in that it is sequentially performed for each of a plurality of 2D images included in the 2D video by the transmitter and the receiver.

3. In paragraph 2, A method further comprising: an operation of acquiring a plurality of RGB-depth data by mapping RGB data acquired from an image sensor of the sensor module and a plurality of depth data acquired from a plurality of depth sensors of the sensor module installed at different locations at a time when a specific 2D image is acquired by the transmitter; 4. In paragraph 1, An operation of extracting facial feature points and head direction information based on 3D modeling of the user's face by applying a deep learning model to the plurality of RGB-depth data by the transmitter; and A method further comprising: an operation of transmitting the extracted facial feature points and head direction information to the receiver together with the encoded RGB-depth data.

5. In paragraph 3, An operation of decoding the encoded RGB-depth data by the receiver to confirm the plurality of RGB-depth data; and A method comprising: an operation of rendering the generated 3D data through a display; 6. In paragraph 5, An operation of generating 2D facial feature points for the user's face by performing 3D modeling on the plurality of RGB-depth data by the receiver; and A method further comprising: an operation of generating 3D facial feature points for the plurality of RGB-depth data using the generated 2D facial feature points and the plurality of depth data.

7. In paragraph 6, By the above receiver, An operation of aligning the generated 2D and 3D facial feature points; and A method further comprising: an operation of aligning the plurality of RGB-depth data by applying an inverse transformation of the received head rotation information to the aligned RGB-depth data.

8. In paragraph 7, A method further comprising: an operation of converting the aligned plurality of RGB-depth data into the 3D data by merging the plurality of aligned RGB-depth data by the receiver; 9. In paragraph 8, An operation of merging the plurality of RGB-depth data acquired from different locations using a volumetric fusion technique by the receiver; and A method further comprising: converting the merged RGB-depth data into the 3D data using a marching cube technique by the receiver; 10. In paragraph 5, A method further comprising: an operation of generating mask data of a corresponding 2D image using the plurality of RGB-depth data by the receiver; 11. In paragraph 10, An operation of identifying a first face part that does not change in real time and a second face part that changes for the user's face from the plurality of RGB-depth data by the receiver; and A method further comprising: an operation of generating an area connecting the first face portion and the second face portion as the mask data by the receiver; 12. In paragraph 11, A method further comprising: an operation of generating the 3D data by using the second face portion among the plurality of RGB-depth data by the receiver.

Citation Information

Patent Citations

  • Display cell, method of manufacturing the same and display device manufactured thereby

    KR1020200108200A

  • The brush roll for cleaning semiconductor wafers

    KR102786433B1