Method, apparatus, and program for streaming 3D objects

By transmitting simplified geometry, color, and alpha information in a container stream, the method addresses bandwidth issues in 3D image transmission, ensuring high-quality and efficient 3D image display on the client.

JP7829244B2Active Publication Date: 2026-03-13MAWARI CORP
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing methods for transmitting 3D images consume high bandwidth and compromise image quality due to the transmission of large video or pixel data.

Method used

A method and system for transmitting 3D objects by extracting color, alpha, and geometry information from the server, simplifying geometry, and encoding these into a container stream for transmission to the client, where the client decodes and reconstructs the 3D object.

Benefits of technology

Reduces data transmission bandwidth while maintaining high-quality 3D image display by using a container stream that is less data-intensive than conventional video or pixel data, ensuring smooth playback and reduced processing load on the client.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007829244000001
    Figure 0007829244000001
  • Figure 0007829244000002
    Figure 0007829244000002
  • Figure 0007829244000003
    Figure 0007829244000003
Patent Text Reader

Abstract

To provide a method, device and program capable of reducing the data transfer amount in transmitting a 3D image from a server to a client.SOLUTION: A method for transmitting a 3D object from a server to a client includes: extracting, from the 3D object on the server, color information, alpha information and geometry information; simplifying the geometry information; and encoding a stream, which includes the color information, the alpha information and the simplified geometry information, and transmitting the encoded stream from the server to the client.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method, apparatus, and program for streaming 3D (three-dimensional) objects.

Background Art

[0002] Conventionally, there is a technique of transmitting a 3D image from a server to a client for display. In this case, for example, on the server side, a method of converting a 3D image into a 2D image is used (Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The conventional problem to be solved is to reduce the bandwidth used for data transmission while maintaining the image quality in the transmission of 3D images.

Means for Solving the Problems

[0005] A method according to one aspect of the present disclosure is a method for transmitting a 3D object from a server to a client, including extracting color information, alpha information, and geometry information from the 3D object on the server, simplifying the geometry information, and encoding and transmitting a stream including the color information, the alpha information, and the simplified geometry information from the server to the client.

[0006] One aspect of the present disclosure is a method for reproducing a 3D object on a server on a client, comprising: receiving an encoded stream from the server containing color information, alpha information, and geometry information of the 3D object; decoding the stream and extracting the color information, alpha information, and geometry information therefrom; reproducing the shape of the 3D object based on the geometry information; and reconstructing the 3D object by projecting information obtained by combining the color information and the alpha information onto the reproduced shape of the 3D object.

[0007] One aspect of the present disclosure is a server comprising one or more processors and memory, wherein the one or more processors extract alpha information and geometry information from 3D objects on the server by executing instructions stored in the memory, simplify the geometry information, encode and transmit a stream containing the alpha information and the simplified geometry information from the server to a client.

[0008] One aspect of the present disclosure is a client comprising one or more processors and memory, wherein the one or more processors receive an encoded stream from a server containing color information, alpha information, and geometry information of a 3D object by executing instructions stored in the memory, decode the stream, extract the color information, alpha information, and geometry information therefrom, reproduce the shape of the 3D object based on the geometry information, and reconstruct the 3D object by projecting information obtained by combining the color information and the alpha information onto the reproduced shape of the 3D object.

[0009] A program according to one aspect of this disclosure includes instructions for a processor to perform any of the methods described above.

[0010] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or recording media, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]

[0011] By sending a container stream as described in this disclosure instead of sending video data or pixels from the server to the client to display a 3D image on the client, the amount of data sent per unit of time from the server to the client is reduced, improving the display quality and responsiveness of the 3D image on the client.

[0012] Further advantages and effects of one embodiment of this disclosure will be made apparent from the specification and drawings. Such advantages and / or effects are provided by several embodiments and features described in the specification and drawings, but not all of them are necessarily provided in order to obtain one or more identical features.

[0013] In this specification, for illustrative purposes, the transmission of 3D images (including video and / or still images) between a server and a client has been used as an example. However, the application of this disclosure is not limited to a client-server system and can be applied to transmission from one computer to another, which may be one or more computers. [Brief explanation of the drawing]

[0014] [Figure 1] This is a functional block diagram of the server and client as disclosed herein. [Figure 2] This flowchart illustrates the server-side processing of the data flow between the server and client, as explained in Figure 1. [Figure 3] This flowchart illustrates the data processing on the client side of the data flow between the server and client, as explained in Figure 1. [Figure 4]It is a flowchart that explains the processing of commands on the client side among the data flows between the server and the client described in FIG. 1. [Figure 5] It is a diagram depicting the data flow for displaying 3D scenes or 3D objects on the client side in a client-server system to which the present disclosure is applied. [Figure 6] It is a diagram showing the process of encoding and decoding geometric information according to the present disclosure. [Figure 7] It is a diagram showing the process of encoding and decoding color information / texture information according to the present disclosure. [Figure 8] It is a diagram showing data synchronization among geometry, color packets, metadata, and commands according to the present disclosure. [Figure 9] It is a diagram showing the decal method according to the present disclosure. [Figure 10] It is a schematic diagram showing an example of the hardware configuration of a client according to the present disclosure. [Figure 11] It is a schematic diagram showing an example of the hardware configuration of a server according to the present disclosure. [Figure 12] It is a schematic diagram showing an example of the configuration of an information processing system according to the present disclosure. [Figure 13] It is a schematic diagram representing the flow of processing on the server side according to the present disclosure. [Figure 14] It is a schematic diagram representing the flow of processing on the client side according to the present disclosure. [Figure 15] It is a diagram showing the arrangement of cameras used in the present disclosure. [Figure 16] It is a diagram showing the pixel configuration in the ARGB system used in the present disclosure.

Mode for Carrying Out the Invention

[0015] <1. 3D Streaming System Architecture> Figure 1 is a functional block diagram of the server and client according to this disclosure. The 3D streaming server 100 includes the functional configuration within the three-dimensional (3D) streaming server, and the 3D streaming client 150 includes the functional configuration within the 3D streaming client. The network 120 represents a wired or wireless network between the server 100 and the client 150.

[0016] One system covered by this disclosure generates a 3D image on the server side, and on the client side, it reconstructs and displays the 3D image based on the feature quantities of the 3D image received from the server. Client devices include any device with display and communication capabilities, such as smartphones, mobile phones, tablets, laptops, smart glasses, head-mounted displays, and headsets. Here, the feature quantities include color information, alpha information, or geometry information of the 3D image.

[0017] <1.2 Processing on the 3D Streaming Server Side> The upper half of Figure 1 is a functional block diagram illustrating the processing in the 3D streaming server 100. The network packet receiving unit 108 receives packets containing commands and / or data from the client 150 via a wired or wireless network 120. The network packet receiving unit 108 extracts the commands and / or data received from the client from the packets and sends the extracted data to the received data processing unit 101, which processes the commands and / or data from the client. The received data processing unit 101, having received the extracted commands and / or data from the client, further extracts the necessary commands and / or data from the received data and sends them to the 3D scene data generation unit 102. Next, the 3D scene data generation unit 102 processes and modifies the 3D scene (or 3D object) data that the server has that corresponds to the request from the client, according to the request sent from the client 150. Next, the extraction unit 103, having received commands and / or data from the 3D scene data generation unit 102, extracts the necessary data from the updated 3D scene data according to the commands from the client and sends them to the 3D stream conversion / encoding unit 104. The 3D stream conversion / encoding unit 104 converts the data received from the extraction unit 103 into a 3D stream and encodes it to generate a 3D stream 105. The 3D stream 105 is then sent to the network packet configuration unit 106, which generates network packets. The network packets are sent to the network packet transmission unit 107. The network packet transmission unit 107 transmits the received network packets to one or more clients 150 via the wired or wireless network 120.

[0018] <1.3 3D Streaming Client-Side Processing> The lower half of Figure 1 is a functional block diagram illustrating the processing in the 3D streaming client 150. The network packet receiving unit 152, which receives packets from the server 100 via the wired or wireless network 120, extracts the encoded 3D stream from the packets and sends it to the 3D stream decoding unit 154. The 3D stream decoding unit 154, having received the encoded 3D stream, decodes the 3D stream and sends the decoded 3D stream to the 3D scene reconstruction unit 155. The 3D scene reconstruction unit 155, having received the decoded 3D stream, reconstructs a 3D scene (or 3D object) from the 3D stream received and decoded from the server 100 and sends the reconstructed 3D scene to the display unit 156. The display unit 156 displays the reconstructed 3D scene and presents it to the user.

[0019] Meanwhile, 3D display (update) requests from the 3D streaming client 150 are sent from the application data output unit 153 to the network packet transmission unit 151. The 3D display (update) request data generated by the application data output unit 153 may include commands such as user input, camera / device position changes, or display updates. Upon receiving the 3D display request, the network packet transmission unit 151 performs necessary processing such as encoding and packetization and sends the 3D display (update) request to the 3D streaming server 100 via the wired or wireless network 120.

[0020] The network packet configuration unit 106 and network packet transmission unit 107 included in the server 100 described above, and the network packet receiving unit 152 and network packet transmission unit 151 included in the client 150 described above, may, for example, be based on the corresponding transmit / receive modules of existing open-source software with necessary modifications, or they may be created specifically from scratch.

[0021] Figure 2 is a flowchart illustrating the server-side processing of the data flow between the server and client described in Figure 1. Processing starts at START (901). First, the network packet receiving unit 108, described in Figure 1, receives packets from the client, including commands to rewrite the 3D scene (902). Next, the received data processing unit 101, described in Figure 1, processes the received commands and outputs the result (903). Next, the 3D scene data generation unit 102, described in Figure 1, generates 3D scene data according to the received commands (904). Next, the extraction unit 103 in Figure 1 extracts the features of the 3D scene (905). Here, features refer to data such as geometry, color, metadata, sound, and commands included in the container stream, which will be described later. Next, the 3D stream conversion / encoding unit 104 in Figure 1 converts the data containing the 3D features into a 3D stream and encodes it (906). Next, the network packet configuration unit 106 in Figure 1 configures a network packet from the 3D stream (907). Then, the network packet transmission unit 107 in Figure 1 transmits the network packet (908). This completes the series of data transmission processes on the server side (909).

[0022] In Figure 2, as an example, the processes from steps 902 to 903 and from steps 904 to 908 are shown to be executed sequentially. However, the processes from steps 902 to 903 and from steps 904 to 908 may be executed in parallel, or the process may start from step 904.

[0023] Figure 3 is a flowchart illustrating the data processing on the client side of the data flow between the server and client described in Figure 1. Processing starts at START (1001). First, the network packet receiving unit 152, described in Figure 1, receives packets sent from the server 100 (1002). Next, the 3D stream decoding unit 154, described in Figure 1, decodes the received packets (1003) and extracts the features of the 3D scene. Next, the 3D scene reconstruction unit 155, described in Figure 1, reconstructs the 3D scene on the client using the features of the 3D scene (1004) and generates 3D scene data (1005). Next, the display unit 156, described in Figure 1, displays the reconstructed 3D scene and presents it to the user (1006). This completes the data processing on the client side (1007).

[0024] In Figure 3, as an example, the processes from steps 1002 to 1004 and from steps 1005 to 1006 are shown to be executed sequentially. However, the processes from steps 1002 to 1004 and from steps 1005 to 1006 may be executed in parallel, or the process may start from step 1005.

[0025] Figure 4 is a flowchart illustrating the command processing on the client side of the data flow between the server and client described in Figure 1. Processing starts at START (1101). The application data output unit 153, described in Figure 1, outputs commands such as 3D scene rewriting commands from an image processing application (1102). The network packet transmission unit 151, described in Figure 1, receives the commands from the application data output unit 153, converts them into packets, and transmits the converted packets to the wired or wireless network 120 (1103). This completes the data processing on the client side (1104).

[0026] In Figure 4, as an example, the processes in step 1102 and step 1103 are shown to be executed sequentially, but the processes in step 1102 and step 1103 may also be executed in parallel.

[0027] <2. 3D Stream Format of This Disclosure> The main features of the 3D stream format disclosed in this information are as follows. The significance of this disclosure lies in the fact that it achieves these features using limited network bandwidth without compromising the quality of the 3D images displayed on the client side.

[0028] (1) The server generates the 3D stream. When generating the 3D stream on the server side, an available engine such as UE4 or Unity is used. Here, UE refers to the Unreal Engine, a game engine developed by Epic Games, and UE5 was announced in May 2020.

[0029] (2) Efficient transmission over a network is supported. Therefore, the amount of data transferred from the server to the client is less compared to conventional methods. To achieve this, this disclosure uses container streams.

[0030] (3) It can operate on a variety of devices. The target devices are, for example, devices that can use Unity (Android, Windows, iOS), WebGL, or UE4 or UE5 (Android, iOS, Windows).

[0031] (4) It is relatively lightweight compared to modern AR (augmented reality) devices. In other words, the processing load on the client side is smaller compared to conventional methods. This is due to the use of the container stream described in this disclosure.

[0032] (5) Support interaction (i.e., two-way communication). In other words, streaming is interactive. This is because commands can be sent and received between the client and the server.

[0033] To achieve the features described above, this disclosure has developed a proprietary container stream as a 3D stream transmitted between the server and the client. This proprietary container stream includes some of the following: geometry, color, metadata, sound, and commands.

[0034] (1) Geometry: Simplified 3D data relating to the outline of objects streamed from a 3D scene on the server. Geometry data is, for example, an array of polygon vertices used to represent the shape of an object.

[0035] (2) Color: This is the color data of an object captured by a camera at a specific location.

[0036] (3) Metadata: This is data that describes the 3D scene, environment, individual objects, data in the stream, etc.

[0037] (4) Sound: Sound (audio) data generated in the 3D scene on the server side or client side. Sound can be communicated bidirectionally between the server and the client.

[0038] (5) Commands: Commands are instructions that include server-side or client-side 3D scenes, system events, status messages, camera and user inputs, and client application events. Commands can be communicated bidirectionally between the server and the client.

[0039] In conventional systems, instead of the container stream described above as disclosed, the video data itself or the pixel data of each frame was sent. Here, a container stream refers to a block of data transferred between a server and a client, and is also called a data stream. The container stream is transferred over the network as a packet stream.

[0040] Conventional video data, or the pixel data of each frame, even when compressed, has a very large data transfer capacity per unit of time. This can cause transmission delays and latency if the network bandwidth between the server and client is not large enough, resulting in problems with smooth 3D image playback on the client side. On the other hand, in the system disclosed herein, the data container used for transfer between the server and client has a data size that is significantly smaller than that of conventional systems. Therefore, the minimum number of frames per unit of time can be secured without having to worry too much about the network bandwidth between the server and client, enabling smooth 3D image playback on the client side.

[0041] Figure 5 is a diagram illustrating the data flow for displaying a 3D scene or 3D object on the client side in a client-server system to which this disclosure is applied. The client side is shown as a terminal device 1221 such as a smartphone, and the terminal device 1221 and smart glasses 1210 are connected via wireless communication such as Wi-Fi or Bluetooth (registered trademark) or wired communication 1222. The smart glasses 1210 represent a view from the user's perspective. Person 1211-1 and cursor 1212-1 are projected onto the left eye of the client's smart glasses, and person 1211-2 and cursor 1212-2 are projected onto the right eye of the smart glasses. To the user of the smart glasses, the images of the right and left eyes are superimposed to create a three-dimensional person 1214 that appears a short distance away. The user of the client-side smart glasses 1210 can use cursor 1212 or other input means to perform operations on the person 1214 displayed in the client-side smart glasses 1210, such as moving, rotating, shrinking / enlarging, changing color / texture, or playing sounds. When such an operation is performed on an object (or scene) on the client, a command (or sound) 1213 or the like is sent from the client to the server via network 120.

[0042] A server that receives commands from a client via network 120 performs operations on the corresponding person 1202 image on a virtual screen 1201 within the application on the server, according to the received commands. Here, the server does not usually need to have a display device and deals with virtual images in a virtual space. Next, the server generates 3D scene data (or 3D object data) after performing the operations on the commands, and sends the extracted features as a container stream 1203 to the client via network 120. The client that receives the container stream 1203 from the server rewrites the data of the corresponding person 1214 on the client's virtual screen according to the geometry, color / texture, metadata, sound, and commands contained in the container stream 1203, and redisplays it, etc. In this example, the object is a person, but the object may be something other than a person, such as a building, car, animal, or still life, and the scene may contain one or more objects.

[0043] The following explains how the "geometry" and "color" data included in the container stream described above are processed, referring to Figures 6-8.

[0044] <3. Geometry Encoding and Decoding Process> Figure 6 shows the process of encoding and decoding geometry data according to this disclosure. In Figure 6, processes 201-205 are performed by the server, and processes 207-211 are performed by the client.

[0045] Each process described in Figure 6 and Figures 7 and 8 below is executed by a processor such as a CPU and / or GPU using the relevant program. The system covered by this disclosure may have either a CPU or a GPU, but for the sake of simplicity, the CPU and GPU will be collectively referred to as the CPU below.

[0046] <3.1 Server-side processing> Let's assume we have a scene with objects. Each object is captured by one or more depth cameras. Here, a depth camera is a camera with a built-in depth sensor that acquires depth information. Using depth cameras, it is possible to add depth information to the two-dimensional (2D) image acquired by a regular camera and obtain three-dimensional information. Here, for example, six depth cameras are used to obtain the complete geometric data of the scene. The camera configuration during shooting will be described later.

[0047] The server generates a streamed 3D object from the captured image and outputs the camera's depth information (201). Next, the camera's depth information is processed to generate a point cloud and outputs an array of points (202). This point cloud is converted into triangles (arrays of triangle vertices) that represent the actual geometry of the object, and a group of triangles is generated by the server (203). Here, a triangle is used as an example to represent the geometry, but other polygons may also be used.

[0048] Then, the geometric data is added to the stream using the array data of each vertex of the triangle group and compressed (204).

[0049] The server sends a container stream containing the compressed geometry data over network 120 (205).

[0050] <3.2 Client-side processing> The client receives a container stream containing compressed data, i.e., geometry data, from the server via network 120 (207). The client decompresses the received compressed data and extracts an array of vertices (208).

[0051] The client places the array of vertices from the decompressed data into a managed geometry data queue and corrects the order of the frame sequence, which may have been distorted during transmission over the network (209). The client reconstructs the scene objects based on the correctly realigned frame sequence (210). The client displays the reconstructed client-side 3D objects on the display (211).

[0052] Geometry data is stored in a managed geometry data queue and synchronized with other data received in the stream (209). This synchronization will be described later using Figure 8.

[0053] Clients to which this disclosure applies generate a mesh based on the received array of vertices. In other words, only the array of vertices is sent from the server to the client as geometry data, so the amount of data per unit time for the vertex array is usually considerably less than that for video data and frame data. On the other hand, other conventional options involve applying a large number of triangles to the data of a given mesh, which has problems as it requires a large amount of processing on the client side.

[0054] Servers to which this disclosure applies will send data only for the parts of a scene (which typically contains one or more objects) that need to be changed (e.g., specific objects) to the client, and will not send data for the parts of the scene that do not need to be changed. This also reduces the amount of data transmitted from the server to the client when a scene is changed.

[0055] The systems and methods to which this disclosure applies are based on the premise that an array of polygon mesh vertices is transmitted from the server to the client. Although the explanation has been based on the premise of triangular polygons, the shape of the polygon is not limited to triangles and may be quadrilateral or other shapes.

[0056] <4. Color / Texture Encoding and Decoding Process> Figure 7 shows the color information / texture information encoding and decoding process according to this disclosure. In Figure 7, processes 301-303 are performed by the server, and processes 305-308 are performed by the client.

[0057] <4.1 Server-side processing for color> Assume there is a scene with objects. Using the view from the camera, the server extracts the color data, alpha data, and depth data of the scene (301). Here, alpha data (or alpha value) is a numerical value that indicates additional information assigned to each pixel, separate from the color information. Alpha data is often used, in particular, to represent transparency. The collection of alpha data is also called an alpha channel.

[0058] Next, the server adds the color data, alpha data, and depth data to the stream and compresses them (302-1, 302-2, 302-3). The server then sends the compressed camera data as part of a container stream to the client via network 120 (303).

[0059] <4.2 Client-side processing of color> The client receives a container stream containing a compressed camera data stream via network 120 (305). The client decompresses the received camera data, as well as preparing a set of frames (306). Next, the client processes the color data, alpha data, and depth data of the video stream from the decompressed camera data, respectively (306-1, 306-2, 306-3). Here, these raw feature data are prepared and queued to be applied to the reconstructed 3D scene. The color data is used to wrap the mesh of the reconstructed 3D scene with textures.

[0060] Additionally, further details, including depth and alpha data, are used. Next, the client synchronizes the color, alpha, and depth data of the video stream (309). The client manages the color data queue by saving the synchronized color, alpha, and depth data to a queue (307). Next, the client projects the color / texture information onto the geometry (308).

[0061] Figure 8 shows the data synchronization between geometry packets, color packets, metadata, and commands according to this disclosure.

[0062] For the data to be available to the client, it is necessary to manage the data in a way that provides the correct content of the data in the stream while the client plays back the received 3D image. Data packets transmitted over a network are not always reliable, and packet delays and / or changes in packet order may occur. Therefore, the client system needs to consider how to manage data synchronization while receiving the data container stream at the client. The basic scheme for synchronizing geometry, color, metadata, and commands in this disclosure is as follows. This scheme may be standard for data formats created for network applications and streams.

[0063] Referring to Figure 8, the 3D stream 410 transmitted from the server side includes geometry packets, color packets, metadata, and commands. The geometry packets, color packets, metadata, and commands included in the 3D stream are synchronized with each other as shown in the frame sequence 410 when the 3D stream is created on the server.

[0064] In this frame sequence 410, time flows from left to right. However, when the frame sequence 410 sent from the server is received by the client, synchronization may be lost or random delays may occur during transmission over the network, as shown in the 3D stream 401 received by the client. In other words, within the 3D stream 401 received by the client, it can be seen that the order or position of geometry packets, color packets, metadata, and commands differs in some places from the 3D stream 410 created by the client.

[0065] The 3D stream 401 received by the client is processed by the packet queue manager 402 to restore its original synchronization, and a frame sequence 403 is generated. In frame sequence 403, where synchronization has been restored by the packet queue manager 402 and the differing delays have been resolved, geometry packets 1, 2, and 3 are in the correct order and arrangement, color packets 1, 2, 3, 4, and 5 are in the correct order and arrangement, metadata packets 1, 2, 3, 4, and 5 are in the correct order and arrangement, and commands 1 and 2 are in the correct order and arrangement. In other words, the frame sequence 403 after alignment within the client has the same order as the frame sequence 410 created within the server.

[0066] Next, the scene is reconstructed using the data for the synchronized current frame (404). Then, the reconstructed frame is rendered (405), and the client displays the scene on the display device (406).

[0067] Figure 9 shows an example of sequence update flow 500. In Figure 9, time progresses from left to right. Referring to Figure 9, first the geometry is updated (501). Here, the color / texture is updated in sync with it (i.e., the horizontal position matches) (505). Next, the color / texture is updated (506), but the geometry is not updated (for example, the color has changed but there is no movement). Next, the geometry is updated (502), and the color / texture is updated in sync with it (507). Next, the color / texture is updated (508), but the geometry is not updated. Next, the geometry is updated (503), and the color / texture is updated in sync with it (509).

[0068] As can be seen from Figure 9, the geometry does not need to be updated every time the color / texture is updated, nor does the color / texture need to be updated every time the geometry is updated. Geometry updates and color / texture updates may be synchronized. Furthermore, color / texture updates do not necessarily have to update both the color and texture; either the color or the texture alone may be updated. In this figure, the frequency of geometry updates is shown as once for every two color / texture updates, but this is just an example, and other frequencies are possible.

[0069] Figure 10 is a schematic diagram showing an example of the hardware configuration of a client according to this disclosure. The client 150 may be a terminal such as a smartphone or mobile phone. The client 150 typically includes a CPU / GPU 601, a display unit 602, an input / output unit 603, a memory 604, a network interface 605, and a storage device 606, and these components are connected to each other via a bus 607 so that they can communicate with one another.

[0070] The CPU / GPU 601 may be a single CPU or a single GPU, or it may consist of one or more components in which the CPU and GPU work together. The display unit 602 is a device for displaying images in normal color, and displays and presents the 3D image according to this disclosure to the user. Referring to Figure 5, as described above, the client may be a combination of a client terminal and smart glasses, in which case the smart glasses have the function of the display unit 602.

[0071] The input / output unit 603 is a device for interacting with external parties such as a user, and may be connected to a keyboard, speaker, buttons, or touch panel located inside or outside the client 150. The memory 604 is volatile memory for storing software and data necessary for the operation of the CPU / GPU 601. The network interface 605 has the function of allowing the client 150 to connect to an external network and communicate. The storage device 606 is non-volatile memory for storing software, firmware, data, etc., necessary for the client 150.

[0072] Figure 11 is a schematic diagram showing an example of the server hardware configuration according to this disclosure. Server 100 typically has a more powerful CPU, faster communication speed, and larger storage capacity than a client. Server 100 typically includes a CPU / GPU 701, an input / output unit 703, memory 704, a network interface 705, and storage device 706, and these components are connected to each other via a bus 707 so that they can communicate with one another.

[0073] The CPU / GPU 701 may consist of a single CPU or a single GPU, or it may consist of one or more components that allow the CPU and GPU to work together. The client device shown in Figure 10 had a display unit 602, but in the case of a server, a display unit is not required. The input / output unit 703 is a device for interacting with the user, etc., and may be connected to a keyboard, speaker, buttons, or touch panel. The memory 704 is a volatile memory for storing software and data necessary for the operation of the CPU / GPU 701. The network interface 705 has the function of allowing the client to connect to an external network and communicate. The storage device 706 is a non-volatile storage device for storing software, firmware, data, etc., required by the client.

[0074] Figure 12 is a schematic diagram showing an example of the configuration of the information processing system according to this disclosure. Server 100, client 150-1, and client 150-2 are connected to each other via a network 120 so that they can communicate with one another.

[0075] Server 100 is, for example, a computer device such as a server, and operates in response to image display requests from clients 150-1 and 150-2, generating and transmitting information related to the images to be displayed by clients 150-1 and 150-2. In this example, two clients are described, but any number of clients, one or more, is acceptable.

[0076] Network 120 may be a wired or wireless LAN (Local Area Network), and clients 150-1 and 150-2 may be smartphones, mobile phones, slate PCs, or game terminals.

[0077] Figure 13 is a schematic diagram showing the server-side processing flow according to this disclosure. Color information (1303) is extracted from object 1302 in server-side scene 1301 using an RGB camera (1310). Alpha information (1304) is extracted from object 1302 in server-side scene 1301 using an RGB camera (1320). Point cloud information (1305) is extracted from object 1320 in server-side scene 1301 using a depth camera (1330). Next, the point cloud information (1305) is simplified to obtain geometry information (1306) (1331).

[0078] Next, the obtained color information (1303), alpha information (1304), and geometry information (1306) are processed into a stream data format (1307) and sent to the client via the network as a container stream for the 3D stream (1340).

[0079] Figure 14 is a schematic diagram showing the client-side processing flow according to this disclosure. Figure 14 relates to the decal method according to this disclosure. Here, decal refers to the process of applying textures or materials to an object. Here, texture refers to data used to represent the texture, pattern, and unevenness of a 3D (three-dimensional) CG model. Material refers to the material of an object, and in 3DCG, for example, it refers to the optical properties and material feel of an object.

[0080] The reason why the decaling method disclosed in this disclosure is less processing-intensive than conventional UV mapping is explained below. Currently, there are several ways to set colors on a mesh. Here, UV refers to the coordinate system used to specify the position, direction, size, etc., when mapping textures to a 3DCG model. It is a two-dimensional Cartesian coordinate system, with the horizontal axis being U and the vertical axis being V. Texture mapping using the UV coordinate system is called UV mapping.

[0081] <5.1 Method for setting a color for each vertex (Traditional Method 1)> For all triangles in the target cloud, the color values ​​are stored at the vertices. However, if the vertex density is low, it results in low-resolution texturing, degrading the user experience. Conversely, if the vertex density is high, it's equivalent to sending color to every pixel on the screen, increasing the amount of data transferred from the server to the client. On the other hand, this can be used as an additional / basic coloring step.

[0082] <5.2 How to set the correct texture to the mesh UV (Traditional Method 2)> Setting the correct texture with UV mapping requires generating a texture of a group of triangles. Then, the UV map of the current mesh needs to be generated and added to the stream. The model's original texture does not contain information such as scene lightning, and a large amount of texture is required for high-quality, detailed 3D models, making it practically unusable. Another reason this method is not adopted is that the original texture works with UVs created on 3D models rendered on the server. Generally, a group of triangles is used to project a coloring texture from different views, and the received UV texture is saved and sent. Furthermore, the amount of data sent and received between the server and client increases because the mesh geometry and topology need to be updated as frequently as the UV texture.

[0083] <5.3 Method for projecting textures onto a mesh (decal) (Method according to this disclosure)> The color / texture from a specific location in the stream, along with metadata about that location, is sent from the server to the client. The client projects this texture onto the mesh from the specified location. In this case, a UV map is not required. With this method, UV generation is not loaded on the streaming side (i.e., the client side). This decal-based approach allows for optimization of the data flow (for example, geometry and color updates can be performed sequentially at different frequencies).

[0084] The client-side processing shown in Figure 14 is essentially the reverse of the server-side processing shown in Figure 13. First, the client receives a container stream, which is a 3D stream sent by the server over the network. Next, it decodes the data from the received 3D stream and restores color information, alpha information, and geometry information in order to reconstruct the object (1410).

[0085] Next, the color information 1431 and alpha information 1432 are combined, resulting in the generation of texture data. Then, this texture data is applied to the geometry data 1433 (1420). In this way, the object that was on the server is reconstructed on the client (1440). If there are multiple objects in the server-side scene, this process is applied to each object.

[0086] Figure 15 shows the camera arrangement used in this disclosure. One or more depth cameras are used to obtain the geometry information of objects in the target scene 1510. The depth cameras acquire depth maps every frame, and these depth maps are then processed into a point cloud. The point cloud is then divided into a predetermined triangular mesh for simplification. The level of detail (graininess) of the triangular mesh can be controlled by changing the resolution of the depth cameras. For example, a typical setup assumed uses six depth cameras with a resolution of 256 x 256 (1521-1526). However, the number of depth cameras required and the resolution of each camera can be further optimized and reduced, and the performance, i.e., image quality and the amount of data transmitted, changes depending on the number of depth cameras and their resolution.

[0087] Figure 15 shows an example configuration consisting of six depth cameras (1521-1526) and one standard camera (1530). The standard camera (i.e., an RGB camera) (1530) is used to acquire the color and alpha information of objects in the target scene 1510.

[0088] Figure 16 shows the pixel configuration in the ARGB system used in this disclosure. The ARGB system adds alpha information (A), which represents transparency, to the conventional RGB (red, green, blue) color information. In the example shown in Figure 16, each of alpha, blue, green, and red is represented by 8 bits (i.e., 256 levels), resulting in a 32-bit configuration for ARGB as a whole. In Figure 16, 1601 shows the number of bits for each color or alpha, 1602 shows each color or alpha, and 1603 shows the 32-bit configuration as a whole. In this example, an ARGB system with 8 bits for each color and alpha, and a total of 32 bits, is described, but these number of bits can be changed according to the desired image quality and amount of data to be transmitted.

[0089] Alpha information can be used as a mask / secondary layer for a color image. Due to the limitations of current hardware encoders, encoding a video stream with color information that also contains alpha information is time-consuming. Furthermore, software encoders for color and alpha on a video stream cannot encode in real time, resulting in delays and failing to achieve the purpose of this disclosure; therefore, they cannot currently serve as an alternative to this disclosure.

[0090] <6.1 Advantages of Reconstructing 3D Stream Scene Geometry as per the Disclosure> The advantages of using the method disclosed for reconstructing the geometry of 3D stream scenes are as follows: The method disclosed reconstructs each scene on the client side by using a "cloud of triangles". The key point of this innovative idea is that the client side is ready to use a large number of triangles. In this case, the number of triangles in the cloud of triangles may be in the hundreds of thousands.

[0091] The client is ready to place each triangle in the appropriate location to create the shape of the 3D scene as soon as it retrieves information from the stream. The advantage of this method is that it reduces the power and time required for processing, as it transfers less data from the server to the client than traditional methods. Instead of the traditional method of generating a mesh frame by frame, it modifies the position of existing geometry. However, by modifying the position of existing geometry, the position of a group of triangles, once generated in the scene, can be changed. Thus, the geometry data provides the coordinates of each triangle, and this change in the object's position is dynamic.

[0092] <6.2 Benefits of 3D Streaming as per this Disclosure> The advantage of the 3D streaming described in this disclosure is that it has less network latency across all six degrees of freedom (DoF). One of the advantages of the 3D streaming format is that the client also has a 3D scene. When navigating in mixed reality (MR) or looking around within an image, the key is how the 3D content connects to the real world and how "it feels like it's actually in place." In other words, when a user is moving around some displayed object and doesn't perceive a delay in the device's position update, the human brain is tricked into believing that the object is actually in that location.

[0093] Currently, client-side devices aim for 70-90 FPS (frames per second) to update 3D content on the display to make the user perceive it as "real." Today, it is impossible to provide a full cycle of frame updates on a remote server with a latency of less than 12ms. In fact, AR device sensors provide information at over 1,000 FPS. And since synchronizing 3D content on the client side is already possible with modern devices, the method disclosed here allows for client-side synchronization of 3D content. Therefore, after reconstructing the 3D scene, it is the client's job to process the positional information of the augmented content, and any reasonable network issues (e.g., transmission delay) that do not affect reality can be resolved.

[0094] <Summary of this disclosure> Several embodiments are provided below to summarize this disclosure.

[0095] A method for transmitting a 3D object from a server to a client, comprising: extracting color information, alpha information, and geometry information from the 3D object on the server; simplifying the geometry information; and encoding and transmitting a stream containing the color information, alpha information, and simplified geometry information from the server to the client.

[0096] The method according to this disclosure, wherein simplifying the geometric information involves converting the point cloud extracted from the 3D object into information about the vertices of a triangle.

[0097] The method according to this disclosure further includes at least one of metadata, sound data, and commands.

[0098] The method according to this disclosure, wherein the server receives a command from the client to rewrite a 3D object on the server.

[0099] The method according to the present disclosure, wherein when the server receives a command from the client to rewrite a 3D object, the server rewrites the 3D object on the server, extracts the color information, the alpha information and the geometry information from the rewritten 3D object, simplifies the geometry information, and encodes and transmits a stream containing the color information, the alpha information and the simplified geometry information from the server to the client.

[0100] The method according to this disclosure, wherein the color information and alpha information are acquired by an RGB camera, and the geometry information is acquired by a depth camera.

[0101] A method for reproducing a 3D object on a server on a client, comprising: receiving an encoded stream from the server containing color information, alpha information, and geometry information of the 3D object; decoding the stream and extracting the color information, alpha information, and geometry information therefrom; reproducing the shape of the 3D object based on the geometry information; and reconstructing the 3D object by projecting information obtained by combining the color information and the alpha information onto the reproduced shape of the 3D object.

[0102] A method according to the present disclosure for displaying the reconstructed 3D object on a display device.

[0103] The method according to this disclosure, wherein the display device is a smart glass.

[0104] A server comprising one or more processors and memory, wherein the one or more processors extract alpha information and geometry information from 3D objects on the server by executing instructions stored in the memory, simplify the geometry information, and encode and transmit a stream containing the alpha information and the simplified geometry information from the server to a client, according to the present disclosure.

[0105] A client comprising one or more processors and memory, wherein the one or more processors receive an encoded stream from a server containing color information, alpha information, and geometry information of a 3D object by executing instructions stored in the memory, decode the stream, extract the color information, alpha information, and geometry information therefrom, reproduce the shape of the 3D object based on the geometry information, and reconstruct the 3D object by projecting information obtained by combining the color information and the alpha information onto the reproduced shape of the 3D object.

[0106] A program including instructions for a processor to perform the method described herein.

[0107] This disclosure can be implemented using software, hardware, or software integrated with hardware. [Industrial applicability]

[0108] This disclosure is applicable to software, programs, systems, devices, client-server systems, terminals, and the like. [Explanation of Symbols]

[0109] 100 servers 101 Received Data Processing Unit 102 3D Scene Data Generation Unit 103 Extraction Unit 104 3D Stream Conversion / Encoding Units 105 3D Streams 106 Network Packet Configuration Unit 107 Network Packet Transmission Unit 108 Network Packet Receiver Unit 120 Wired or wireless network 150 clients 150-1 Client 150-2 Client 151 Network Packet Transmission Unit 152 Network Packet Receiver Unit 153 Application Data Output Unit 154 3D stream decoding units 155 3D Scene Reconstruction Unit 156 Display Unit 602 Display section 603 Input / output section 604 memory 605 Network Interface 606 Storage device 607 Bus 703 Input / output section 704 memory 705 Network Interface 706 Storage device 707 Bus 1201 screen 1202 people 1203 Container Stream 1210 Smart Glasses 1211-1 people 1211-2 people 1212-1 Cursor 1212-2 Cursor 1213 command 1214 people 1221 Terminal device 1521-1526 Depth Camera 1530 RGB Camera

Claims

1. A method for sending a virtual 3D object from a server to a client, The server receives commands from the client regarding a virtual 3D object on the server, and modifies the virtual 3D object according to the commands from the client. Extracting color information, alpha information, and geometry information from the modified virtual 3D object on the server, To simplify the aforementioned geometric information, The server encodes and transmits a stream containing the color information, alpha information, and simplified geometry information from the server to the client. Methods that include...

2. The method according to claim 1, wherein simplifying the geometric information involves converting the point cloud extracted from the virtual 3D object into information of the vertices of a triangle, and compressing the converted information of the vertices of a triangle before sending it to the client.

3. The method according to claim 1 or 2, wherein the stream further includes at least one of metadata, sound data, and commands.

4. The method according to any one of claims 1 to 3, wherein the server receives a command from the client to rewrite a virtual 3D object on the server.

5. The method according to any one of claims 1 to 4, wherein when the server receives a command from the client to rewrite a virtual 3D object, the server rewrites the virtual 3D object on the server, extracts the color information, the alpha information and the geometry information from the rewritten virtual 3D object, simplifies the geometry information, and encodes and transmits a stream containing the color information, the alpha information and the simplified geometry information from the server to the client.

6. The method according to any one of claims 1 to 5, wherein the color information and the alpha information are acquired by an RGBA camera, and the geometry information is acquired by a depth camera.

7. A server comprising one or more processors and memory, The one or more processors execute instructions stored in the memory, The server receives commands from a client regarding a virtual 3D object on the server, and modifies the virtual 3D object according to the commands from the client. Color information, alpha information, and geometry information are extracted from the modified virtual 3D object on the server. The aforementioned geometric information is simplified, A server that encodes and transmits a stream containing the color information, the alpha information, and the simplified geometry information from the server to the client.

8. A program comprising instructions for causing a processor to perform any of the methods described in claims 1 to 6.

Citation Information

Patent Citations

  • Image photographing method

    JP1997200599A

  • Image processing apparatus which processes graphics by dividing space, and image processing method

    JP2014219739A

  • Video synthesizing apparatus, program and method for synthesizing viewpoint video by projecting object information onto plural surfaces

    JP2019046077A

  • Image forming device

    JP2020052369A

  • Remote shading-based 3D streaming apparatus and method

    US20100134494A1