Content server, client terminal, image display system, display data transmission method, and display image generation method

By using a content server to generate learning images and transmit 3D scene information through machine learning, the system addresses the challenges of maintaining high-quality user experiences in three-dimensional display worlds, achieving reduced latency and robustness against packet loss.

WO2025120714A1PCT designated stage expired Publication Date: 2025-06-12SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/043347
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing content distribution systems face challenges in maintaining high-quality user experiences due to potential delays and packet loss during real-time image processing and transmission, especially in three-dimensional display worlds.

Method used

A content server generates learning images from multiple viewpoints of a three-dimensional display world and uses machine learning to acquire and transmit 3D scene information to client terminals, allowing for the generation of high-quality display images with reduced latency and robustness against packet loss.

Benefits of technology

The proposed solution effectively reduces the impact of distribution delays on user experience by enabling high-quality, low-latency image rendering in three-dimensional display worlds, even in the presence of packet loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023043347_12062025_PF_FP_ABST
    Figure JP2023043347_12062025_PF_FP_ABST
Patent Text Reader

Abstract

This content server 20 acquires neural networks 230a, 23b,... representing information about a plurality of types of 3D scenes by generating a learning image representing each scene in a display world and performing machine learning. The content server 20 divides the neural networks 230a, 23b,... to generate a plurality of neural networks 232a, 232b, etc., randomly switches the order of packets, and transmits the packets to a client terminal 10. The client terminal 10 draws a display image 238 corresponding to the most recent display viewpoint by, for example, returning the divided neural networks 232a, 232b, etc. to the original neural network 230a.
Need to check novelty before this filing date? Find Prior Art

Description

Content server, client terminal, image display system, display data transmission method, and display image generation method

[0001] The present invention relates to a content server, a client terminal, an image display system, a display data transmission method, and a display image generation method for displaying an image of a three-dimensional display world.

[0002] With the recent expansion of communication networks and advances in image processing technology, it has become possible to enjoy a wide variety of electronic content regardless of the viewing environment. For example, in the field of electronic games, a system has become widespread in which a server collects operation information entered into each client terminal and distributes game images that reflect that information as needed, allowing multiple players to participate in the same game regardless of location.

[0003] This is not limited to electronic games, but any electronic content that generates video images in real time in response to user operations and distributes them from a server can utilize the server's abundant processing environment, making it easier to display high-quality images while minimizing the impact of the client terminal's processing performance. However, the constant transmission of operation information from the client terminal and the processing of data transmission from the server that receives it can cause the display to be unable to keep up with viewpoint operations, or images to be lost due to packet loss, which can degrade the quality of the user experience.

[0004] The present invention has been made in consideration of these problems, and its purpose is to provide a technology that reduces the impact of distribution on the quality of the user experience when processing images of content that is distributed from a server.

[0005] To solve the above problems, one aspect of the present invention relates to a content server, which includes: a training image generation unit that generates training images representing scenes in a three-dimensional display world, the situation of which changes in response to user operations, as viewed from multiple viewpoints; a type-specific 3D scene information acquisition unit that acquires multiple types of 3D scene information representing three-dimensional information about each scene through machine learning using the training images as training data; and a 3D scene information transmission unit that transmits data on the multiple types of 3D scene information to a client terminal that renders a display image using the 3D scene information in the order in which the machine learning for each scene is completed.

[0006] Another aspect of the present invention relates to a client terminal, comprising: an input information acquisition unit that acquires information on user operations and information on a display viewpoint for a three-dimensional display world; a 3D scene information acquisition unit that acquires, from a server, data on multiple types of 3D scene information representing three-dimensional information acquired by machine learning for each scene in the display world whose situation changes in response to the user operations; and an image generation unit that renders at least a portion of a frame of a display image using the most recently acquired 3D scene information based on the latest display viewpoint.

[0007] Yet another aspect of the present invention relates to an image display system including: a client terminal that displays images of a three-dimensional display world whose situation changes in response to user operations; and a content server that transmits data used to generate the display images, wherein the content server includes a learning image generation unit that generates images representing each scene of the display world as viewed from multiple viewpoints as learning images, a type-specific 3D scene information acquisition unit that acquires multiple types of 3D scene information representing three-dimensional information of each scene through machine learning using the learning images as training data, and a 3D scene information transmission unit that transmits the multiple types of 3D scene information data to the client terminal in the order in which the machine learning for each scene is completed, and the client terminal includes an input information acquisition unit that acquires information on user operations and information on a display viewpoint for the display world, a 3D scene information acquisition unit that acquires the multiple types of 3D scene information data from the content server, and an image generation unit that draws at least a portion of a frame of a display image using the most recently acquired 3D scene information based on the latest display viewpoint.

[0008] Yet another aspect of the present invention relates to a display data transmission method, the display data transmission method including the steps of: generating, as learning images, images representing how each scene in a three-dimensional display world, whose situation changes in response to a user operation, is viewed from a plurality of viewpoints; acquiring, by machine learning using the learning images as training data, a plurality of types of 3D scene information representing three-dimensional information of each scene; and transmitting, to a client terminal that renders a display image using the 3D scene information, data on the plurality of types of 3D scene information in the order in which the machine learning for each scene is completed.

[0009] Yet another aspect of the present invention relates to a display image generation method, the display image generation method including the steps of: acquiring information on a user operation and information on a display viewpoint for a three-dimensional display world; acquiring, from a server, data of multiple types of 3D scene information representing three-dimensional information acquired by machine learning for each scene in the display world whose situation changes in response to the user operation; and rendering at least a part of a frame of a display image using the most recently acquired 3D scene information based on the latest display viewpoint.

[0010] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention.

[0011] According to the present invention, in image processing of content that involves distribution from a server, the impact of distribution on the quality of the user experience can be reduced.

[0012] FIG. 1 is a diagram showing an example of the configuration of an image display system to which the present embodiment can be applied. FIG. 2 is a diagram showing the internal circuit configuration of a client terminal of the present embodiment. FIG. 3 is a diagram showing the basic flow of image processing of the present embodiment in comparison with the prior art. FIG. 4 is a diagram showing the configuration of functional blocks of a client terminal and a content server of the present embodiment. FIG. 5 is a diagram showing a schematic diagram of a procedure by which a content server acquires 3D scene information in the present embodiment. FIG. 6 is a diagram for explaining a mode in which 3D scene information of different ranges is acquired in the present embodiment. FIG. 7 is a diagram showing a schematic diagram of data transition in the transmission and reception of 3D scene information in the present embodiment. FIG. 8 is a diagram for explaining the temporal relationship between machine learning in the content server and image drawing in the client terminal in the present embodiment. FIG. 9 is a diagram for explaining image correction processing by an image generation unit of a client terminal in the present embodiment.

[0013] 1 shows an example of the configuration of an image display system to which this embodiment can be applied. The image display system 1 includes client terminals 10a, 10b, and 10c that display images in response to user operations, and a content server 20 that provides data used for display. Input devices 14a, 14b, and 14c for user operations and display devices 16a, 16b, and 16c for displaying images are connected to the client terminals 10a, 10b, and 10c, respectively. Communication can be established between the client terminals 10a, 10b, and 10c and the content server 20 via a network 8 such as a WAN (World Area Network) or a LAN (Local Area Network).

[0014] The client terminals 10a, 10b, and 10c may be connected to the display devices 16a, 16b, and 16c and the input devices 14a, 14b, and 14c either wired or wirelessly. Alternatively, two or more of these devices may be integrated. For example, in the figure, the client terminal 10b is connected to a head-mounted display, which is the display device 16b. The head-mounted display can change the field of view of the displayed image by the movement of the user wearing it on their head, so it also functions as the input device 14b.

[0015] Furthermore, the client terminal 10c is a portable terminal that is integrally configured with a display device 16c and an input device 14c, which is a touchpad that covers the display device 16c's screen. As such, the external shape and connection form of the illustrated devices are not limited. The number of client terminals 10a, 10b, 10c and content servers 20 that can be connected to the network 8 is also not limited. Hereinafter, the client terminals 10a, 10b, 10c will be collectively referred to as client terminals 10, the input devices 14a, 14b, 14c as input device 14, and the display devices 16a, 16b, 16c as display device 16.

[0016] The input device 14 may be any one or a combination of general input devices such as a controller, keyboard, mouse, touchpad, or joystick, or various sensors such as a motion sensor or camera provided in a head-mounted display, and supplies the content of user operations to the client terminal 10. The display device 16 may be a general display such as a liquid crystal display, plasma display, organic EL display, wearable display, or projector, and displays images output from the client terminal 10.

[0017] The content server 20 provides data of content accompanied by image display to the client terminal 10. The type of content is not particularly limited, and may be any of electronic games, decorative images, web pages, video chat using avatars, etc. In this embodiment, the content server 20 sequentially acquires information on user operations on the input device 14 from the client terminal 10, reflects the information in the world to be displayed, and transmits the necessary data so that an image representing the information can be displayed on the client terminal 10 side.

[0018] 2 shows the internal circuit configuration of the client terminal 10. The client terminal 10 includes a CPU (Central Processing Unit) 122, a GPU (Graphics Processing Unit) 124, and a main memory 126. These components are interconnected via a bus 130. An input / output interface 128 is also connected to the bus 130. Connected to the input / output interface 128 are a communication unit 132 including a peripheral device interface such as a USB or a network interface for a wired or wireless LAN, a storage unit 134 such as a hard disk drive or nonvolatile memory, an output unit 136 that outputs data to the display device 16, an input unit 138 that inputs data from the input device 14, and a recording medium drive unit 140 that drives a removable recording medium such as a magnetic disk, optical disk, or semiconductor memory.

[0019] The CPU 122 executes an operating system stored in the storage unit 134 to control the entire client terminal 10. The CPU 122 also executes various programs read from a removable recording medium and loaded into the main memory 126, or downloaded via the communication unit 132. The GPU 124 performs drawing processing in accordance with drawing commands from the CPU 122 and stores the display image in a frame buffer (not shown). The GPU 124 then converts the display image stored in the frame buffer into a video signal and outputs it to the output unit 136. The main memory 126 is composed of RAM (Random Access Memory) and stores programs and data required for processing. The content server 20 may also have a similar internal circuit configuration.

[0020] FIG. 3 shows the basic flow of image processing in this embodiment, compared with the prior art. In this embodiment, a three-dimensional world containing various objects is the primary display target. The state of this world changes depending on program specifications and user operations. In the general processing shown in (a), the content server continually acquires information on the content of user operations, the position of the viewpoint relative to the displayed world, and the direction of the line of sight. Hereinafter, the entire three-dimensional space to be displayed is referred to as the "display world," and the state of the displayed world within or near the display field of view is referred to as the "scene." The position of the viewpoint relative to the scene and the direction of the line of sight may also be collectively referred to simply as the "viewpoint." The viewpoint may be manually controlled by the user via the input device 14, or may be derived from the movement of the user's head using a motion sensor or the like provided in the head-mounted display.

[0021] The content server renders the display image 200 in a field of view corresponding to the viewpoint information while changing the scene in response to user operations. The content server generates the display image 200 using well-known computer graphics rendering techniques such as ray tracing and rasterization. The content server transmits the generated display image 200 to the client terminal, which then displays it on a display device. By repeating the illustrated process at a predetermined frame rate, a moving image representing changes in the scene in response to user operations, etc., is displayed on the client terminal.

[0022] In this embodiment shown in (b), the content server 20 also acquires information on the content of user operations, the position of the viewpoint relative to the displayed world, and the direction of the line of sight as needed, and similarly renders images while changing the scene accordingly. Meanwhile, in this embodiment, the content server 20 uses the images as training images 202 and as training data for machine learning. The content server 20 collects the training images 202 and performs machine learning to generate 3D scene information 204 representing three-dimensional information about the scene.

[0023] Neural Radiance Fields (NeRF) is a method for acquiring information about three-dimensional space through machine learning (see, for example, Ben Mildenhall and five others, "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis," Communications of the ACM, January 2022, Vol. 65, No. 1, pp. 99-106). When NeRF is introduced in this embodiment, first, each piece of viewpoint information determined when generating the training images 202, i.e., the position of the virtual viewpoint and the direction of the line of sight, is input, and the corresponding training images 202 are used as training data to obtain data representing the three-dimensional information of the scene through regression using a multilayer perceptron (MLP).

[0024] This data is a neural network that takes five-dimensional parameters consisting of position coordinates (x, y, z) and a direction vector d (θ, φ) in three-dimensional space as input, and outputs volume density σ and color information c (RGB) of the three primary colors. In this embodiment, the data of this neural network is called "3D scene information." Note that various improvement methods for NeRF have been proposed, and the specific method is not particularly limited in this embodiment. Furthermore, it is not intended that the machine learning method be limited to NeRF.

[0025] In this embodiment, the content represented by the learning image 202, and therefore the 3D scene information 204, may change from moment to moment. The diagram illustrates the generation of 3D scene information 204 for a scene at a certain time, or at a very short time that can be considered as a time. Hereinafter, a scene at a certain time, or at a very short time that can be considered as a time, may be referred to as "one scene." Note that if there is no movement in the displayed world, it can be treated as "one scene" regardless of time. Furthermore, when 3D scene information 204 has been obtained for a previous scene, the content server 20 updates the 3D scene information 204 to correspond to the next scene. Hereinafter, the generation and update of 3D scene information may be collectively referred to as "acquiring" 3D scene information.

[0026] To obtain accurate 3D scene information 204, it is desirable for the content server 20 to collect learning images 202 of one scene from as many viewpoints as possible. Therefore, the content server 20 may collect learning images 202 in the following manner, for example: (1) Generate viewpoints suitable for learning around the viewpoint that defines the field of view of the image that is actually displayed, and generate corresponding images. (2) Reuse images displayed on the terminals of multiple users viewing the same scene.

[0027] Hereinafter, the viewpoint that defines the actual display field of view on the client terminal 10 will be referred to as the "display viewpoint," and will be distinguished from the "learning viewpoint" that is set when generating a learning image. The content server 20 may implement either (1) or (2), or both. For example, a viewpoint that is missing due to (2) may be supplemented by (1). The content server 20 transmits 3D scene information 204, which is the learning result, to the client terminal 10. The client terminal 10 generates a display image 206 using the transmitted 3D scene information 204.

[0028] By using the 3D scene information 204, the client terminal 10 can display a high-quality representation of the scene as seen from any viewpoint, with a relatively low load. When NeRF is applied, the client terminal 10 generates a ray r that passes through each pixel on the view screen from the display viewpoint, and calculates the pixel value C(r) of the display image by volume rendering that integrates color along that direction, as follows:

[0029]

[0030] where t n , t f are the proximal and distal ends of the ray r, respectively, and T(t) is the cumulative transmittance in the direction of the ray, which can be expressed as follows:

[0031]

[0032] The content server 20 repeats the illustrated process in response to movement in the scene, updating the 3D scene information 204 at a predetermined rate and sequentially transmitting the updated 3D scene information to the client terminal 10. The client terminal 10 generates display images 206 while updating the 3D scene information used, thereby enabling moving images to be displayed from any viewpoint that have the same changes as the learning images 202 generated by the content server 20. For example, the client terminal 10 can display an image with little delay in response to changes in viewpoint by drawing the display image 206, which shows how the scene looks from the viewpoint immediately before display, using the latest 3D scene information 204.

[0033] 4 shows the functional block configuration of the client terminal 10 and the content server 20 in this embodiment. The functions of the components in the illustrated functional blocks may be implemented in circuits or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, configured or programmed to implement the functions described herein. A processor is considered to be a circuit or processing circuitry including transistors and other circuits. A processor may also be a programmed processor that executes a program stored in a memory.

[0034] In this specification, a circuit, unit, or means is hardware that is programmed to realize or performs a described function. The hardware may be any hardware disclosed in this specification or any hardware known to be programmed to realize or perform the described function. If the hardware is a processor, which is considered to be a type of circuit, the circuit, means, or unit is a combination of hardware and software used to configure the hardware and / or processor.

[0035] The client terminal 10 includes an input information acquisition unit 50 that acquires input information such as user operations, a 3D scene information acquisition unit 52 that acquires 3D scene information data from the content server 20, a 3D scene information storage unit 54 that stores the acquired 3D scene information data, an image generation unit 56 that generates a display image, and an output unit 58 that outputs the display image data. The input information acquisition unit 50 acquires the content of user operations from the input device 14 as needed. User operations include selecting and launching content, and inputting commands for content that is currently being played.

[0036] The input information acquisition unit 50 also acquires information about a display viewpoint for the displayed world from the input device 14 or the head-mounted display at any time or at predetermined time intervals. Techniques for detecting the position and posture of the head of a user wearing a head-mounted display and acquiring viewpoint information based on the detected information are well known, and these may also be applied to this embodiment. The input information acquisition unit 50 supplies the acquired information to the content server 20 and the image generation unit 56 as appropriate. The 3D scene information acquisition unit 52 sequentially acquires continuously updated 3D scene information data from the content server 20.

[0037] As will be described later, the 3D scene information acquisition unit 52 acquires multiple types of 3D scene information representing one scene from the content server 20. The 3D scene information storage unit 54 stores data of the multiple types of 3D scene information acquired by the 3D scene information acquisition unit 52. When the 3D scene information acquisition unit 52 acquires new 3D scene information, it updates the data of the same type of 3D scene information stored in the 3D scene information storage unit 54.

[0038] The image generation unit 56 draws a display image at a predetermined frame rate using the 3D scene information data most recently stored in the 3D scene information storage unit 54. Here, the image generation unit 56 acquires the latest display viewpoint from the input information acquisition unit 50 and draws an image in the corresponding field of view using a technique such as the volume rendering described above. If 3D scene information corresponding to scene transitions can be prepared using machine learning, even if the client terminal 10 generates a display image, it can draw a high-quality image with a lighter load than with normal processing such as ray tracing.

[0039] The image generation unit 56 changes the display image representing a single scene over time, space, or both by drawing it using multiple types of 3D scene information. For example, in a mode in which the content server 20 transmits 3D scene information data in descending order of information density, the image generation unit 56 sequentially updates the display image using the transmitted 3D scene information. This allows the image representing a single scene to be displayed with low latency, while maintaining visual image quality by gradually increasing the resolution. The output unit 58 outputs the image drawn by the image correction unit 92 to the display device 16 at a predetermined rate for display.

[0040] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, a learning viewpoint generation unit 72 that generates a learning viewpoint, a display world control unit 74 that controls the display world, a 3D model storage unit 76 that stores 3D models of objects, a learning image generation unit 78 that generates learning images, a type-specific 3D scene information acquisition unit 80 that acquires data on multiple types of 3D scene information, a 3D scene information storage unit 84 that stores the acquired data on the 3D scene information, and a 3D scene information transmission unit 86 that transmits the data on the 3D scene information to the client terminal 10.

[0041] The input information acquisition unit 70 acquires user operation content and viewpoint information from the client terminal 10 at any time or at predetermined time intervals. The training viewpoint generation unit 72 generates a plurality of training viewpoints for generating training images. The training viewpoint generation unit 72 generates training viewpoints according to a predetermined rule around the latest display viewpoint acquired by the input information acquisition unit 70. For example, the training viewpoint generation unit 72 arranges a predetermined number of training viewpoints evenly within a sphere of a predetermined radius centered on the latest display viewpoint. Here, the predetermined radius may be a value obtained by multiplying the maximum speed expected for viewpoint movement by the longest time until display using the data is performed, for example.

[0042] The learning viewpoint generation unit 72 is not limited to distributing the learning viewpoints evenly, but may distribute more learning viewpoints in a range in which the viewpoint is expected to move, depending on the situation in the displayed world, etc. Furthermore, the learning viewpoint generation unit 72 may uniformly set line of sight in a predetermined number of directions for the viewpoint at each position, or may set more line of sight in directions in which the line of sight is expected to move. Note that the learning viewpoint generation unit 72 may also use the latest display viewpoint itself as a learning viewpoint.

[0043] The display world control unit 74 controls the three-dimensional display world represented as content in accordance with the details of the user operations acquired by the input information acquisition unit 70. For example, if the content is an electronic game, the display world control unit 74 places necessary objects such as a user character in the virtual space that serves as the stage for the electronic game, and gives them movement in accordance with commands input by the user and program specifications.

[0044] The three-dimensional model storage unit 76 stores three-dimensional models of objects existing in the display world, and the display world control unit 74 uses these models to construct the display world by reading them out as appropriate. The training image generation unit 78 generates training images from images of the scene viewed from multiple training viewpoints generated by the training viewpoint generation unit 72. The training image generation unit 78 preferably generates training images using a technique capable of rendering high-quality images, such as ray tracing.

[0045] The type-specific 3D scene information acquisition unit 80 uses the learning images generated by the learning image generation unit 78 to generate 3D scene information through machine learning as described above, and updates the 3D scene information to respond to changes in the scene. Here, the type-specific 3D scene information acquisition unit 80 acquires multiple types of 3D scene information representing a single scene. For example, the type-specific 3D scene information acquisition unit 80 acquires multiple types of 3D scene information that differ in the spatial density of the represented information. Hereinafter, the spatial density of the information contained in the 3D scene information will be referred to as "information density." Information density may also be referred to as the resolution of the information contained in the 3D scene information, or the spatial frequency of the information.

[0046] For example, according to the NeRF document mentioned above, positional encoding is performed to convert vectors input to a neural network during training into vectors in a high-dimensional space that include high frequencies, thereby enabling the high-frequency components of the output vector to be represented more accurately. The function γ used for the conversion is expressed as follows:

[0047]

[0048] In the above formula, the larger the parameter L, the more detailed 3D scene information, including higher frequency components, can be acquired. Using this, the type-specific 3D scene information acquisition unit 80 sets multiple parameters L and performs machine learning individually to acquire multiple pieces of 3D scene information with different information densities in parallel. However, the means for controlling the information density of the 3D scene information is not limited to this. For example, instead of positional encoding, multiresolution hash encoding may be applied, which sets a grid with multiple resolutions and represents an input vector based on its positional relationship with the vertices (see, for example, Thomas Muller et al., "Instant Neural Graphics Primitives with a Multiresolution Hash Encoding," ACM Transactions on Graphics, July 2022, Vol. 41, No. 4, Article 102, pp. 1-15).

[0049] In this case, the type-specific 3D scene information acquisition unit 80 can acquire multiple pieces of 3D scene information with different information densities in parallel by setting multiple levels L corresponding to the number of grid resolutions and performing machine learning individually. In these cases, multiple values ​​to be set as the parameter L are prepared in the internal memory of the type-specific 3D scene information acquisition unit 80. The type-specific 3D scene information acquisition unit 80 then sequentially uses multiple learning images generated for one scene by the learning image generation unit 78 to acquire multiple pieces of 3D scene information with different information densities in parallel, and stores the pieces of 3D scene information in the 3D scene information storage unit 84 or updates previously stored 3D scene information of the same type.

[0050] In this case, even for the same scene, the learning speed varies depending on the information density, and the lower the information density, the faster learning is completed, and therefore the faster the 3D scene information is updated. When performing machine learning by setting grids, the type-specific 3D scene information acquisition unit 80 may set only one level L for learning, and then individually select and read data from grids of different levels (e.g., grids of levels 0 to 1, grids of levels 0 to 2, ..., grids of levels 0 to L-1, etc.) when transmitting data to the client terminal 10. In this case, too, the faster learning is completed for lower-level grids, i.e., grids with lower information density, and the same effects as those described below are achieved.

[0051] The types of 3D scene information acquired by the type-specific 3D scene information acquisition unit 80 are not limited to distinctions in information density. For example, the type-specific 3D scene information acquisition unit 80 may acquire multiple pieces of 3D scene information for different learning target areas or objects in the display world. In this case, the smaller the range of the learning target in the display world, the faster the learning is completed, and therefore the faster the 3D scene information is updated.

[0052] In this way, even if the multiple types of 3D scene information acquired by the type-specific 3D scene information acquisition unit 80 are for the same scene, the time required to complete the update may differ depending on the density of the information and the size of the range represented by the 3D scene information. For this reason, the type-specific 3D scene information acquisition unit 80 records the time of the reflected scene in association with each of the multiple types of 3D scene information stored in the 3D scene information storage unit 84. Note that the type-specific 3D scene information acquisition unit 80 may acquire multiple types of 3D scene information with different combinations of information density and range of the learning target.

[0053] The 3D scene information transmitting unit 86 transmits the latest 3D scene information data stored in the 3D scene information storage unit 84 to the client terminal 10. The 3D scene information transmitting unit 86 transmits the 3D scene information to the client terminal 10 in order, starting with the 3D scene information for which learning of one scene has been completed. For example, when data of 3D scene information with different information densities is to be transmitted, learning is completed in order starting with the 3D scene information with the lowest information density, as described above. Therefore, the 3D scene information transmitting unit 86 transmits the data of the 3D scene information with the lowest information density when learning of that information has been completed, and then transmits the data of the 3D scene information with the next lowest information density when learning of that information has been completed, and so on, gradually transmitting up to the 3D scene information with the highest information density.

[0054] The 3D scene information transmission unit 86 includes a division unit 88. The division unit 88 divides each of multiple types of 3D scene information representing one scene into multiple pieces of data. Each type of 3D scene information representing one scene is composed of a separate neural network. The division unit 88 divides each neural network into multiple neural networks. Here, division means removing some nodes that are different from each other while maintaining the structure of nodes that are associated by a hash table or the like. The nodes to be removed in the divided neural network are determined randomly.

[0055] A technique called dropout, which randomly deactivates some of the nodes in a neural network to alleviate overfitting, is known (see, for example, Nitish Srivastava et al., "Dropout: A Simple Way to Prevent Neural Networks from Overfitting," Journal of Machine Learning Research, June 2014, Vol. 15, pp. 1919-1958). As is clear from dropout, the accuracy of learning results can be maintained at a certain level even if some nodes are deactivated.

[0056] By utilizing this characteristic, the client terminal 10 can generate a display image with relatively high accuracy even when using 3D scene information with some nodes removed. Therefore, by having the content server 20 transmit multiple neural networks representing the same 3D scene information with some nodes removed, it is possible to suppress an increase in data size while increasing robustness against packet loss. If there is no packet loss, the image generation unit 56 of the client terminal 10 can restore the neural network before division and draw the display image. If there is packet loss, the image generation unit 56 draws the display image using the neural network that was acquired. As described above, it is possible to generate a display image in this case as well.

[0057] As described above, multiple types of 3D scene information representing a single scene may arrive from the content server 20 with a time lag. In this case, the image generation unit 56 of the client terminal 10 updates the display image using the newly acquired 3D scene information. Since the resolution of the display image corresponds to the information density of the 3D scene information, the resolution of the display image gradually increases by acquiring 3D scene information in order from the lowest information density. Here, 3D scene information with low information density requires a smaller sampling number in volume rendering, and therefore the rendering speed of the display image increases.

[0058] As a result, by first acquiring 3D scene information with low information density that requires a short learning time in the content server 20 and then drawing low-resolution images in a short time, the time from the start of learning each scene to its display can be significantly reduced. Furthermore, because high-resolution 3D scene information can ultimately be used to display high-definition images, low-latency display can be achieved while minimizing the impact on the visual appearance. As described above, the types of 3D scene information are not limited to distinctions in resolution. For example, the content server 20 may learn only a partial area that the user is gazing at in a short time and transmit it first, and then transmit 3D scene information for the entire scene later. In this case, using the same principle, it is possible to display the area of ​​interest with low latency and update it to track the surrounding images.

[0059] The 3D scene information transmitting unit 86 may transmit 3D scene information representing one scene to the client terminal 10 using different communication protocols depending on the type of 3D scene information. For example, the 3D scene information transmitting unit 86 may transmit 3D scene information for an area where image resolution should be prioritized using the highly reliable TCP / IP (Transmission Control Protocol / Internet Protocol), and transmit 3D scene information for an area where low latency in movement should be prioritized using the high transfer rate UDP (User Datagram Protocol). In this way, the correspondence between the type of 3D scene information and the appropriate communication protocol is set in advance in the internal memory of the 3D scene information transmitting unit 86, for example.

[0060] 5 is a schematic diagram showing the procedure by which the content server 20 acquires 3D scene information. First, the display world control unit 74 constructs a display world 210 in which, for example, an enemy character 212 exists. As described above, the display world 210 may be dynamic, but here, the display world 210 corresponding to one scene is shown. In response to this, the training viewpoint generation unit 72 generates multiple training viewpoints based on the most recent display viewpoint, and the training image generation unit 78 generates training images 214a, 214b, 214c, etc. corresponding to each training viewpoint. Note that there is no limit to the number of training images to be generated.

[0061] The type-specific 3D scene information acquisition unit 80 performs machine learning using learning images 214a, 214b, 214c, etc., and updates multiple types of 3D scene information. The substance of each piece of 3D scene information is a neural network 216a, 216b, 216c, etc. The neural networks 216a, 216b, 216c, etc. differ in at least one of the information density and the range they represent. When the information densities are to be varied, the type-specific 3D scene information acquisition unit 80 performs learning by, for example, respectively setting the above-mentioned parameter L. In this case, the lower the information density of 3D scene information, the faster learning is completed.

[0062] When the representation range is varied, the type-specific 3D scene information acquisition unit 80 performs learning by, for example, extracting corresponding regions from the training images 214a, 214b, and 214c. In this case, the narrower the representation range of the 3D scene information, the faster the learning is completed. When changing the information density and the representation range in combination, the time until learning is completed varies depending on the balance between them. Qualitatively, the higher the information density and the wider the representation range, the longer the learning completion time. Therefore, the balance between information density and range width is optimized to obtain an appropriate delay time. By repeating the illustrated process at a predetermined rate, the neural networks 216a, 216b, 216c, etc. are each updated in accordance with changes in the scene.

[0063] FIG. 6 is a diagram illustrating how 3D scene information of different ranges is acquired. Assume that a portion of the scene in the display world 210 shown in FIG. 5 is represented as the display image 222. The type-specific 3D scene information acquisition unit 80 separately acquires, for example, 3D scene information representing a range 226 within the display world 210 corresponding to an important region 224 on the display image 222, and 3D scene information representing the entire scene including the area outside that range. Here, the important region 224 may be, for example, a region within a predetermined range from the user's gaze point, a region within a predetermined range from the center of the display image, a region showing the battle situation or acquired items, or a region where major objects such as enemy characters or user characters are present. Area selection rules are set in advance depending on the content, etc.

[0064] Alternatively, the type-specific 3D scene information acquisition unit 80 may directly identify a major object itself or a predetermined-sized area including the object present in the display world 210, and acquire 3D scene information as a learning target. When determining the learning area based on the user's gaze point, a well-known gaze point detector is provided in the client terminal 10. The input information acquisition unit 70 of the content server 20 then acquires gaze point information from the client terminal 10 at a predetermined rate and individually determines the learning area.

[0065] The type-specific 3D scene information acquisition unit 80 may separately learn the range of the display world corresponding to the display image 222 currently being displayed on the client terminal 10 and a wider range including the outside of that range. In either case, the type-specific 3D scene information acquisition unit 80 performs machine learning by extracting corresponding partial areas from the multiple training images generated by the training viewpoint generation unit 72, and also performs machine learning of a wider range by using the entire training images, for example.

[0066] The number of variations in ranges represented by the 3D scene information is not limited to two, and may be three or more. The inclusion relationship between the ranges is not limited, and 3D scene information may be acquired for each of multiple independent ranges. In any case, 3D scene information is acquired for the number of set ranges. The narrower the range, the shorter the learning time and rendering time, enabling display with low latency. By utilizing this characteristic, even if the information density of the 3D scene information corresponding to the important region 224 is increased to a certain extent, if the range is narrow, rendering can be completed in the same amount of time as rendering a wide-area image using 3D scene information with low information density. As a result, it is possible to display the important region 224 with high resolution and low latency.

[0067] 7 is a schematic diagram showing the transition of data during transmission and reception of 3D scene information. The passage of time is shown from top to bottom of the diagram, with (a) to (c) showing the transition of data states in the content server 20 and (c) to (e) showing the transition of data states in the client terminal 10. This diagram also shows 3D scene information representing one scene. First, the type-specific 3D scene information acquisition unit 80 of the content server 20 acquires multiple neural networks 230a, 230b, ... corresponding to multiple types of 3D scene information, respectively, based on newly generated learning images, as shown in (a).

[0068] Next, the 3D scene information transmitting unit 86 of the content server 20 divides the neural networks 230a and 230b as shown in (b). That is, the 3D scene information transmitting unit 86 generates a plurality of neural networks 232a and 232b from the neural network 230a, excluding some mutually different nodes. The 3D scene information transmitting unit 86 also generates a plurality of neural networks 234a and 234b from the neural network 230b, excluding some mutually different nodes.

[0069] In the figure, nodes excluded from the original neural networks 230a and 230b are indicated by dotted lines in the neural networks 232a, 232b, 234a, and 234b after division. If the nodes excluded from one neural network 232a in the divided neural networks 232a and 232b are left in the other neural network 232b, the original neural network 230a can be completely restored by combining the two.

[0070] However, as mentioned above, even if missing nodes occur due to packet loss or the like, it is possible to generate an image using the remaining neural networks. The number of neural network divisions is not limited to two. The division unit 88 of the 3D scene information transmission unit 86 preferably randomly determines nodes to be excluded using a method similar to dropout. The 3D scene information transmission unit 86 packetizes the divided neural networks 232a, 232b, 234a, and 234b and transmits them to the client terminal 10.

[0071] At this time, the 3D scene information transmission unit 86 randomly changes the transmission order of packets from multiple neural networks, as shown in (c). In the figure, the horizontal arrangement indicates transmission in the order of neural networks 232b, 234b, 232a, .... By changing the transmission order, it is possible to reduce the possibility that all 3D scene information of a certain type will be lost due to consecutive packet losses. However, the switching is performed between neural networks that have been updated in parallel within a certain allowable time. This prevents unnecessary delays in the rendering process due to the switching.

[0072] The 3D scene information acquisition unit 52 of the client terminal 10 sequentially acquires the packets and then reconstructs the neural network before division by restoring the order of the extracted neural networks, as shown in (d). For this reason, the content server 20 adds metadata to each of the neural networks 232b, 234b, 232a, ... to be transmitted, indicating which of the original neural networks 230a, 230b, ... it has divided.

[0073] In the illustrated example, the 3D scene information acquisition unit 52 was able to acquire neural networks 232a and 232b, allowing the original neural network 230a to be completely restored. Meanwhile, as shown by the dotted-line frame 236, one of the divided neural networks, neural network 234a, was unable to be acquired due to packet loss for the original neural network 230b. In this case, the original neural network 230b cannot be completely restored. In either case, the image generation unit 56 of the client terminal 10 uses the acquired neural networks 232a, 232b, and 234b to render a display image 238, as shown in FIG. 1(e).

[0074] Specifically, the image generation unit 56 draws an image of the scene seen from the latest display viewpoint by volume rendering using the neural networks 232 a, 232 b, and 234 b. When multiple types of 3D scene information are transmitted with a time difference, the image generation unit 56 updates at least a portion of the display image 238 using the latest 3D scene information. This makes it possible to realize a display in which the image resolution gradually increases and important images move with particularly low latency.

[0075] 8 is a diagram illustrating the temporal relationship between machine learning in the content server 20 and image rendering in the client terminal 10. Here, as an example, assume a case in which multiple types of 3D scene information with different information densities are transmitted. The horizontal direction of the diagram is the time axis, with the learning time in the content server 20 and the rendering time of each frame of the display image in the client terminal 10 indicated by rectangles. The number attached to each rendering time rectangle indicates the frame order.

[0076] As shown in the upper part, the type-specific 3D scene information acquisition unit 80 of the content server 20 individually learns the first, second, ..., nth types of 3D scene information. Here, the ordinal numbers correspond to the information density, with the first type having the lowest information density and the nth type having the highest information density. As shown in the figure, even if the type-specific 3D scene information acquisition unit 80 starts learning multiple types of 3D scene information simultaneously, the higher the information density, the longer it takes to complete the learning. First, when the content server 20 has completed learning the first type of 3D scene information, it transmits it to the client terminal 10, as indicated by arrow A1.

[0077] The client terminal 10 uses the first type of 3D scene information to draw images of frames numbered 0, 1, and 2 at the lowest resolution, starting from time t1. Meanwhile, when the content server 20 has completed learning the second type of 3D scene information, it transmits it to the client terminal 10, as indicated by arrow A2. The client terminal 10 begins drawing images at a resolution corresponding to the second type of information density, starting with frame number 3, immediately after time t2 when the second type of 3D scene information is acquired. By repeating the same process, the frame resolution gradually increases.

[0078] When the content server 20 has completed learning the nth type of 3D scene information, it transmits it to the client terminal 10, as indicated by arrow An. The client terminal 10 begins drawing the image at the highest resolution corresponding to the nth type of information density, starting with frame number m+2, immediately after time tn when the 3D scene information was acquired. In this embodiment, multiple types of 3D scene information representing the same scene are learned, and the 3D scene information that is learned earliest is immediately transmitted to the client terminal 10. This allows the frame drawing start time to be significantly earlier than when only 3D scene information with a high information density, such as the nth type, is transmitted.

[0079] That is, based on the time when learning begins in the content server 20, rendering can begin with a delay of only the time Ta required for data transmission and the time Tb required for learning the first type of 3D scene information. In practice, before time t1, a frame may be rendered using the 3D scene information of the previous scene transmitted immediately before. Furthermore, when rendering each frame, the image generator 56 performs volume rendering based on the most recent display viewpoint acquired immediately before. This not only increases the resolution from frame number 0 to m+2, but also allows images to be displayed that reflect the movement of the viewpoint with low latency. Here, the lower the information density of the first type of 3D scene information, the fewer the number of samples in volume rendering and the higher the rendering speed, enabling display with less latency.

[0080] In this way, the image generation unit 56 can be said to have the function of correcting a generated image to match the latest display viewpoint, regardless of whether the image has already been displayed. Figure 9 is a diagram for explaining the image correction process performed by the image generation unit 56 of the client terminal 10. Image 250a is a frame or a portion of a display image, and shows an image 252a of a cylindrical object in the foreground. If the position of the cylindrical object changes relatively due to movement of the viewpoint, and the image 252b of the object shifts in image 250b after the elapse of time Δt, it becomes necessary to newly draw a background region 254 that was hidden in image 250a.

[0081] In conventional technology that displays images transmitted from the content server 20 without using 3D scene information, it is difficult to create a new region 254 that is not represented in the original image 250a. By using 3D scene information that includes information about the object and its surrounding area, all pixels of image 250b can be determined to correspond to the latest viewpoint, which naturally makes it possible to render region 254. Furthermore, if the display viewpoint moves and the positional relationship between the object and the light source changes, the color of the object image 252b may also change. In this case, too, the change in color can be accurately represented by using the 3D scene information to render an image seen from a new display viewpoint. This allows image 250c after the time Δt has elapsed to be generated with high accuracy.

[0082] As described above, assuming that the client terminal 10 uses 3D scene information to draw each frame of the display image, the images 250a and 250c are each images drawn by the client terminal 10 using the latest 3D scene information at that time. On the other hand, assuming a case in which the content server 20 transmits the display image together with the 3D scene information, the client terminal 10 can correct the image 250a transmitted from the content server 20 using the 3D scene information.

[0083] For example, the client terminal 10 generates an image 250c corrected to reflect the movement of the viewpoint that occurs during the time difference Δt between when the image 250a is generated on the content server 20 side and when it is displayed on the client terminal 10. In this case as well, the image generation unit 56 of the client terminal 10 redraws an area 254 that is not shown in the image 250a generated by the content server 20 and an image 252b whose color has changed, using the latest 3D scene information.

[0084] In this case, since the high-quality image 250a generated by the content server 20 is combined with an area newly drawn by the client terminal 10, it is desirable that the image generation unit 56 draw the necessary area using 3D scene information with high information density. Note that the image generation unit 56 can omit new drawing processing using 3D scene information in a situation where simply moving or deforming the image in the image 250a transmitted from the content server 20 does not create a sense of incongruity. For example, for a distant object where the amount of image shift or color change is small when the display viewpoint is changed, the image generation unit 56 may correct the image by directly processing the image 250a.

[0085] For this reason, the image generation unit 56 stores in advance in an internal memory or the like rules for determining areas that need to be redrawn using 3D scene information. For example, when the speed of the display viewpoint exceeds a threshold value or is predicted to exceed it, the image generation unit 56 may redraw, using 3D scene information, objects whose distance from the viewpoint is less than or equal to the threshold value and their surrounding areas. In this case, the content server 20 also transmits geometry information of the display world to the client terminal 10. This allows the image generation unit 56 to obtain changes in the distance between the display viewpoint and the object.

[0086] According to the present embodiment described above, the content server 20 generates training images representing the display world of the content and acquires multiple types of 3D scene information for each scene. The content server 20 sequentially transmits the 3D scene information for which training has been completed to the client terminal 10, and the client terminal 10 uses the information to draw frames of the display image corresponding to the latest display viewpoint. By setting various types of 3D scene information, such as the density of information and the width of the range to be represented, the training time can be changed and the delay time until display can also be controlled. Furthermore, the level of detail can be increased in certain areas, allowing for low-latency, high-quality images to be displayed that take into account the importance of the image.

[0087] Furthermore, the content server 20 divides a neural network that constitutes one piece of 3D scene information into multiple neural networks and transmits them to the client terminal 10. This increases robustness against packet loss while suppressing increases in data size. As a result, in image processing involving distribution from the content server 20, the impact of distribution can be reduced, improving the quality of the user experience.

[0088] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention.

[0089] As described above, the present invention can be used in various information processing devices such as content servers, game devices, head-mounted displays, display devices, mobile terminals, and personal computers, as well as image display systems including any of these.

[0090] The present disclosure may include the following aspects. [Item 1] A content server including a circuit configured to: generate, as training images, images representing how each scene in a three-dimensional display world, whose situation changes in response to user operations, is viewed from multiple viewpoints; acquire multiple types of 3D scene information representing three-dimensional information of each scene through machine learning using the training images as training data; and transmit data of the multiple types of 3D scene information to a client terminal that draws a display image using the 3D scene information in the order in which machine learning for each scene is completed. [Item 2] The content server according to item 1, wherein the circuit acquires multiple pieces of 3D scene information having different spatial information densities. [Item 3] The content server according to item 1, wherein the circuit divides a neural network constituting one piece of 3D scene information into multiple neural networks from which some different nodes are excluded, and transmits the multiple neural networks to the client terminal. [Item 4] The content server according to item 3, wherein the circuit packets the neural network after division and randomly changes the transmission order. [Item 5] The content server according to item 1, wherein the circuit acquires a plurality of pieces of 3D scene information representing different ranges in the display world. [Item 6] The content server according to item 5, wherein the circuit acquires the 3D scene information representing a range of the display world corresponding to a defined area in an image displayed on the client terminal. [Item 7] The content server according to item 5, wherein the circuit acquires 3D scene information targeting objects existing in the display world. [Item 8] The content server according to item 1, wherein the circuit transmits data of the plurality of types of 3D scene information to the client terminal using different communication protocols.[Item 9] A client terminal including a circuit configured to: acquire information about a user operation and information about a display viewpoint for a three-dimensional display world; acquire from a server data of multiple types of 3D scene information representing three-dimensional information acquired by machine learning for each scene in the display world whose situation changes in response to the user operation; and render at least a portion of a frame of a display image using the most recently acquired 3D scene information based on the latest display viewpoint. [Item 10] The client terminal according to Item 9, wherein the circuit acquires multiple neural networks obtained by dividing a neural network constituting one piece of 3D scene information, with some mutually different nodes excluded, and reconstructs the neural network before division. [Item 11] The client terminal according to Item 9, wherein the circuit acquires multiple pieces of 3D scene information with different spatial information densities in order of decreasing density of the information, and changes the resolution of the frame to correspond to the information density of the acquired 3D scene information. [Item 12] The client terminal according to Item 9, wherein the circuitry uses the 3D scene information to draw an area of ​​the display image transmitted from the server that is determined to need to be drawn due to a change in the display viewpoint.[Item 13] An image display system including a client terminal that displays images of a three-dimensional display world in which a situation changes in response to user operations, and a content server that transmits data used to generate display images, wherein the client terminal and the content server have circuits configured to: generate images that represent each scene of the display world as seen from a plurality of viewpoints as learning images; acquire multiple types of 3D scene information that represent three-dimensional information of each scene by machine learning using the learning images as training data; transmit the data of the multiple types of 3D scene information to the client terminal in the order in which machine learning for each scene is completed; and acquire information on the user operation and information on a display viewpoint for the display world; acquire the data of the multiple types of 3D scene information from the content server; and draw at least a portion of a frame of a display image using the 3D scene information most recently acquired based on the latest display viewpoint. [Item 14] A display data transmission method comprising: generating, as training images, images that represent each scene in a three-dimensional display world, the situation of which changes in response to user operation, as viewed from multiple viewpoints; acquiring, by machine learning using the training images as training data, multiple types of 3D scene information that represent three-dimensional information for each scene; and transmitting data of the multiple types of 3D scene information to a client terminal that draws a display image using the 3D scene information in the order in which machine learning for each scene is completed. [Item 15] A display image generation method comprising: acquiring information on user operation and information on a display viewpoint for the three-dimensional display world; acquiring, from a server, data of the multiple types of 3D scene information that represent three-dimensional information, acquired by machine learning, for each scene in the display world, the situation of which changes in response to the user operation; and drawing at least a portion of a frame of a display image using the most recently acquired 3D scene information based on the latest display viewpoint.[Item 16] A recording medium having recorded thereon a program for causing a computer to realize the following functions: generating images as training images that represent how each scene in a three-dimensional display world, the situation of which changes in response to user operation, is viewed from multiple viewpoints; acquiring multiple types of 3D scene information that represent three-dimensional information for each scene through machine learning using the training images as training data; and transmitting data of the multiple types of 3D scene information to a client terminal that draws a display image using the 3D scene information in the order in which machine learning for each scene is completed. [Item 17] A recording medium having recorded thereon a program for causing a computer to realize the following functions: acquiring information on user operation and information on a display viewpoint for the three-dimensional display world; acquiring data of multiple types of 3D scene information that represent three-dimensional information, acquired through machine learning, for each scene in the display world, the situation of which changes in response to the user operation, from a server; and drawing at least a portion of a frame of a display image using the 3D scene information most recently acquired based on the latest display viewpoint.

[0091] 1 Image display system, 10 Client terminal, 14 Input device, 16 Display device, 20 Content server, 50 Input information acquisition unit, 52 3D scene information acquisition unit, 54 3D scene information storage unit, 56 Image generation unit, 58 Output unit, 70 Input information acquisition unit, 72 Learning viewpoint generation unit, 74 Display world control unit, 76 3D model storage unit, 78 Learning image generation unit, 80 Type-specific 3D scene information acquisition unit, 84 Type-specific 3D scene information storage unit, 86 3D scene information transmission unit, 88 Splitting unit, 122 CPU, 124 GPU, 126 Main memory.

Claims

1. A content server, comprising: a learning image generation unit that generates, as learning images, images representing how each scene of a three-dimensional display world whose situation changes according to user operations is viewed from a plurality of viewpoints; a type-specific 3D scene information acquisition unit that acquires a plurality of types of 3D scene information representing three-dimensional information of each scene by machine learning using the learning images as teacher data; and a 3D scene information transmission unit that transmits data of the plurality of types of 3D scene information to a client terminal that draws a display image using the 3D scene information, in the order in which machine learning for each scene is completed.

2. The content server according to claim 1, wherein the type-specific 3D scene information acquisition unit acquires a plurality of the 3D scene information having different densities of spatial information.

3. The content server according to claim 1 or 2, wherein the 3D scene information transmission unit divides a neural network constituting one of the 3D scene information into a plurality of neural networks from which different partial nodes are excluded, and then transmits the divided neural networks to the client terminal.

4. The content server according to claim 3, wherein the 3D scene information transmission unit packetizes the divided neural networks and randomly rearranges the transmission order thereof.

5. The content server according to claim 1 or 2, wherein the type-specific 3D scene information acquisition unit acquires a plurality of the 3D scene information representing different ranges in the display world.

6. The content server according to claim 5, wherein the type-specific 3D scene information acquisition unit acquires the 3D scene information representing the range of the display world corresponding to a region defined in an image displayed on the client terminal.

7. The content server according to claim 5, wherein the type-specific 3D scene information acquisition unit acquires 3D scene information targeting an object existing in the display world.

8. The content server according to claim 1, wherein the 3D scene information transmission unit transmits data of the plurality of types of 3D scene information to the client terminal using different communication protocols.

9. An input information acquisition unit that acquires information on user operations and information on a display viewpoint for a three-dimensional display world; a 3D scene information acquisition unit that acquires, from a server, data of a plurality of types of 3D scene information representing three-dimensional information, which is obtained by machine learning, for each scene of the display world whose situation changes according to the user operation; and an image generation unit that draws at least a part of a frame of a display image using the most recently acquired 3D scene information based on the latest display viewpoint. A client terminal characterized by comprising the above components.

10. The 3D scene information acquisition unit of the client terminal according to claim 9, wherein the 3D scene information acquisition unit acquires a plurality of neural networks obtained by dividing a neural network constituting one piece of the 3D scene information and excluding different partial nodes from each other, and reconstructs the neural network before the division.

11. The 3D scene information acquisition unit of the client terminal according to claim 9 or 10, wherein the 3D scene information acquisition unit acquires a plurality of the 3D scene information with different spatial information densities in ascending order of the density of the information, and the image generation unit changes the resolution of the frame so as to correspond to the density of the information of the acquired 3D scene information.

12. The image generation unit of the client terminal according to claim 9 or 10, wherein the image generation unit draws, using the 3D scene information, a region determined to require drawing due to a change in the display viewpoint among the display images transmitted from the server.

13. A client terminal that displays an image of a three-dimensional display world whose situation changes according to a user operation, and a content server that transmits data used for generating the display image, wherein the content server includes: a learning image generation unit that generates, as learning images, images representing how each scene of the display world is viewed from a plurality of viewpoints; a type-specific 3D scene information acquisition unit that acquires a plurality of types of 3D scene information representing the three-dimensional information of each scene by machine learning using the learning images as teacher data; and a 3D scene information transmission unit that transmits the data of the plurality of types of 3D scene information to the client terminal in the order in which the machine learning for each scene is completed, and the client terminal includes: an input information acquisition unit that acquires information on the user operation and information on a display viewpoint for the display world; a 3D scene information acquisition unit that acquires the data of the plurality of types of 3D scene information from the content server; and an image generation unit that draws at least a part of a frame of the display image using the 3D scene information acquired most recently based on the latest display viewpoint. An image display system characterized by the above.

14. A step of generating, as learning images, images representing how each scene of a three-dimensional display world whose situation changes according to a user operation is viewed from a plurality of viewpoints; a step of acquiring, by machine learning using the learning images as teacher data, a plurality of types of 3D scene information representing the three-dimensional information of each scene; and a step of transmitting the data of the plurality of types of 3D scene information to a client terminal that draws a display image using the 3D scene information in the order in which the machine learning for each scene is completed. A display data transmission method characterized by including the above.

15. A step of acquiring information on a user operation and information on a display viewpoint for a three-dimensional display world; a step of acquiring, from a server, data of a plurality of types of 3D scene information representing three-dimensional information, which is acquired by machine learning, for each scene of the display world whose situation changes according to the user operation; and a step of drawing at least a part of a frame of a display image using the 3D scene information acquired most recently based on the latest display viewpoint. A display image generation method characterized by including the above.

16. A computer program, characterized in that it causes a computer to realize: a function of generating, as learning images, images representing how each scene of a three-dimensional display world whose situation changes according to a user operation is viewed from a plurality of viewpoints; a function of acquiring, by machine learning using the learning images as teacher data, a plurality of types of 3D scene information representing three-dimensional information of each scene; and a function of transmitting, to a client terminal that draws a display image using the 3D scene information, data of the plurality of types of 3D scene information in the order in which machine learning for each scene is completed.

17. A computer program, characterized in that it causes a computer to realize: a function of acquiring information on a user operation and information on a display viewpoint for a three-dimensional display world; a function of acquiring, from a server, data of a plurality of types of 3D scene information representing three-dimensional information, which is acquired by machine learning, for each scene of the display world whose situation changes according to the user operation; and a function of drawing at least a part of a frame of a display image using the 3D scene information acquired most recently based on the latest display viewpoint.

Citation Information

Patent Citations

  • Indoor scene three-dimensional reconstruction system and method based on neural radiation field

    CN114004941A

  • Transmitting device and receiving device

    JP2020005201A

  • Robust View Synthesis for Unconstrained Image Data

    JP2023543538A

  • Processing and / or transmitting 3D data

    US20160300385A1

  • System and method for dynamically adjusting level of details of point clouds

    US20210035352A1