Content server, client terminal, image display system, draw data transmission method, and display image generation method
The image processing system addresses the challenge of displaying high-quality images with responsive changes in viewpoints and gazes by using a content server to generate and transmit texture models and geometry data to client terminals, resulting in improved user experience with high-quality and timely image rendering.
Patent Information
- Application Number
- PCT/JP2023/039246
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
Existing image processing systems struggle to display high-quality images with high responsiveness to changes in viewpoints and gazes, particularly in systems where content is distributed from a server to client terminals.
The system comprises a content server equipped with an input information acquisition unit, a learning texture generating unit, a texture model generating unit, and a drawing data transmitting unit, which work together to generate and transmit texture models and geometry data to client terminals, enabling the generation of high-quality images that adapt to changes in viewpoints and gazes.
This solution allows for the display of high-quality images with minimal delay in response to changes in viewpoints and gazes, effectively improving the user experience by enhancing the responsiveness and quality of the image processing.
Smart Images

Figure JP2023039246_08052025_PF_FP_ABST
Abstract
Description
Content server, client terminal, image display system, drawing data transmission method, and display image generation method
[0001] The present invention relates to a content server, a client terminal, an image display system, a drawing data transmission method, and a display image generation method for displaying an image of a three-dimensional display world.
[0002] With the recent expansion of communication networks and advances in image processing technology, it has become possible to enjoy a wide variety of electronic content regardless of the viewing environment. For example, in the field of electronic games, a system has become widespread in which a server collects operation information entered into each client terminal and distributes game images that reflect that information as needed, allowing multiple players to participate in the same game regardless of location.
[0003] This is not limited to electronic games, but any electronic content that generates video in real time in response to user operations and distributes it from a server can utilize the server's abundant processing environment, making it easier to display high-quality images while minimizing the impact of the client terminal's processing performance. However, the transmission of operation information from the client terminal and the processing of video distribution from the server that receives it are constantly involved, making it difficult to immediately reflect changes in viewpoint or line of sight that occur during this time in the display.
[0004] The present invention has been made in consideration of these problems, and its purpose is to provide a technology that displays high-quality images with high responsiveness to changes in viewpoint or line of sight when processing images of content that is distributed from a server.
[0005] To solve the above problems, one aspect of the present invention relates to a content server, comprising: an input information acquisition unit that acquires, from a client terminal that generates a display image representing a three-dimensional display world, information on a display viewpoint that defines the display image; a training texture generation unit that generates a plurality of training textures that represent a distribution of color values on the surface of the object by drawing an image of the object as seen from a plurality of training viewpoints generated based on the display viewpoints; a texture model generation unit that generates a texture model that outputs color values corresponding to changes in viewpoint through machine learning using the training textures as training data; and a drawing data transmission unit that associates geometry data of the display world with the texture model and transmits the associated data to the client terminal.
[0006] Another aspect of the present invention relates to a client terminal, comprising: an input information acquisition unit that acquires information on a display viewpoint that defines a display image representing a three-dimensional display world; a drawing data acquisition unit that acquires, from a server, a texture model that outputs color values corresponding to changes in viewpoint, the texture model being generated by machine learning using a training texture that represents a distribution of color values of the surface of an object viewed from a plurality of training viewpoints, along with geometry data of the display world; and an image generation unit that generates a display image by acquiring color values corresponding to the latest display viewpoint using the texture model and outputs the display image to a display device.
[0007] Yet another aspect of the present invention relates to an image display system including a client terminal that generates and displays a display image representing a three-dimensional display world, and a content server that transmits drawing data used to generate the display image, wherein the content server includes an input information acquisition unit that acquires, from the client terminal, information about a display viewpoint that defines the display image, a learning texture generation unit that generates a plurality of learning textures that represent a distribution of color values on the surface of an object by drawing images of the object as seen from a plurality of learning viewpoints generated based on the display viewpoints, a texture model generation unit that generates a texture model that outputs color values corresponding to changes in the viewpoint through machine learning using the learning textures as training data, and a drawing data transmission unit that associates geometry data of the display world with the texture model and transmits the associated data to the client terminal, and wherein the client terminal includes an input information acquisition unit that acquires information about the display viewpoint, a drawing data acquisition unit that acquires the texture model and the geometry data, and an image generation unit that generates a display image by acquiring color values corresponding to the latest display viewpoint using the texture model, and outputs the display image to a display device.
[0008] Yet another aspect of the present invention relates to a drawing data transmission method, the drawing data transmission method comprising the steps of: acquiring, from a client terminal that generates a display image representing a three-dimensional display world, information on a display viewpoint that defines the display image; generating a plurality of training textures that represent a distribution of color values on the surface of the object by drawing images of the object as seen from a plurality of training viewpoints generated based on the display viewpoints; generating a texture model that outputs color values corresponding to changes in the viewpoint by machine learning using the training textures as training data; and transmitting geometry data of the display world and the texture model in association with each other to the client terminal.
[0009] Yet another aspect of the present invention relates to a display image generation method, the display image generation method including the steps of: acquiring information on a display viewpoint that defines a display image representing a three-dimensional display world; acquiring, from a server, a texture model that outputs color values corresponding to a change in viewpoint, the texture model being generated by machine learning using a training texture that represents a distribution of color values of the surface of an object viewed from a plurality of training viewpoints, together with geometry data of the display world; and acquiring color values corresponding to the latest display viewpoint using the texture model, thereby generating a display image and outputting the display image to a display device.
[0010] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention.
[0011] According to the present invention, in image processing of content distributed from a server, high-quality images can be displayed with high responsiveness to changes in viewpoint or line of sight.
[0012] FIG. 1 is a diagram illustrating an example configuration of an image display system to which the present embodiment can be applied. FIG. 1 is a diagram illustrating a display delay in response to a change in the display viewpoint in a mode in which a content server generates a display image. FIG. 2 is a diagram illustrating an image determined by ray tracing. FIG. 3 is a diagram illustrating data transmitted from a content server to a client terminal in the present embodiment. FIG. 4 is a diagram illustrating a visible range texture generated by a content server in the present embodiment. FIG. 5 is a diagram illustrating an internal circuit configuration of a client terminal in the present embodiment. FIG. 6 is a diagram illustrating functional block configurations of a client terminal and a content server in the present embodiment. FIG. 7 is a diagram schematically illustrating how a learning texture generation unit of a content server generates a learning texture in the present embodiment. FIG. 8 is a diagram illustrating a procedure in which an image generation unit of a client terminal generates a display image using a texture model in the present embodiment. FIG. 9 is a flowchart illustrating a processing procedure in which a content server generates drawing data and transmits it to a client terminal in the present embodiment. FIG. 10 is a flowchart illustrating a processing procedure in which a client terminal generates a display image using drawing data and outputs it in the present embodiment.
[0013] 1 shows an example of the configuration of an image display system to which this embodiment can be applied. The image display system 1 includes client terminals 10a, 10b, and 10c that display images in response to user operations, and a content server 20 that provides data used for display. Input devices 14a, 14b, and 14c for user operations and display devices 16a, 16b, and 16c for displaying images are connected to the client terminals 10a, 10b, and 10c, respectively. Communication can be established between the client terminals 10a, 10b, and 10c and the content server 20 via a network 8 such as a WAN (World Area Network) or a LAN (Local Area Network).
[0014] The client terminals 10a, 10b, and 10c may be connected to the display devices 16a, 16b, and 16c and the input devices 14a, 14b, and 14c either wired or wirelessly. Alternatively, two or more of these devices may be integrated. For example, in the figure, the client terminal 10b is connected to a head-mounted display, which is the display device 16b. The head-mounted display can change the field of view of the displayed image by the movement of the user wearing it on their head, so it also functions as the input device 14b.
[0015] Furthermore, the client terminal 10c is a portable terminal that is integrally configured with a display device 16c and an input device 14c, which is a touchpad that covers the display device 16c's screen. As such, the external shape and connection form of the illustrated devices are not limited. The number of client terminals 10a, 10b, 10c and content servers 20 that can be connected to the network 8 is also not limited. Hereinafter, the client terminals 10a, 10b, 10c will be collectively referred to as client terminals 10, the input devices 14a, 14b, 14c as input device 14, and the display devices 16a, 16b, 16c as display device 16.
[0016] The input device 14 may be any one or a combination of general input devices such as a controller, keyboard, mouse, touchpad, or joystick, or various sensors such as a motion sensor or camera provided in a head-mounted display, and supplies the content of user operations to the client terminal 10. The display device 16 may be a general display such as a liquid crystal display, plasma display, organic EL display, wearable display, or projector, and displays images output from the client terminal 10.
[0017] The content server 20 provides data of content accompanied by image display to the client terminal 10. The type of content is not particularly limited and may be any of electronic games, decorative images, web pages, video chat using avatars, etc. In this embodiment, the content server 20 sequentially acquires information on user operations on the input device 14 from the client terminal 10, reflects the information in the world to be displayed, and displays an image representing the information on the client terminal 10 side.
[0018] In this embodiment, the display image is rendered using three-dimensional computer graphics (3DCG). In the field of 3DCG, realistic image representation is possible by more accurately representing physical phenomena occurring in the space of the display target. Ray tracing is known as a physically based rendering method that achieves this. Ray tracing accurately calculates the propagation of various types of light that reach the viewpoint, such as light from a light source as well as diffuse reflection and specular reflection on the surface of an object, making it possible to more realistically represent changes in color and brightness due to movement of the viewpoint or the object itself.
[0019] By utilizing the abundant processing environment of the content server 20 and generating high-resolution images at a high rate using ray tracing, it is possible to enjoy content with high quality regardless of the processing performance of the client terminal 10. On the other hand, it becomes necessary to send and receive various types of data between the content server 20 and the client terminal 10, which poses a problem of a tendency for delays in display to occur in response to user operations on the client terminal 10 side or changes in the viewpoint or line of sight relative to the displayed world. Hereinafter, the position of the viewpoint or line of sight relative to the displayed world will simply be referred to as the "viewpoint," and the viewpoint corresponding to the display field of view will sometimes be referred to as the "display viewpoint."
[0020] 2 is a diagram illustrating a delay in display in response to a change in the display viewpoint in a mode in which the content server 20 generates a display image. When an image generated by the content server 20 is displayed on the client terminal 10 side, as described above, it takes a certain amount of time from when the client terminal 10 transmits information about the display viewpoint to the content server 20 until the image generated in response to that information is displayed on the client terminal 10 side. As a result, a delay occurs in the field of view of the displayed image in response to an actual change in the display viewpoint.
[0021] 1A shows how content server 20 generates an image. Content server 20 sets view screen 280a to correspond to the display viewpoint recognized at that time, and renders image 284 contained in corresponding view frustum 282a on view screen 280a. Assume here that the viewpoint at the time of display has shifted to the left, as indicated by the arrow. In this case, when the transmitted image is displayed on client terminal 10, view screen 280b will have shifted to the left, as shown in 1B.
[0022] That is, the original field of view 290 corresponding to the viewing frustum 282b at the time of display will be shifted from the field of view 286 of the image transmitted from the content server 20. If this is ignored and the image transmitted from the content server 20 is displayed as is, a delay will occur in the field of view of the displayed image relative to changes in the display viewpoint, which may cause an unacceptable sense of discomfort. In particular, when the display device 16 is a head-mounted display, this may impair the sense of immersion in virtual reality or cause visually induced motion sickness, thereby degrading the quality of the user experience.
[0023] One solution to this problem is reprojection, in which the client terminal 10 corrects the field of view to match the viewpoint at the time of display. For example, a wide-range image that takes into account the movement of the display viewpoint can be transmitted from the content server 20, and the client terminal 10 can crop and display the field of view at the time of display. This at least corrects the delay in the display field of view, but because the information supplied to the client terminal 10 is limited to two-dimensional information, it is difficult to accurately represent changes in color that depend on the direction from which an object is viewed, as can be achieved with ray tracing.
[0024] FIG. 3 is a diagram illustrating an image determined by ray tracing. In ray tracing, rays are generated from the display viewpoint 102 to pass through each pixel on the view screen 106, and pixel values are determined by sampling the color at the point where the rays reach. This allows accurate representation of not only the color of the object itself due to diffuse reflection, but also shadows, reflections due to specular reflection, and images transmitted through translucent objects. In the example shown, a ray 156 arriving at point 154 on the surface of a spherical object 150a may probabilistically reach light sources 152a and 152b (rays 158a and 158b), or may reach another object 150c due to specular reflection (ray 158c).
[0025] If object 150a is semi-transparent, ray 150d passes through the object from point 154 and is refracted, reaching another object 150b. The rays that reach the other objects 150b and 150c eventually reach light sources 152a and 152b. The color of point 154 is represented by the superposition of the colors of these rays. In other words, the color of point 154 reflects not only the color of object 150a itself, but also the colors of the other objects 150b and 150c.
[0026] As a result, the surface of object 150a displays a reflected image of another object 150c and an image of object 150b being seen through it. In the case of an object such as a floor (not shown), a shadow image is formed depending on the positional relationship between the light source and other object 150a. If the position of display viewpoint 102, the direction of the line of sight, and ultimately the incident direction of ray 156 reaching spherical object 150a change, the subsequent ray path also changes, and as a result, the color of point 154 also changes. Hereinafter, this change in color of the same point in response to a change in viewpoint will be referred to as a "viewpoint-dependent effect."
[0027] If such ray tracing were performed on the client terminal 10 side, the display world could be expressed with accurate colors for the viewpoint at the time of display, but this would increase the processing load and could be difficult to achieve depending on the processing capacity of the client terminal 10. Therefore, in this embodiment, by generating color information that can respond to changes in the display viewpoint in the content server 20, it becomes possible for the client terminal 10 to generate a display image that matches both the field of view and color to the viewpoint at the time of display with a light processing load.
[0028] 4 is a diagram illustrating data transmitted from content server 20 to client terminal 10 in this embodiment. Content server 20 controls three-dimensional space 200 of the display world in accordance with the settings of content such as an electronic game. In the example shown in the figure, multiple objects 200a, 200b, and 200c exist in three-dimensional space 200. Content server 20 generates geometry data representing the positions and shapes of objects 202a, 202b, and 202c in three-dimensional space 200, and textures 204a, 204b, and 204c representing the color distribution of the surfaces of objects 202a, 202b, and 202c.
[0029] As shown in the figure, textures 204a, 204b, and 204c are data in which the color value c of each point is mapped onto a planar area formed by two-dimensionally UV-expanding the surface of the object. The color value c is composed of, for example, the three primary colors R (red), G (green), and B (blue). The client terminal 10 acquires the geometry data and textures 204a, 204b, and 204c and generates a display image 206 by ray tracing or rasterization.
[0030] When ray tracing is used, the client terminal 10 sets a view screen corresponding to the viewpoint immediately before display, generates rays for each pixel on the view screen, and obtains the destination on the object using geometry data. The client terminal 10 then obtains pixel values for the display image 206 by sampling the color values of the destination from textures 204 a, 204 b, and 204 c. When rasterization is used, the client terminal 10 uses geometry data to project polygons of the object onto the view screen corresponding to the viewpoint immediately before display, and applies areas of textures 204 a, 204 b, and 204 c corresponding to each polygon to obtain pixel values for the display image 206.
[0031] By generating high-resolution textures 204a, 204b, and 204c in the content server 20, it is possible to generate a high-quality display image 206 on the client terminal 10 with a field of view that matches the viewpoint at the time of display, even if data transfer time occurs. Furthermore, in order to express viewpoint-dependent effects as described above, the content server 20 of this embodiment prepares texture data with variations that take into account changes in the viewpoint over transfer time. In the figure, data with such variations is represented by the overlay of multiple textures (e.g., texture data 208).
[0032] The multiple textures are generated from different viewpoints relative to the object. Therefore, viewpoint-dependent effects can be expressed by the client terminal 10 changing the reference texture in accordance with changes in the display viewpoint. When the objects 200a, 200b, and 200c move or change shape in response to user operations, etc., the content server 20 updates the geometry data and texture data at predetermined time steps and transmits them to the client terminal 10. This makes it possible to express changes in color in response to both changes in the viewpoint and changes in the object itself.
[0033] In this embodiment, the texture data 208 is actually a neural network representing the results of machine learning using multiple textures from different viewpoints as training data. The client terminal 10 inputs information about the viewpoint immediately before display into the neural network to obtain appropriate color values at that time and use these as pixel values for the display image 206. This makes it possible to naturally represent changes in the color of objects in response to changes in the display viewpoint, even if the content server 20 actually generates a limited number of textures.
[0034] Neural Radiance Fields (NeRF) is a method for representing three-dimensional space using a neural network (see, for example, Ben Mildenhall et al., "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis," Communications of the ACM, January 2022, Vol. 65, No. 1, pp. 99-106). This technology uses multiple images representing a common scene as training data and performs regression using a multilayer perceptron (MLP) to obtain data representing three-dimensional information about the scene. This data is a function f consisting of a neural network that takes five-dimensional parameters consisting of position coordinates (x, y, z) and direction vectors (θ, φ) in three-dimensional space as input and outputs volume density σ and color value c (RGB), as shown below.
[0035]
[0036] If the content server 20 generates such NeRF data at a predetermined time step and the client terminal 10 that receives it performs volume rendering, it is possible to display an image corresponding to the viewpoint at the time of display. However, in this case, the learning cost in the content server 20 and the rendering cost in the client terminal 10 are large, which may actually result in a decrease in processing efficiency. Therefore, as described above, the content server 20 of this embodiment derives the function F that outputs only the color value c by machine learning. That is, the function F used as texture data in this embodiment can be expressed as follows:
[0037]
[0038] Here, (u, v) are the position coordinates of the texture in the UV coordinate system. In other words, function F is a neural network that takes four-dimensional parameters consisting of position coordinates (u, v) and a direction vector (θ, φ) as input and outputs a color value c (RGB). Function F makes it possible to appropriately change the color value c of the position coordinates (u, v) on the texture corresponding to a point on the object surface using the line-of-sight direction vector (θ, φ). Furthermore, since the order of the parameters is lower than in NeRF, it is possible to reduce learning costs and rendering costs.
[0039] A method of applying NeRF to obtain the texture of an object through machine learning is disclosed, for example, in "MobileNeRF: Exploiting the Polygon Rasterization Pipeline for Efficient Neural Field Rendering on Mobile Architectures" by Zhiqin Chen and three others, 2023 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 16569-16578, and this method can also be applied in this embodiment. Hereinafter, the neural network representing the function F will be referred to as a "texture model."
[0040] The content server 20 generates multiple textures to be used for machine learning, based on the display viewpoint at the time when information is supplied from the client terminal 10, and then anticipates changes in the viewpoint over the short time until display. In other words, if all anticipated changes in viewpoint over that short time are covered, color information for objects visible from other viewpoints is not required. Therefore, in this embodiment, the textures used for learning on the object surface are limited to the visible range, thereby speeding up the texture data generation process in the content server 20.
[0041] 5 is a diagram for explaining the visible range textures generated by the content server 20 in this embodiment. The content server 20 basically generates viewpoints for generating textures to be used for learning around the display viewpoint acquired at that time. Hereinafter, the viewpoints generated, including the display viewpoint at that time, will be referred to as "learning viewpoints," and the textures generated for those viewpoints will be referred to as "learning textures." The figure shows a bird's-eye view of a three-dimensional space including a virtual camera 210 corresponding to a certain learning viewpoint and objects 212a, 212b, 212c, and 214.
[0042] The range of the surfaces of objects 212a, 212b, 212c, and 214 that is visible from virtual camera 210 is limited depending on their position and direction. In the figure, this range is indicated by a thick line (e.g., thick line 216). For each of objects 212a, 212b, and 212c, content server 20 generates learning textures 218a, 218b, and 218c, each of which has color information only within the visible range. For example, content server 20 sets a view screen 220 corresponding to virtual camera 210 and performs ray tracing. During this process, color information of the destination of rays reaching each of objects 212a, 212b, and 212c is obtained using the principle shown in FIG. 3 and written to a texture buffer.
[0043] As a result, as shown in the figure, learning textures 218a, 218b, and 218c are generated, with color values stored only in a portion of the texture. The content server 20 generates similar textures by changing the position and direction of the virtual camera 210, and then performs machine learning for each object to generate a texture model. Note that in the state shown in the figure, the object 214 is hidden by the other objects 212b and 212c and is not visible to the virtual camera 210. In this case, the content server 20 does not need to generate a texture for the object 214. Even if the other objects 212b and 212c are transparent and the object 214 can be seen through, the image created through such transparency is reflected in the textures 218b and 218c of the other objects 212b and 212c.
[0044] The area on the object surface represented by one pixel on the view screen 220 depends on the distance of the object from the virtual camera 210, with the area becoming smaller the closer the object is to the virtual camera 210. In other words, as described above, the resolution of the texture generated by the content server 20 changes depending on the distance of the object. As a result, the density of the information volume of the texture model obtained by learning the texture becomes greater the closer the object is to the virtual camera 210, and this trend coincides with the trend of the level of detail required for each object when generating a display image. As a result, the process from generation to display of the texture model can be performed quickly with high processing efficiency without incurring excessive learning or rendering costs.
[0045] The content server 20 may limit the objects within the visible range that exhibit viewpoint-dependent effects. For example, objects located close to the display viewpoint are likely to experience significant changes in surface color with a slight change in viewpoint, while objects located farther away are less likely to experience changes in surface color with a change in viewpoint. Therefore, for distant objects that do not exhibit viewpoint-dependent effects and therefore cause little discomfort to the user, the content server 20 may not perform texture machine learning and may instead transmit a single texture generated for the display viewpoint to the client terminal 10. By generating a texture model by limiting processing to the necessary and sufficient steps as described above, accurate images can be displayed with low latency, even after undergoing a process of preparing images from multiple viewpoints.
[0046] 6 shows the internal circuit configuration of the client terminal 10. The client terminal 10 includes a CPU (Central Processing Unit) 22, a GPU (Graphics Processing Unit) 24, and a main memory 26. These components are interconnected via a bus 30. An input / output interface 28 is also connected to the bus 30. Connected to the input / output interface 28 are a communication unit 32 including a peripheral device interface such as USB or IEEE 1394 or a network interface for a wired or wireless LAN, a storage unit 34 such as a hard disk drive or nonvolatile memory, an output unit 36 that outputs data to the display device 16, an input unit 38 that inputs data from the input device 14, and a recording medium drive unit 40 that drives a removable recording medium such as a magnetic disk, optical disk, or semiconductor memory.
[0047] The CPU 22 controls the entire client terminal 10 by executing an operating system stored in the storage unit 34. The CPU 22 also executes various programs read from a removable recording medium and loaded into the main memory 26, or downloaded via the communication unit 32. The GPU 24 has the functions of a geometry engine and a rendering processor, performs drawing processing in accordance with drawing commands from the CPU 22, and stores display images in a frame buffer (not shown). The GPU 24 then converts the display images stored in the frame buffer into video signals and outputs them to the output unit 36. The main memory 26 is composed of RAM (Random Access Memory) and stores programs and data required for processing. The content server 20 may also have a similar internal circuit configuration.
[0048] 7 shows the functional block configuration of the client terminal 10 and the content server 20 in this embodiment. Note that the client terminal 10 and the content server 20 may perform various processes necessary for implementing the content, such as audio processing, but the diagram shows functional blocks related to image processing.
[0049] The illustrated functional blocks can be realized in hardware terms by the configuration of the CPU, GPU, various memories, etc. shown in Figure 6, and in software terms by programs that are loaded into memory from a recording medium, etc. and that perform various functions such as data input function, data storage function, image processing function, and communication function. Therefore, it will be understood by those skilled in the art that these functional blocks can be realized in various forms using only hardware, only software, or a combination thereof, and are not limited to any one of them.
[0050] The client terminal 10 includes an input information acquisition unit 50 that acquires input information such as user operations, a drawing data acquisition unit 52 that acquires data for image drawing from the content server 20, an image generation unit 54 that generates a display image, and an output unit 56 that outputs display image data. The input information acquisition unit 50 acquires the content of user operations from the input device 14 as needed. User operations include selecting and launching content, and inputting commands for content currently being executed.
[0051] The input information acquisition unit 50 also acquires information about a display viewpoint for the displayed world from the input device 14 or the head-mounted display at any time or at predetermined time intervals. Techniques for detecting the position and posture of the head of a user wearing a head-mounted display and acquiring viewpoint information based on the detected information are well known, and such techniques may also be applied to this embodiment. The input information acquisition unit 50 supplies the acquired information to the content server 20 and the image generation unit 54 as appropriate.
[0052] The drawing data acquisition unit 52 acquires drawing data used to generate a display image from the content server 20 and decodes and expands it as appropriate. Here, the drawing data includes geometry data representing the positions and three-dimensional shapes of objects existing in the display world, texture models associated with objects that express viewpoint-dependent effects, and texture data associated with objects that do not express viewpoint-dependent effects. The drawing data acquisition unit 52 acquires sets of drawing data at the display frame rate or a predetermined rate set separately.
[0053] The image generation unit 54 generates a display image at the display frame rate using the rendering data. That is, the image generation unit 54 uses geometry data from the most recently acquired rendering data to place an object in the three-dimensional space to be displayed, acquires the latest display viewpoint information from the input information acquisition unit 50, and sets the corresponding viewscreen. The image generation unit 54 then draws an image of the object on the viewscreen by ray tracing or rasterization. At this time, the image generation unit 54 determines the color of an object that expresses a viewpoint-dependent effect using a texture model. Specifically, the image generation unit 54 acquires the color value of a point on the object surface that corresponds to a pixel to be rendered using the texture model based on the corresponding UV coordinates and the line of sight direction.
[0054] On the other hand, for objects that do not exhibit viewpoint-dependent effects, the image generation unit 54 determines the color using a texture. Specifically, the image generation unit 54 samples the color values of points on the object surface that correspond to the pixels to be drawn from the texture based on the corresponding UV coordinates. For objects that exhibit viewpoint-dependent effects, the image generation unit 54 may generate the entire texture corresponding to the latest display viewpoint using a texture model and use it to draw the image. In this case, the image generation unit 54 may sample the color values of points on the object surface that correspond to the pixels to be drawn from the texture it has generated. The output unit 56 outputs the display image data generated by the image generation unit 54 to the display device 16 at the display frame rate and displays the image.
[0055] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, a learning viewpoint generation unit 72 that generates a learning viewpoint, a display world control unit 74 that controls the display world, a three-dimensional model storage unit 76 that stores three-dimensional models of objects, a target object selection unit 78 that selects an object that expresses a viewpoint-dependent effect, a learning texture generation unit 80 that generates learning textures, a texture model generation unit 82 that generates a texture model, a drawing data formation unit 84 that forms a data set for image drawing, and a drawing data transmission unit 86 that transmits image drawing data to the client terminal 10.
[0056] The input information acquisition unit 70 acquires information on the content of user operations and viewpoints from the client terminal 10 at any time or at predetermined time intervals. The training viewpoint generation unit 72 generates a plurality of training viewpoints for generating training textures. The training viewpoint generation unit 72 generates the most recent display viewpoint acquired by the input information acquisition unit 70 and training viewpoints around it according to predetermined rules. As described above, it is sufficient for the texture model to be information that covers the movable range of the viewpoint for the longest period of time until the client terminal 10 displays the data using that information.
[0057] Therefore, the training viewpoint generation unit 72 may, for example, evenly distribute a predetermined number of training viewpoints within a sphere of a predetermined radius centered on the latest display viewpoint. Here, the predetermined radius may be, for example, a value obtained by multiplying the maximum speed expected for viewpoint movement by the longest time until display using the data is performed. The training viewpoint generation unit 72 is not limited to distributing the training viewpoints evenly, and may also distribute more training viewpoints in a range where the viewpoint is expected to move, depending on the situation in the displayed world, etc. Furthermore, the training viewpoint generation unit 72 may evenly set line of sight in a predetermined number of directions for the viewpoint at each position, or may set more line of sight in directions where the line of sight is expected to move.
[0058] The display world control unit 74 controls the three-dimensional display world represented as content in accordance with the details of user operations acquired by the input information acquisition unit 70. For example, if the content is an electronic game, the display world control unit 74 places necessary objects such as a user character in the virtual space that serves as the stage for the electronic game, and gives them movement in accordance with commands entered by the user and program specifications. The three-dimensional model storage unit 76 stores three-dimensional models of objects existing in the display world, which the display world control unit 74 reads out as needed to use in constructing the display world.
[0059] The display world control unit 74 further sets, for the constructed display world, the display viewpoint acquired by the input information acquisition unit 70. The display world control unit 74 appropriately supplies model data and the like of objects that make up the latest display world to the learning texture generation unit 80 and the drawing data formation unit 84.
[0060] The target object selection unit 78 selects an object that exhibits a viewpoint-dependent effect based on the latest state of the display world controlled by the display world control unit 74. For example, the target object selection unit 78 selects an object that exists within a predetermined range from the display viewpoint in the display world as the target object. However, as long as an object that is highly likely to exhibit a viewpoint-dependent effect can be selected, the selection criteria are not limited to this, and may also be based on the object's original size, apparent size, material, original color, etc. The selection criteria may be one or a combination of multiple criteria.
[0061] The training texture generation unit 80 generates training textures for an object selected as a target for expressing viewpoint-dependent effects for the multiple training viewpoints generated by the training viewpoint generation unit 72. The training texture generation unit 80 preferably uses a technique capable of rendering high-quality images, such as path tracing, to generate multiple training textures for each object that represent only color information in the visible range, as described above. The texture model generation unit 82 performs machine learning using the training textures as training data, thereby generating a texture model for each object, as described above.
[0062] The drawing data generation unit 84 generates a drawing data set to be transmitted to the client terminal 10. To this end, the drawing data generation unit 84 generates a texture corresponding to the latest display viewpoint for objects that do not express viewpoint-dependent effects. In this case, the drawing data generation unit 84 also generates a texture representing only color information of the visible region for each object using a technique capable of rendering high-quality images, such as path tracing. The drawing data generation unit 84 also generates geometry data for the display world. Here, the drawing data generation unit 84 may generate geometry data representing only information about portions of objects present in the display world that are visible from the display viewpoint.
[0063] As a technique for transmitting such geometry data, for example, the technique disclosed in "Shading Atlas Streaming" by Joerg H. Muller and six others, ACM Transactions on Graphics, November 2018, Vol. 37, No. 6, Article No. 199 can be applied. This technique can further reduce the size of data transmitted from the content server 20 to the client terminal 10. The drawing data formation unit 84 associates the texture model generated by the texture model generation unit 82 for objects that express viewpoint-dependent effects, and the texture generated by the drawing data formation unit 84 for other objects, with the geometry data, and generates the resulting drawing data.
[0064] The rendering data generation unit 84 may set a difference in the level of detail of the geometry data, and therefore the coarseness of the polygons, between objects that express viewpoint-dependent effects and other objects. If viewpoint-dependent effects can be expressed precisely, the user can recognize the shape by changes in color, so even if the polygons are coarser, i.e., the number of polygons per unit area is reduced, the impact on appearance is small. Therefore, the rendering data generation unit 84 may adjust the polygons of objects that express viewpoint-dependent effects so that they are coarser than those of other objects. Existing technologies can be applied to adjust the coarseness of the polygons.
[0065] This further reduces the size of the geometry data transmitted from the content server 20 to the client terminal 10. The drawing data transmitting unit 86 appropriately compresses and encodes the set of drawing data formed by the drawing data forming unit 84, and transmits the compressed data to the client terminal 10 at a predetermined rate.
[0066] 8 is a schematic diagram showing how the training texture generator 80 of the content server 20 generates training textures. The upper part of the figure, like Fig. 5, shows a bird's-eye view of a three-dimensional space including a virtual camera 210 corresponding to a viewpoint and objects 212a, 212b, 212c, and 214. As shown in (a), (b), (c), and so on, the training texture generator 80 sets the virtual camera 210 to correspond to a plurality of training viewpoints, and generates high-quality images of objects 212a, 212b, and 212c within the visible range.
[0067] As a result, multiple training textures with different viewpoints are generated for each object. In the figure, training textures 240a, 242a, 244a, ... are generated for object 212a, training textures 240b, 242b, 244b, ... are generated for object 212b, and training textures 240c, 242c, 244c, ... are generated for object 212c. Due to the viewpoint-dependent effect, training textures are generated in which the color changes depending on the viewpoint, even for the same point on the surface of the same object.
[0068] Furthermore, each training texture represents color information within the visible range as seen from a respective viewpoint. Therefore, even training textures for the same object have different pixel ranges containing color information. By performing machine learning using these textures as training data, the discrete states of (a), (b), (c), etc. are smoothly interpolated, color values that change in response to changes in the viewpoint are obtained, and changes in newly visible parts can be naturally represented. Note that, as described above, in the illustrated example, object 214 is not visible from virtual camera 210, so content server 20 does not generate a training texture.
[0069] 9 is a diagram illustrating the procedure by which the image generation unit 54 of the client terminal 10 generates a display image using a texture model. First, the image generation unit 54 places objects 248a, 248b, and 248c in the three-dimensional space to be displayed, based on the geometry data transmitted from the content server 20. As described above, the geometry data of the objects 248a, 248b, and 248c may include only information about the visible range. Furthermore, the image generation unit 54 acquires information about the latest display viewpoint from the input information acquisition unit 50, and sets the virtual camera 250 and view screen 252 corresponding to that information.
[0070] The image generation unit 54 uses ray tracing or rasterization to draw images of objects 248a, 248b, and 248c on the view screen 252. For example, when determining the color of pixel 256, the image generation unit 54 identifies the object 248c corresponding to the image 254, and then obtains the UV coordinates (u, v) of a point 258 on the surface that corresponds to pixel 256. The image generation unit 54 also obtains the direction (θ, φ) of a line of sight 260 relative to the surface of object 248c.
[0071] The image generation unit 54 then inputs the obtained parameters (u, v, θ, φ) into a texture model 262 associated with the object 248c to obtain a color value 264 for the pixel 256. The image generation unit 54 sequentially determines pixel values for objects that express viewpoint-dependent effects using a similar process, thereby rendering the image. For objects that do not express viewpoint-dependent effects, the image generation unit 54 sequentially determines pixel values and renders the image by sampling the texture associated with the object based on the UV coordinates corresponding to the pixels.
[0072] Next, the operation of the content server 20 and the client terminal 10, which can be realized by the above-described configuration, will be described. Fig. 10 is a flowchart showing the processing procedure in which the content server 20 generates drawing data and transmits it to the client terminal 10. This flowchart starts when a user has selected content to be displayed on the client terminal 10 with which communication has been established and is displaying an initial image. Note that although the diagram shows each processing step being performed in order, any processing step may be performed in parallel with the others.
[0073] In this state, the input information acquisition unit 70 of the content server 20 acquires viewpoint information and the content of the user operation from the client terminal 10 (S10). The display world control unit 74 updates the state of the objects in the display world as appropriate in accordance with the content of the user operation, etc. (S12). The target object selection unit 78 selects an object that expresses a viewpoint-dependent effect in accordance with the distance from the latest display viewpoint, etc. (S14). For example, the target object selection unit 78 selects an object that exists within a predetermined range from the display viewpoint as the target object.
[0074] The training texture generation unit 80 then uses ray tracing or the like to draw an image representing the target object as viewed from one of the multiple training viewpoints generated by the training viewpoint generation unit 72 based on the display viewpoint. This allows the training texture generation unit 80 to generate training textures for each target object (S16). The training texture generation unit 80 repeats the process of S16 until training textures have been generated for all training viewpoints (N in S18). Once training textures have been generated for all training viewpoints (Y in S18), the texture model generation unit 82 performs machine learning using the generated training textures as training data to generate a texture model for each target object (S20).
[0075] On the other hand, for objects that do not exhibit viewpoint-dependent effects, the drawing data formation unit 84 generates textures by drawing images that represent the appearance as seen from the display viewpoint using ray tracing or the like, and also generates geometry data for the visible range (S22).The drawing data formation unit 84 then associates the texture models of objects that exhibit viewpoint-dependent effects with the textures and geometry data of objects that do not exhibit viewpoint-dependent effects, to form a set of drawing data.
[0076] The drawing data transmission unit 86 appropriately compresses and encodes the drawing data and transmits it to the client terminal 10 (S24). If there is no need to end the image display, such as when the content ends, the content server 20 repeats the processes from S10 to S24 (N in S26). If it becomes necessary to end the image display, the content server 20 ends all processes (Y in S26).
[0077] 11 is a flowchart showing the processing steps by which the client terminal 10 generates and outputs a display image using drawing data. This flowchart is executed when the user selects content to be displayed, and the client terminal 10 displays an initial image on the display device 16 while transmitting viewpoint information and user operation details to the content server 20. Note that although the figure shows each processing step being performed in order, any processing step may be performed in parallel with others.
[0078] In this state, the drawing data acquisition unit 52 of the client terminal 10 acquires drawing data from the content server 20 and decodes and expands it as appropriate (S30). The image generation unit 54 acquires information about the current display viewpoint via the input information acquisition unit 50 (S32). The image generation unit 54 then sets an object in three-dimensional space based on the geometry data included in the drawing data, and sets a viewscreen corresponding to the display viewpoint (S34). Next, the image generation unit 54 acquires pixel values representing the image of the object for each pixel on the viewscreen (S36).
[0079] In particular, for objects that exhibit viewpoint-dependent effects, the image generation unit 54 acquires pixel values using a texture model in the procedure shown in Fig. 9. For objects that do not exhibit viewpoint-dependent effects, the image generation unit 54 acquires pixel values by sampling the associated texture. The image generation unit 54 repeats the process of S36 until it has acquired pixel values for all pixels on the viewscreen and the image is complete (N in S38).
[0080] When the image is completed (Y in S38), the output unit 56 outputs the image to the display device 16 as one frame of the display image (S40). If there is no need to end the image display, such as when the content has ended, the client terminal 10 repeats the processes from S30 to S40 (N in S42). If it becomes necessary to end the image display, the client terminal 10 ends all processes (Y in S42).
[0081] According to the present embodiment described above, in image processing of electronic content, the content server 20 prepares a texture model for each object that derives viewpoint-dependent effects, i.e., viewpoint-dependent color changes, and transmits this to the client terminal 10 along with the geometry data. This makes it possible to generate an image that represents the displayed world with a field of view and colors that correspond to the viewpoint immediately before display, enabling high-quality image representation with little delay.
[0082] The content server 20 utilizes its abundant resources to generate a texture model by performing machine learning on images viewed from multiple training viewpoints with high resolution. Compared to the general NeRF, which takes into account the density of three-dimensional objects, the texture model allows for a lower order of input and output parameters, thereby reducing the processing load for training and rendering. Furthermore, because the content server 20 uses images drawn from a common training viewpoint for training, the amount of information in the resulting texture model increases the closer an object is to the display viewpoint. This trend is consistent with the level of detail of objects required for display, allowing texture models to be generated with a necessary and sufficient amount of processing.
[0083] Furthermore, for objects that do not require viewpoint-dependent effects, texture models are not generated, and instead, textures drawn from the display viewpoint are associated with them, minimizing the processing load for generating texture models and using them to obtain color values. Furthermore, because texture models allow for highly accurate representation of object colors and contours, the impact on appearance can be minimized even if polygons are made coarse, resulting in reduced data transfer volume. These features enable high-quality images to be displayed with high responsiveness to changes in viewpoint and line of sight.
[0084] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention.
[0085] For example, in this embodiment, the content server 20 generates new learning textures for the required viewpoint based on the latest display viewpoint transmitted from the client terminal 10. On the other hand, the content server 20 may reuse textures that have been generated in the past or textures generated for other client terminals 10 as at least a part of the learning textures. The smaller the movement of objects in the display world and the smaller the change in the display viewpoint, the smaller the change in the texture of each object. Therefore, even if a texture model generated by reusing a past texture is used, a display image that has little impact on appearance can be generated.
[0086] This reduces the processing load of texture model generation on the content server 20, enabling display with lower latency. Whether or not to reuse textures generated in the past, the proportion to reuse, and the allowable range of generation times for the reused textures may be adaptively determined depending on the movement of the object and the magnitude of changes in the display viewpoint. If the object is not moving, the content server 20 can first generate training textures from a training viewpoint that covers the movable range of the display viewpoint and generate a texture model, and then omit the generation of training textures and the machine learning process itself.
[0087] As described above, the present invention can be used in various information processing devices such as content servers, game devices, head-mounted displays, display devices, mobile terminals, and personal computers, as well as image display systems including any of these.
[0088] 1 Image display system, 10 Client terminal, 14 Input device, 16 Display device, 20 Content server, 22 CPU, 24 GPU, 26 Main memory, 50 Input information acquisition unit, 52 Drawing data acquisition unit, 54 Image generation unit, 56 Output unit, 70 Input information acquisition unit, 72 Learning viewpoint generation unit, 74 Display world control unit, 76 Three-dimensional model storage unit, 78 Target object selection unit, 80 Learning texture generation unit, 82 Texture model generation unit, 84 Drawing data formation unit, 86 Drawing data transmission unit.
Claims
1. A content server comprising: an input information acquisition unit that acquires, from a client terminal that generates a display image representing a three-dimensional display world, information on a display viewpoint that defines the display image; a learning texture generation unit that generates a plurality of learning textures that represent the distribution of color values on the surface of an object by drawing an image of the object as seen from a plurality of learning viewpoints generated based on the display viewpoints; a texture model generation unit that generates a texture model that outputs color values corresponding to changes in viewpoint through machine learning using the learning textures as training data; and a drawing data transmission unit that associates geometry data of the display world with the texture model and transmits them to the client terminal.
2. A content server as described in claim 1, characterized in that the texture model generation unit generates, for each object, a neural network as the texture model, which outputs color values based on the UV coordinates of the object surface and the line of sight direction.
3. A content server as described in claim 1 or 2, characterized in that the learning texture generation unit generates, for each learning viewpoint, the learning texture representing the distribution of color values of the area of the object surface that is visible from that learning viewpoint.
4. A content server as described in claim 1 or 2, characterized in that the learning texture generation unit generates the learning texture for objects selected based on a predetermined criterion from among objects present in the display world.
5. A content server according to claim 4, characterized in that said learning texture generating section generates said learning texture for an object existing within a predetermined range from said display viewpoint.
6. A content server as described in claim 4, further comprising a drawing data forming unit which generates a texture representing the distribution of color values on the surface of an object by drawing an image of an object other than the selected object as seen from the display viewpoint, and wherein the drawing data transmitting unit further associates the texture with the geometry data and transmits the texture to the client terminal.
7. The contents server according to claim 6, wherein said drawing data forming section sets the polygon density of said selected object to be smaller than that of other objects.
8. A content server as described in claim 1 or 2, characterized in that the input information acquisition unit acquires the content of user operations from the client terminal, and further includes a display world control unit that changes the display world based on the content of the user operations, and the drawing data transmission unit transmits the texture model and the geometry data, which are updated in accordance with changes in the display world, at a predetermined rate.
9. A client terminal comprising: an input information acquisition unit that acquires information on a display viewpoint that defines a display image representing a three-dimensional display world; a drawing data acquisition unit that acquires from a server a texture model that outputs color values corresponding to changes in viewpoint, the texture model being generated by machine learning using a learning texture that represents the distribution of color values of the surface of an object viewed from a plurality of learning viewpoints, together with geometry data of the display world; and an image generation unit that generates a display image by acquiring color values corresponding to the latest display viewpoint using the texture model, and outputs the display image to a display device.
10. An image display system comprising: a client terminal which generates and displays a display image representing a three-dimensional display world; and a content server which transmits drawing data used to generate the display image, wherein the content server comprises: an input information acquisition unit which acquires, from the client terminal, information on a display viewpoint which defines the display image; a learning texture generation unit which generates a plurality of learning textures which represent a distribution of color values on the surface of an object by drawing an image of the object as seen from a plurality of learning viewpoints which are generated based on the display viewpoints; a texture model generation unit which generates a texture model which outputs color values corresponding to changes in viewpoint by machine learning using the learning textures as teacher data; and a drawing data transmission unit which associates geometry data of the display world with the texture model and transmits the corresponding data to the client terminal, wherein the client terminal comprises: an input information acquisition unit which acquires information on the display viewpoint; a drawing data acquisition unit which acquires the texture model and the geometry data; and an image generation unit which generates a display image by acquiring color values corresponding to the latest display viewpoint using the texture model, and outputs the display image to a display device.
11. A drawing data transmission method comprising the steps of: acquiring, from a client terminal that generates a display image representing a three-dimensional display world, information on a display viewpoint that defines the display image; generating a plurality of learning textures that represent the distribution of color values on the surface of the object by drawing images of the object as seen from a plurality of learning viewpoints generated based on the display viewpoints; generating a texture model that outputs color values corresponding to changes in viewpoint by machine learning using the learning textures as training data; and transmitting the geometry data of the display world and the texture model to the client terminal in association with each other.
12. A display image generating method comprising the steps of: acquiring information on a display viewpoint that defines a display image representing a three-dimensional display world; acquiring from a server, together with geometry data of the display world, a texture model that outputs color values corresponding to changes in viewpoint, the texture model being generated by machine learning using a learning texture that represents the distribution of color values of the surface of an object viewed from a plurality of learning viewpoints; and generating a display image by acquiring color values corresponding to the latest display viewpoint using the texture model, and outputting the display image to a display device.
13. A computer program that causes a computer to realize the following functions: a function to obtain, from a client terminal that generates a display image representing a three-dimensional display world, information on a display viewpoint that defines the display image; a function to generate a plurality of learning textures that represent the distribution of color values on the surface of an object by drawing images of the object as seen from a plurality of learning viewpoints generated based on the display viewpoints; a function to generate a texture model that outputs color values corresponding to changes in viewpoint through machine learning using the learning textures as training data; and a function to associate geometry data of the display world with the texture model and transmit the result to the client terminal.
14. A computer program that causes a computer to realize the following functions: a function to acquire information on a display viewpoint that defines a display image that represents a three-dimensional display world; a function to acquire from a server a texture model that outputs color values corresponding to changes in viewpoint, the texture model being generated by machine learning using a learning texture that represents the distribution of color values of the surface of an object as viewed from multiple learning viewpoints, together with geometry data of the display world; and a function to generate a display image by acquiring color values corresponding to the latest display viewpoint using the texture model, and output the display image to a display device.
Citation Information
Patent Citations
Information processing apparatus, information processing program, and information processing method
JP2020013390A
Image rendering method and apparatus
JP2022151746A
Artificial Intelligence (AI) controlled camera perspective generator and AI broadcaster
JP2022545128A