Image processing device, image processing method, and data structure of 3D scene information for display

The image processing device and method address the challenge of applying machine learning to dynamic content by generating learning images and creating 3D scene information, while controlling viewpoints to ensure accurate and controlled display.

WO2025094267A1PCT designated stage expired Publication Date: 2025-05-08SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/039247
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing image processing technologies struggle to apply machine learning effectively to content that changes dynamically in response to user operations, particularly in real-time scenarios, due to challenges in acquiring and utilizing learning images appropriately.

Method used

An image processing device and method that generate learning images representing a dynamic display world, using these images to create 3D scene information through machine learning, and restrict viewpoints to ensure proper display control and adherence to intended viewing angles.

Benefits of technology

Enables the application of machine learning to dynamic content, allowing for controlled and accurate viewpoint management, thereby enhancing the quality and integrity of displayed 3D scene information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023039247_08052025_PF_FP_ABST
    Figure JP2023039247_08052025_PF_FP_ABST
Patent Text Reader

Abstract

An additional viewpoint setting unit 110 of a content server 20 acquires viewpoint restriction information from an application execution unit 74 and sets viewpoints for generating training images within a restricted range. The application execution unit 74 generates images of a display world corresponding to the set viewpoints. A 3D scene information generation unit 76 performs machine learning using the images generated by the application execution unit 74, and generates 3D scene information about the display world. A free-viewpoint image generation unit 114 acquires the viewpoint restriction information from the application execution unit 74 and generates an image representing the display world from a free viewpoint within the restricted range.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method, and data structure of 3D scene information for display

[0001] The present invention relates to an image processing device and an image processing method for processing images of content that reflect user operations, and a data structure of 3D scene information for display.

[0002] With the recent expansion of communication networks and advances in image processing technology, it has become possible to enjoy a wide variety of electronic content regardless of the viewing environment. For example, in the field of electronic games, a system has become widespread in which a server collects information related to the status of each client device, such as the content of user operations and location information, and distributes image data that reflects this information as needed, allowing multiple players to participate in the same game regardless of location.

[0003] Meanwhile, in recent years, advances in machine learning technologies such as deep learning have made it possible to acquire various types of information from images. For example, NeRF (Neural Radiance Fields) is a method for representing three-dimensional space using a neural network. NeRF is a method for representing the volume density and radiance of an object in three-dimensional space as a five-dimensional function consisting of position coordinates and direction using a neural network. For example, if an NeRF representation is obtained based on images of an object captured from multiple directions, it is possible to represent the appearance of the object as seen from any viewpoint using volume rendering (see, for example, Non-Patent Document 1).

[0004] Ben Mildenhall and five others, "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis," Communications of the ACM, January 2022, Vol. 65, No. 1, pp. 99-106

[0005] While image processing using machine learning as described above makes it possible to obtain highly flexible images from limited information, it requires learning using appropriate and sufficient images, which limits the scope of application. For example, in the case of content in which the displayed scene changes in real time in response to user operations, it is difficult to implement due to issues such as when to acquire learning images for the constantly changing scene and how to use the learned information. If easy implementation increases the freedom of the viewpoint of the displayed world, there is a risk that the displayed world will be exposed from an unintended angle of view.

[0006] The present invention was made in consideration of these problems, and its purpose is to provide a technology that applies machine learning to content in which the state of the displayed world can change in response to user operations, and that can appropriately control the viewpoint when displaying using the obtained three-dimensional information.

[0007] To achieve the above object, one aspect of the present invention relates to an image processing device, comprising: an application execution unit that executes an application program and generates, at a predetermined rate, display image frames representing a three-dimensional display world whose situation changes in response to user operations; and a system unit that causes the application execution unit to generate training images that represent the display world and are different from the display images, performs machine learning using the training images as training data to generate 3D scene information representing three-dimensional information about the display world and use it for display, and in this process restricts a viewpoint set for the display world based on viewpoint restriction information associated with the application program.

[0008] Another aspect of the present invention also relates to an image processing device, comprising: a 3D scene information storage unit that stores 3D scene information formed by a neural network that represents three-dimensional information of a display world, and viewpoint restriction information that corresponds to the 3D scene information, and an arbitrary viewpoint image generation unit that reads the 3D scene information and the viewpoint restriction information from the 3D scene information storage unit and generates, by volume rendering using the 3D scene information, a display image that represents a state of the display world as viewed from an arbitrary viewpoint within a range of restrictions indicated by the viewpoint restriction information.

[0009] Yet another aspect of the present invention relates to an image processing method, the image processing method including the steps of: an application execution unit executing an application program and generating, at a predetermined rate, frames of a display image representing a three-dimensional display world whose situation changes in response to a user operation; and a system unit causing the application execution unit to generate learning images that represent the display world and are different from the display images, and performing a process of generating 3D scene information representing three-dimensional information of the display world by machine learning using the learning images as training data, and using the 3D scene information for display, and limiting a viewpoint set for the display world based on viewpoint restriction information associated with the application program.

[0010] Yet another aspect of the present invention relates to a data structure of 3D scene information for display, characterized in that the data structure of 3D scene information for display associates data of 3D scene information formed by a neural network representing three-dimensional information of a display world with viewpoint restriction information that indicates restriction information to be imposed on an arbitrary viewpoint when a display image representing a state of the display world as viewed from the arbitrary viewpoint is generated by volume rendering using the 3D scene information, the data being read from a storage device together with the 3D scene information by an image processing device.

[0011] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention.

[0012] According to the present invention, machine learning is applied to content in which the state of the displayed world can change in response to user operations, and the viewpoint can be appropriately controlled when displaying using the obtained three-dimensional information.

[0013] 1 is a diagram showing an example of the configuration of an image display system to which the present embodiment can be applied. FIG. 2 is a diagram showing the internal circuit configuration of a client terminal in the present embodiment. FIG. 3 is a diagram showing a basic flow of image processing in the present embodiment in comparison with the conventional technology. FIG. 4 is a diagram showing an overview of the processing flow in an aspect in which a user saves a desired scene as 3D scene information. FIG. 5 is a diagram showing the configuration of functional blocks of a client terminal and a content server that realize scene saving in the present embodiment. FIG. 6 is a diagram showing a schematic sequence of images generated in a main image output phase of the present embodiment. FIG. 7 is a diagram showing an example of the arrangement of pseudo viewpoints generated by a pseudo viewpoint generation unit in the present embodiment. FIG. 8 is a diagram showing a schematic diagram of switching between a main image and a standby image displayed on a display device in the present embodiment. FIG. 9 is a diagram for explaining an aspect in which a 3D scene information generation unit extracts an area to be used for learning from a learning image in the present embodiment. FIG. 10 is a diagram showing an overview of the processing flow in an aspect in which 3D scene information is used to correct a display image. FIG. 11 is a diagram for explaining reprojection as an example of correction of a main image in the present embodiment. FIG. 12 is a diagram showing the configuration of functional blocks of a client terminal and a content server that realize correction of a display image in the present embodiment. FIG. 13 is a diagram showing a schematic sequence of images generated in the present embodiment. 1 is a diagram for explaining an aspect in which an image is generated by shifting a display viewpoint when images for the left eye and the right eye are displayed on a head-mounted display in this embodiment. FIG. 2 is a diagram illustrating an overview of a processing flow in an aspect in which 3D scene information is used to distribute replay images. FIG. 3 is a diagram illustrating the functional block configuration of a client terminal and a content server that realizes the distribution of replay videos in this embodiment. FIG. 4 is a diagram schematically illustrating a sequence of images generated in the main image output phase of this embodiment. FIG. 5 is a diagram illustrating an example of a screen that an additional viewpoint setting unit of a content server in this embodiment displays to accept the setting of an additional viewpoint by a user. FIG. 6 is a diagram illustrating an example of a heat map generated by a heat map generation unit of a content server in this embodiment. FIG. 7 is a diagram illustrating an example of a display screen for replay images that is displayed on a display device in the replay image distribution phase of this embodiment.10 is a diagram illustrating a configuration of functional blocks of a content server in a mode in which a display viewpoint is restricted by an application.FIG. 11 is a diagram illustrating an example of a data structure of 3D scene information in this embodiment.FIG.

[0014] 1 shows an example of the configuration of an image display system to which this embodiment can be applied. The image processing system 1 includes client terminals 10a, 10b, and 10c that display images in response to user operations, etc., and a content server 20 that provides image data used for display. Input devices 14a, 14b, and 14c for user operations and display devices 16a, 16b, and 16c for displaying images are connected to the client terminals 10a, 10b, and 10c, respectively. Communication can be established between the client terminals 10a, 10b, and 10c and the content server 20 via a network 8 such as a WAN (World Area Network) or a LAN (Local Area Network).

[0015] The client terminals 10a, 10b, and 10c may be connected to the display devices 16a, 16b, and 16c and the input devices 14a, 14b, and 14c either wired or wirelessly. Alternatively, two or more of these devices may be integrated. For example, in the figure, the client terminal 10b is connected to a head-mounted display, which is the display device 16b. The head-mounted display can change the field of view of the displayed image by the movement of the user wearing it on their head, so it also functions as the input device 14b.

[0016] The client terminal 10c is a mobile terminal, tablet terminal, or the like, and is integrally configured with a display device 16c and an input device 14c, which is a touchpad covering the display device 16c's screen. Thus, the external shape and connection form of the illustrated devices are not limited. The number of client terminals 10a, 10b, and 10c and content servers 20 connected to the network 8 is also not limited. Hereinafter, the client terminals 10a, 10b, and 10c will be collectively referred to as client terminals 10, the input devices 14a, 14b, and 14c as input device 14, and the display devices 16a, 16b, and 16c as display device 16.

[0017] The input device 14 is a general input device such as a controller, keyboard, mouse, touchpad, or joystick, and receives user operations and supplies the operations to the client terminal 10. The input device 14 may also be various sensors such as a motion sensor or camera provided in a head-mounted display, mobile terminal, or tablet terminal, and may supply the sensor data to the client terminal 10. The display device 16 may be a general display such as a liquid crystal display, plasma display, organic EL display, wearable display, or projector, and displays an image output from the client terminal 10.

[0018] The content server 20 provides data of content accompanied by image display to the client terminal 10. The type of content is not particularly limited, and may be any of electronic games, decorative images, promotional images, web pages, video chat using avatars, etc. In this embodiment, the content server 20 basically generates moving image and audio data representing the content and immediately transmits the data to the client terminal 10 to realize streaming.

[0019] In this case, the content server 20 may sequentially acquire information on user operations on the input device 14 or sensor data acquired by various sensors from the client terminal 10 and reflect the information in images and sounds. This allows multiple users to participate in the same game or communicate in the virtual world. However, the configuration of the image processing system is not limited to that shown in the figure. For example, the image generation entity is not limited to the content server 20, but may be performed by the client terminal 10 itself, or the two may work together.

[0020] 2 shows the internal circuit configuration of the client terminal 10. The client terminal 10 includes a CPU (Central Processing Unit) 122, a GPU (Graphics Processing Unit) 124, and a main memory 126. These components are interconnected via a bus 130. An input / output interface 128 is also connected to the bus 130. Connected to the input / output interface 128 are a communication unit 132 including a peripheral device interface such as a USB or a network interface for a wired or wireless LAN, a storage unit 134 such as a hard disk drive or nonvolatile memory, an output unit 136 that outputs data to the display device 16, an input unit 138 that inputs data from the input device 14, and a recording medium drive unit 140 that drives a removable recording medium such as a magnetic disk, optical disk, or semiconductor memory.

[0021] The CPU 122 executes an operating system stored in the storage unit 134 to control the entire client terminal 10. The CPU 122 also executes various programs read from a removable recording medium and loaded into the main memory 126, or downloaded via the communication unit 132. The GPU 124 has the functions of a geometry engine and a rendering processor, performs drawing processing in accordance with drawing commands from the CPU 122, and stores display images in a frame buffer (not shown). The GPU 124 then converts the display images stored in the frame buffer into video signals and outputs them to the output unit 136. The main memory 126 is composed of RAM (Random Access Memory) and stores programs and data required for processing. The content server 20 may also have a similar internal circuit configuration.

[0022] 3 shows the basic flow of image processing in this embodiment in comparison with the prior art. As mentioned above, the main processing can be performed by either the content server 20 or the client terminal 10, or by both working together, so here we will not distinguish between them and will describe the processing as being performed by an "image processing device." In this embodiment, the main display target is a three-dimensional world containing various objects. The state of this world changes depending on the program specifications and user operations.

[0023] In the case of the general processing shown in (a), the image processing device first acquires information on the content of user operations, the position of the viewpoint relative to the displayed world, and the direction of the line of sight as needed. Hereinafter, the entire three-dimensional space of the display target will be referred to as the "display world," and the state of the displayed world within or near the display field of view will be referred to as the "scene." The position of the viewpoint relative to the scene and the direction of the line of sight will sometimes be collectively referred to simply as the "viewpoint." The viewpoint may be manually operated by the user via the input device 14, or may be derived from the movement of the user's head using a motion sensor or the like provided in the head-mounted display.

[0024] The image processing device draws a display image 200 in a field of view corresponding to viewpoint information while changing the scene in response to user operations. The image processing device generates the display image 200 using well-known computer graphics drawing techniques such as ray tracing and rasterization, and outputs it to the display device 16. The image processing device continues to generate the display image 200 at a predetermined frame rate, thereby displaying a moving image that shows changes in the scene in response to user operations, etc. In other words, the display image 200 is a frame of a moving image that can be interactively changed based on user operations and viewpoint information.

[0025] Hereinafter, a moving image generated in parallel with the acquisition of user operations and viewpoint information will be referred to as a "main image." A typical example of a main image is a game image during play. The image processing device may acquire the contents of user operations in parallel from multiple users, as in a multiplayer game, and reflect them in the display image 200. In this embodiment shown in (b), the image processing device also generates a main image in a similar manner. Meanwhile, in this embodiment, the image processing device uses the main image as a training image 202 and as training data for machine learning. The image processing device collects the training images 202 and performs machine learning to generate 3D scene information 204 representing three-dimensional information about a scene.

[0026] When applying NeRF to machine learning, first, the viewpoint information determined when generating the training images 202, i.e., the virtual viewpoint position and gaze direction, is input, and the corresponding training images 202 are used as training data to obtain data representing three-dimensional information of the scene through regression using a multilayer perceptron (MLP). This data is a neural network that inputs five-dimensional parameters consisting of position coordinates (x, y, z) in three-dimensional space and a direction vector d(θ, φ), and outputs volume density σ and color information c (RGB) of the three primary colors.

[0027] In this embodiment, the neural network data is referred to as "3D scene information." However, any technology that can estimate 3D information from multiple 2D images can be introduced, not just NeRF, and the representation format of the 3D scene information is not limited. In this embodiment, the training image 202 is the main image. In other words, the content represented by the training image 202, and therefore the 3D scene information 204, can change from moment to moment. The figure shows the situation in which 3D scene information 204 of a scene is generated at a certain time, or at a very short time that can be considered as a single time.

[0028] To obtain accurate 3D scene information 204, it is desirable for the image processing device to collect learning images 202 of a scene at a single time, or at a very small time that can be considered as a single time, from as many viewpoints as possible. For this reason, the image processing device collects learning images 202, for example, in the following manner: (1) In addition to the viewpoint that defines the field of view of the image that is actually displayed, the image processing device generates a viewpoint suitable for learning and generates a corresponding image. (2) The image processing device reuses display images from various viewpoints that are distributed to the terminals of multiple users viewing the same scene.

[0029] Hereinafter, the viewpoint generated by the image processing device itself in (a) will be referred to as the "pseudo viewpoint," and the viewpoint that defines the actual display will be referred to as the "display viewpoint." The image processing device may implement either (1) or (2), or both. For example, a viewpoint that is missing in (2) may be supplemented by (1). In either case, the training image 202 may include a general display image 200 as shown in (a) of the figure. Therefore, the image processing device may output at least a portion of the training image 202 to the display device 16 as a display image.

[0030] On the other hand, the image processing device may use the 3D scene information 204 to separately generate a display image 206 or to correct the display image. By using the 3D scene information 204, it is possible to represent the scene as seen from any viewpoint with high quality, with a relatively low load. When NeRF is applied, the image processing device generates a ray r that passes through a pixel on the view screen from the display viewpoint, and calculates the pixel value C(r) of the display image by volume rendering that integrates color along that direction, as follows:

[0031]

[0032] where t n , t f are the proximal and distal ends of the ray r, respectively, and T(t) is the cumulative transmittance in the direction of the ray, which can be expressed as follows:

[0033]

[0034] Regarding NeRF, in addition to the basic technique disclosed in, for example, Non-Patent Document 1, various improved techniques have been proposed, and any of these may be adopted in this embodiment. Therefore, detailed description will be omitted here. The image processing device may generate a single piece of 3D scene information 204 representing a scene at a single time or a very short time, or may continuously update the 3D scene information 204 at a predetermined rate by repeating the illustrated process. In the former case, the image processing device can use the 3D scene information 204 to represent a scene captured at a moment in time from the main image from any viewpoint. In the latter case, the chronological order is also preserved in the 3D scene information group. Therefore, by applying the relevant time to the 3D scene information used to generate the display image 206, the image processing device can represent a moving image having the same changes as the main image from any viewpoint.

[0035] For example, the image processing device displays the 3D scene information 204 in accordance with a user request at a timing different from the display period of the main image, such as after the end of a game, and accepts an operation of the display viewpoint from the user. This provides, for example, a function that allows a user to view a fleeting scene saved as 3D scene information 204 during game play from various angles after play ends, or to share it with other users. The image processing device can also provide a function for distributing replay videos that can be viewed from any viewpoint.

[0036] When the 3D scene information 204 is continuously updated at a predetermined rate, the image processing device may use the 3D scene information 204 for correction when displaying a main image. For example, in a mode in which streaming images are viewed on a head-mounted display, the image processing device uses the 3D scene information 204 to correct the image to match the position and orientation of the user's head immediately before display. Examples of modes that can be realized in this embodiment will be described below. For ease of understanding, each mode will be described individually, but in practice, multiple modes may be combined and implemented.

[0037] 2. Saving a Scene Figure 4 shows an overview of the processing flow in a mode in which a user saves a desired scene as 3D scene information. This mode is realized in two periods: a main image output phase 210 and a saved scene viewing phase 212. The main image output phase 210 is a period in which the main image of the content is output, such as during game play. During this period, an image processing device, for example, the content server 20, accepts a user operation to save the scene (S10).

[0038] In response to this, the content server 20 generates training images that represent the scene at the time the user operation was performed from multiple viewpoints (S12), and generates 3D scene information 220 representing the scene by performing machine learning (S14). In practice, the generation of training images and learning using the training images may be performed in parallel. The saved scene viewing phase 212 begins when the user requests viewing, at any time, such as after the end of game play. During this period, an image processing device, for example, the content server 20, generates an image of the scene using the saved 3D scene information 220 and outputs it for display (S16).

[0039] Alternatively, the content server 20 performs processing to share the saved scene with other users in response to a user request (S18). For example, the content server 20 uses an existing SNS (Social Networking Service) mechanism to send and display an image of the scene to the client terminal 10 of another user specified by the sharing user. In either case, the content server 20 uses the 3D scene information 220 to generate a display image of the scene while changing the display viewpoint in response to viewpoint manipulation by the user viewing the image.

[0040] Figure 5 shows the functional block configuration of the client terminal 10 and content server 20 that realizes scene storage. The functional blocks shown in Figure 5 and Figures 12, 16, 21, and 23 described below can be realized in hardware terms using the CPU, GPU, various memories, and other components shown in Figure 2, and in software terms using programs loaded into memory from a recording medium or the like that perform various functions such as data input, data storage, image processing, and communication. Therefore, those skilled in the art will understand that these functional blocks can be realized in various forms using only hardware, only software, or a combination thereof, and are not limited to any one of these. In the following description, the content server 20 is primarily responsible for image processing, but at least some of this processing may also be performed by the client terminal 10.

[0041] The client terminal 10 includes an input information acquisition unit 50 that acquires input information such as user operations, an image data acquisition unit 52 that acquires image data from the content server 20, and an output unit 54 that outputs display image data. The input information acquisition unit 50 acquires the content of user operations from the input device 14 at any time. User operations include content selection and launch, and command input for content currently being played. The input information acquisition unit 50 also accepts operations to save a desired scene from the main image of the content, and operations to request viewing of a saved scene or sharing with other users. In this embodiment, the operation to save a scene is sufficient if the timing is specified. Therefore, it is preferably realized by a simple operation such as pressing a button on the input device 14.

[0042] The input information acquisition unit 50 also acquires information on the display viewpoint from the input device 14 or the head-mounted display at any time or at predetermined time intervals. Technology for detecting the position and posture of the head of a user wearing a head-mounted display and acquiring information on the display viewpoint based on the detected position and posture is well known, and this technology may also be applied to this embodiment. The display viewpoint here includes the display viewpoint for viewing a saved scene as well as the display viewpoint for the main image. The input information acquisition unit 50 supplies the acquired information to the content server 20 as appropriate.

[0043] The image data acquisition unit 52 acquires display image data from the content server 20. Here, the display image data may include data of the main image, data of the image of the saved scene, and data of the standby image used during the period in which the scene to be saved is learned. The output unit 54 outputs the display image acquired by the image data acquisition unit 52 to the display device 16 for display.

[0044] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, a pseudo viewpoint generation unit 72 that generates a pseudo viewpoint for generating learning images, an application execution unit 74 that executes applications such as electronic games, a 3D scene information generation unit 76 that generates data of 3D scene information, a 3D scene information storage unit 78 that stores the generated data of the 3D scene information, a standby image generation unit 80 that generates a standby image indicating a learning image generation period, a saved scene image generation unit 81 that generates an image representing a saved scene, and an image data transmission unit 82 that transmits data of the display image to the client terminal 10.

[0045] The input information acquisition unit 70 acquires the content of a user operation and information on a display viewpoint from the client terminal 10 at any time or at predetermined time intervals. The input information acquisition unit 70 basically supplies the acquired information to the application execution unit 74. When a user operation to save a scene is acquired, the input information acquisition unit 70 also supplies the information and the latest display viewpoint information to the pseudo viewpoint generation unit 72. At this time, the pseudo viewpoint generation unit 72 generates a pseudo viewpoint for generating a learning image based on the latest display viewpoint. The pseudo viewpoint generation unit 72 supplies the generated pseudo viewpoint information to the application execution unit 74.

[0046] In the main image output phase, the application execution unit 74 processes an application for content such as an electronic game based on the content of a user operation. The application execution unit 74 includes a main image generation unit 84, which generates main image frames corresponding to a display viewpoint at a predetermined rate. When a user operation to save a scene is performed, the main image generation unit 84 generates, as a learning image, an image representing the scene as seen from the pseudo viewpoint generated by the pseudo viewpoint generation unit 72.

[0047] In the illustrated example, it is assumed that the application execution unit 74 generates a main image based on viewpoint information supplied from the input information acquisition unit 70. In this case, the pseudo viewpoint generation unit 72 generates pseudo viewpoint information in the same format as the viewpoint information supplied by the input information acquisition unit 70 and supplies the pseudo viewpoint information to the application execution unit 74, thereby allowing the application execution unit 74 to generate learning images through normal processing without distinguishing between a true display viewpoint and a pseudo viewpoint. As a result, this embodiment can be easily implemented even in conventional content that does not support machine learning.

[0048] However, the present embodiment is not limited to this. An API (Application Programming Interface) having a function for generating a pseudo viewpoint may be prepared and specified in the application program, so that the application execution unit 74 includes the pseudo viewpoint generation unit 72. In any case, it is desirable for the application execution unit 74 to pause the progress of the content until the main image generation unit 84 generates a sufficient number of learning images. This allows the scene at the time the user performs the save operation to be treated as a static scene, allowing a sufficient number of learning images to be generated and 3D scene information to be generated with high accuracy.

[0049] If the progress of the content is paused, the application execution unit 74 resumes the progress of the content when all images corresponding to the pseudo viewpoints have been generated. In the main image output phase, the 3D scene information generation unit 76 acquires the learning images generated by the application execution unit 74 and generates 3D scene information of the scene to be saved by machine learning as described above. Note that the 3D scene information generation unit 76 may extract only the area to be saved from the learning images generated by the main image generation unit 84 and use it for machine learning.

[0050] The 3D scene information storage unit 78 stores the 3D scene information generated by the 3D scene information generation unit 76. The 3D scene information storage unit 78 stores the 3D scene information in association with information such as the identification information of the user who requested the scene to be saved and the timing of the save on the time axis of the main image. This facilitates searching for the scene to be displayed in the saved scene viewing phase. When a user operation to save a scene is performed in the main image output phase, the standby image generation unit 80 generates a standby image to be displayed during the image learning period. Displaying the standby image allows the user to recognize that the scene saving is progressing. Furthermore, if the display device 16 is a head-mounted display, this can reduce sickness caused by pausing the scene and the field of view no longer following head movement.

[0051] In the saved scene viewing phase, when a user operation is performed to request viewing of a saved scene, the saved scene image generation unit 81 generates a display image representing the scene by the above-mentioned volume rendering using the 3D scene information stored in the 3D scene information storage unit 78. At this time, the saved scene image generation unit 81 acquires a display viewpoint from the input information acquisition unit 70 and generates a display image while changing the viewpoint for the saved scene accordingly. In the main image output phase, the image data transmission unit 82 sequentially transmits data of the main image generated by the main image generation unit 84 and the standby image generated by the standby image generation unit 80 to the client terminal 10.

[0052] The image data transmission unit 82 also transmits data of the image of the saved scene generated by the saved scene image generation unit 81 in the saved scene viewing phase to the client terminal 10. When a user operation to share the saved scene with other users is performed, the image data transmission unit 82 transmits data of the image of the saved scene to the shared client terminal 10. In this case, since a general SNS platform can actually be used, detailed functional blocks are omitted in the figure.

[0053] 6 shows a schematic diagram of a sequence of images generated in the main image output phase of this embodiment. The diagram shows the relationship between the viewpoint recognized or generated by the content server 20 and the destination of each frame generated thereby, with the horizontal axis representing time. The content server 20 basically generates display image frames (e.g., frame 232) at a predetermined rate so as to correspond to the display viewpoint (e.g., display viewpoint 230) indicated by the white circle, and transmits them to the client terminal 10.

[0054] As a result, when a scene that the user wants to save appears in the main image displayed on the client terminal 10, the user performs a scene save operation, for example, by pressing a predetermined button provided on the input device 14. In the figure, when the save operation is performed at time t1, the content server 20 generates a pseudo viewpoint (e.g., pseudo viewpoint 234) indicated by a black circle and generates a corresponding training image (e.g., training image 236). The content server 20 temporarily suspends the generation of display image frames while the training image is being generated. As shown in the figure, the rate at which training images are generated may be higher than the rate at which display frames are generated, depending on the processing capacity of the content server 20.

[0055] As an example, if the rendering processing capability of the main image generation unit 84 is 120 fps, 120 pseudo viewpoints can be prepared and processed sequentially by the main image generation unit 84, thereby generating 120 training images per second. The content server 20 pauses the progress of the content while generating the training images, generates a standby image (e.g., standby image 238) indicated by the hatching, and transmits it to the client terminal 10. As described above, the standby image may be a still image or a video. The standby image may also be generated on the client terminal 10 side. The display of the standby image continues until time t2, when the content server 20 has finished generating the predetermined number of training images. As described above, in an environment where 120 training images can be generated per second, the display time of the standby image may be on the order of a few seconds.

[0056] The content server 20 generates 3D scene information of the scene based on the learning images generated up to time t2, and stores the information in the 3D scene information storage unit 78. The content server 20 resumes the progress of the content at time t2, and generates frames of display images at a predetermined rate so as to correspond to the latest display viewpoint, and transmits the frames to the client terminal 10.

[0057] 7 illustrates an example of the arrangement of pseudo viewpoints generated by the pseudo viewpoint generator 72. In this example, multiple pseudo viewpoints (e.g., viewpoint 242) are arranged to surround a scene, such as an object 240, that is included in the display field of view at the time the scene is saved. For example, the pseudo viewpoint generator 72 arranges pseudo viewpoints at regular intervals on the surface of a sphere 244 of a predetermined radius centered at a position in the scene that corresponds to the center of the display field of view. Then, a line of sight from each pseudo viewpoint toward the center of the sphere 244 is set.

[0058] This allows for the generation of learning images that represent the scene the user was viewing at the time of the save operation from various directions. However, the arrangement of the pseudo viewpoints is not limited to that shown in the figure. For example, if the scene includes the ground, a hemisphere may be used instead of the sphere 244 so that only the ground is valid. Furthermore, the surface on which the viewpoints are placed is not limited to a spherical surface; it may be the surface of a rectangular parallelepiped, cylinder, ellipsoid, or the like, and in some cases, it may not even be a specific three-dimensional surface. Furthermore, the viewpoints are not limited to being evenly distributed; a biased distribution may be provided, such as by placing more viewpoints in an area where the display viewpoint is likely to be located during the saved scene viewing phase or in an area where important objects are visible. This allows for the efficient generation of 3D scene information with high accuracy for important areas of the scene.

[0059] The pseudo viewpoint generator 72 may also set pseudo viewpoints on multiple three-dimensional surfaces. For example, the pseudo viewpoint generator 72 may place pseudo viewpoints on the surfaces of concentric spheres of different sizes. This allows for the generation of learning images that represent scenes viewed from various distances. The direction of the line of sight is not limited to the center of the scene. For example, the pseudo viewpoint generator 72 may set lines of sight radially from the position of a virtual user in the scene as the starting point.

[0060] This allows for the generation of 3D scene information that can accommodate large rotations of the display field of view during the saved scene viewing phase. In any case, the more pseudo viewpoints are added, the more accurate the resulting 3D scene information becomes, and the higher the quality of the displayed image becomes. However, since the time and memory usage required to generate learning images increases, it is desirable to determine the number of pseudo viewpoints generated by the pseudo viewpoint generator 72 depending on the processing power of the content server 20, the content of the scene, the purpose of generating the 3D scene information, and the like.

[0061] 8 is a schematic diagram showing how the main image and standby image displayed on the display device 16 are switched in this embodiment. As described above, while the main content is progressing, such as during game play, frames 250a of the main image are displayed on the display device 16 at a predetermined rate. In contrast, when the user performs a scene save operation at any time, the display switches to a standby image 252. In the example shown, the saturation and brightness of the main image frame 250a displayed during the save operation are reduced, and a progress indicator 254 indicating that processing is in progress is superimposed on it.

[0062] However, the configuration of the standby image is not limited to that shown in the figure, and it may be a simple solid image, or an image that does not include the image of frame 250a. Alternatively, the image of frame 250a itself may be subjected to some processing. Once the learning image has been generated, display resumes from frame 250b of the main image immediately following.

[0063] FIG. 9 is a diagram illustrating how the 3D scene information generation unit 76 extracts areas used for learning from a learning image in this embodiment. In this example, the main image 260 generated by the main image generation unit 84 of the application execution unit 74 includes, in addition to the scene image, superimposed additional images necessary for the content, such as a field 262a showing the game score and a field 262b showing icons of weapons held. If the main image generation unit 84 generates images without distinguishing between a display viewpoint and a pseudo viewpoint, the learning image may also have the same configuration. Therefore, the 3D scene information generation unit 76 excludes areas showing these additional images and uses only areas showing the scene itself for machine learning.

[0064] This can eliminate problems such as the inclusion of unnecessary information in the 3D scene information or the generation of false objects. The size and position of the region 264 can be set in advance based on the size and position of the additional image to be superimposed. However, the basis for setting the region 264 is not limited to the presence of the additional image, and consideration may also be given to the appropriateness of the scene to be viewed later. For example, the extracted region may be widened or narrowed depending on the range occupied by the image of the main object in the main image currently being displayed. In other words, the extracted region may be fixed, or may be variable depending on changes in the display content.

[0065] According to the above-described aspect of saving a user's desired scene, the content server 20 generates 3D scene information for a scene through machine learning in response to a user's operation to save a scene at a certain timing in the displayed main image. This allows the user to view a fleeting scene that appears during the progress of the content from any viewpoint at another time. The saved scene can also be shared with other users, such as friends. By being able to view the saved scene from any viewpoint, the user can look back on and verify the saved situation with a level of realism that cannot be achieved with conventional techniques such as image screenshots.

[0066] When saving a scene, multiple pseudo-viewpoints are generated according to the display situation at that time, and training images are generated intensively. This allows users without technical knowledge to efficiently generate images suitable for training with simple operations, ultimately enabling highly accurate 3D scene information to be generated in a short time. Furthermore, since pseudo-viewpoint information is generated in a format similar to normal application processing and supplied to the application to generate training images, it can be easily applied even to conventional applications that do not support machine learning.

[0067] 2. Display Image Correction FIG. 10 shows an overview of the processing flow for a mode in which 3D scene information is used to correct a display image. This mode is implemented in a main image output phase 270, during which a main image of content is output, such as during gameplay. During this phase, an image processing device, e.g., the content server 20, generates training images along with the main image to be displayed (S20), and performs machine learning to generate 3D scene information 272 representing the scene for each time step (S22). In other words, the 3D scene information 272 is updated over time. The image processing device, e.g., the client terminal 10, then corrects the main image to be displayed using the latest 3D scene information 272 (S24). Correcting an image consisting of two-dimensional information using 3D scene information containing three-dimensional information enables high-precision correction, improving the quality of the displayed image.

[0068] 11 is a diagram illustrating reprojection as an example of main image correction. Reprojection refers to a process of correcting a generated main image so that the field of view matches the position and orientation of the user's head immediately before display, for example, when the display device 16 is a head-mounted display. When a main image generated by the content server 20 is displayed on the client terminal 10, as shown in FIG. 6, it takes a certain amount of time from when the content server 20 recognizes the display viewpoint until the frame generated accordingly is displayed on the client terminal 10. In reality, it also takes time for the client terminal 10 to transmit the display viewpoint to the content server 20.

[0069] As a result, a delay occurs in the change in the field of view of the displayed main image relative to a change in the actual viewpoint, which can cause a noticeable sense of discomfort. This can degrade the quality of the user experience, especially when the display device 16 is a head-mounted display, by damaging the sense of immersion in virtual reality and causing motion sickness. Therefore, the client terminal 10 corrects the frame of the main image transmitted from the content server 20 to match the field of view immediately before display.

[0070] Illustrated in (a) is how the content server 20 generates a main image. The content server 20 sets a view screen 280a to correspond to the display viewpoint recognized at that time, and renders an image 284 contained in the corresponding view frustum 282a on the view screen 280a. Assume that the viewpoint at the time of display is shifted to the left, as indicated by the arrow. In this case, the client terminal 10 corrects the image so that it matches the field of view when the view screen 280b is shifted to the left, as shown in (b).

[0071] The view frustum 282b corresponding to the newly set viewscreen 280b does not include area 288 of the field of view 286 of the transmitted main image, but now includes area 290. Therefore, the client terminal 10 discards the image of area 288 and additionally renders the image in the newly required area 290, resulting in a corrected display image. At this time, the client terminal 10 adds the image using the latest 3D scene information generated by the content server 20, thereby generating a high-quality image that takes into account changes in color due to movement of the viewpoint.

[0072] 12 shows the functional block configuration of the client terminal 10 and the content server 20 that realizes correction of the displayed image. Note that blocks having the same functions as the functional blocks shown in FIG. 6 are assigned the same reference numerals, and descriptions thereof will be omitted where appropriate. The client terminal 10 includes an input information acquisition unit 50 that acquires input information such as a user operation, an image data acquisition unit 52 that acquires image data from the content server 20, a 3D scene information data acquisition unit 88 that acquires 3D scene information data from the content server 20, a 3D scene information storage unit 90 that stores the 3D scene information data, an image correction unit 92 that corrects the displayed image using the 3D scene information, and an output unit 54 that outputs the displayed image data.

[0073] The input information acquisition unit 50 acquires the details of user operations and information on display viewpoints as described above, and supplies the acquired information to the content server 20 and the image correction unit 92 as appropriate. The image data acquisition unit 52 acquires data for each frame of the main image from the content server 20. The 3D scene information data acquisition unit 88 sequentially acquires 3D scene information data that is continuously generated at predetermined time steps from the content server 20. The 3D scene information storage unit 90 stores the 3D scene information data acquired by the 3D scene information data acquisition unit 88.

[0074] The image correction unit 92 corrects the main image transmitted from the content server 20 using the 3D scene information data stored in the 3D scene information storage unit 58. That is, as described above, the latest display viewpoint is acquired from the input information acquisition unit 50, and the 3D scene information is used to additionally render the missing area of ​​the corresponding field of view. To this end, the content server 20 adds a timestamp to the main image data before transmitting it, and the image correction unit 92 acquires the amount of change in the display viewpoint based on the time difference between the timestamp and the time of correction, and identifies the missing part of the display image.

[0075] The image correction unit 92 then renders the missing area using the latest 3D scene information. Furthermore, the image correction unit 92 removes areas outside the field of view from the frame of the main image transmitted from the content server 20, and connects the removed areas with the area it has rendered to create a display image. However, the correction performed by the image correction unit 92 is not limited to adding or removing areas from the field of view. For example, the image correction unit 92 may re-render images of objects that are close and susceptible to changes in viewpoint, and their surrounding areas, using the 3D scene information. This allows the display of an image with colors adjusted to accommodate changes in viewpoint. Alternatively, the image correction unit 92 may render the entire display image using the 3D scene information.

[0076] If 3D scene information corresponding to scene transitions can be prepared using machine learning, the client terminal 10 can generate display images with a lighter load than with normal ray tracing or other processing, and can render high-quality images. Using this, on the premise that the client terminal 10 can ultimately generate display images using 3D scene information, the content server 20 does not need to generate a main image that strictly matches the display viewpoint. Therefore, the content server 20 may intentionally generate a main image from a viewpoint that is shifted from the display viewpoint to increase the efficiency of collecting learning images.

[0077] As an example, if the display device 16 is a head-mounted display, the image correction unit 92 may use 3D scene information to render at least one of the main images for the left eye and the right eye based on the latest display viewpoint. This eliminates the need to impose on the content server 20 the constraint of always generating a pair of highly redundant main images for the left eye and the right eye. For example, the content server 20 may generate a pair of main images with little overlapping fields of view by setting the distance between the left and right viewpoints wider than the actual distance. This allows a variety of learning images to be collected in a short period of time. The output unit 54 outputs the display image corrected or generated by the image correction unit 92 to the display device 16 for display.

[0078] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, a pseudo viewpoint generation unit 72 that generates a pseudo viewpoint for generating learning images, an application execution unit 74 that executes applications such as electronic games, a 3D scene information generation unit 76 that generates 3D scene information data, a 3D scene information storage unit 78 that stores the generated 3D scene information data, an image data transmission unit 82 that transmits main image data to the client terminal 10, and a 3D scene information data transmission unit 86 that transmits 3D scene information data to the client terminal 10.

[0079] The input information acquisition unit 70 acquires the content of user operations and information on the display viewpoint from the client terminal 10 at any time or at predetermined time intervals, and supplies the information to the application execution unit 74. The input information acquisition unit 70 also supplies the information on the display viewpoint to the pseudo viewpoint generation unit 72. The pseudo viewpoint generation unit 72 generates a pseudo viewpoint for generating learning images based on the latest display viewpoint. In this aspect, 3D scene information of a scene is learned while the main image is displayed, so opportunities to generate learning images are limited.

[0080] For this reason, the input information acquisition unit 70 may supply the information on the display viewpoint acquired at that time only to the pseudo viewpoint generation unit 72, and the pseudo viewpoint generation unit 72 may purposely shift the display viewpoint or add more pseudo viewpoints and supply them to the application execution unit 74. The pseudo viewpoint generation unit 72 may predict future display viewpoints according to a history of changes in the display viewpoint up to that point in time, and generate pseudo viewpoints with a distribution according to that prediction.

[0081] The application execution unit 74 processes the content application based on the details of the user operation. The application execution unit 74 includes a main image generation unit 84, which generates main image frames corresponding to a display viewpoint at a predetermined rate. However, as described above, the main image generation unit 84 may also generate images corresponding to a pseudo viewpoint obtained by shifting the display viewpoint as frames of the main image for display. The main image generation unit 84 also generates images of the scene viewed from the pseudo viewpoint generated by the pseudo viewpoint generation unit 72 as learning images.

[0082] In this embodiment, the pseudo viewpoint generation unit 72 generates pseudo viewpoint information in the same format as the viewpoint information supplied by the input information acquisition unit 70 and supplies the pseudo viewpoint information to the application execution unit 74. This allows the application execution unit 74 to generate learning images through normal processing without distinguishing between a true display viewpoint and a pseudo viewpoint. As a result, this embodiment can be easily implemented even in conventional content that does not support machine learning. However, as described above, the function of the pseudo viewpoint generation unit 72 may be provided in the application execution unit 74 using an API or the like.

[0083] The 3D scene information generation unit 76 acquires learning images, including a main image to be displayed, from the application execution unit 74 and generates 3D scene information for the scene at predetermined time steps using the machine learning described above. In this case, the 3D scene information generation unit 76 may also extract only the area of ​​the image generated by the main image generation unit 84 that is necessary for correcting the display image and use it for machine learning. The 3D scene information storage unit 78 temporarily stores the 3D scene information generated by the 3D scene information generation unit 76. The image data transmission unit 82 transmits data of the main image generated by the main image generation unit 84 to the client terminal 10 at a predetermined rate. The 3D scene information data transmission unit 86 transmits data of the 3D scene information stored in the 3D scene information storage unit 78 to the client terminal 10 at a predetermined rate.

[0084] FIG. 13 is a schematic diagram illustrating a sequence of images generated in this embodiment. The diagram illustrates the relationship between the viewpoint recognized or generated by the content server 20 and the destination of each frame generated thereby, with the horizontal axis representing time. As in FIG. 6 , the content server 20 basically generates display image frames (e.g., frames 302a and 302b) at a predetermined rate to correspond to the display viewpoints (e.g., display viewpoints 300a and 300b) indicated by white circles, and transmits them to the client terminal 10. However, as described above, the display viewpoints in this case may be essentially pseudo-viewpoints that are shifted from the actual display viewpoints. The client terminal 10 appropriately corrects and displays the transmitted images.

[0085] Furthermore, the content server 20 generates training images between the generation of display image frames, i.e., during the cycle until the next frame is generated. For example, the content server 20 generates pseudo viewpoints 304a and 304b, indicated by black circles, to be processed between the processing of display viewpoints 300a and 300b, and generates training images 306a and 306b corresponding to the pseudo viewpoints. The content server 20 also uses frames of the display image to be transmitted to the client terminal 10 as training images. As shown in the figure, if images are drawn at a rate higher than the display frame rate, the training images required to generate 3D scene information can be efficiently obtained.

[0086] For example, if the display frame rate is 60 fps, and the main image generation unit 84 operates at 120 fps, it is possible to acquire twice as many training images as the frames of the display image. If the main image generation unit 84 operates at 180 fps, it is possible to acquire three times as many training images as the frames of the display image. Note that the example shown in the figure shows a case where a display image is sent to one client terminal 10, but if the display viewpoint of the image sent to another user's client terminal 10 is different, such as in a multiplayer game, that image can also be used as a training image. By efficiently collecting training images in this way, it is possible to improve the accuracy of the 3D scene information representing the scene at each time step, and ultimately to display high-quality images.

[0087] 14 is a diagram illustrating a manner in which images are generated by shifting the display viewpoint when displaying images for the left and right eyes on a head-mounted display. The diagram schematically illustrates the display viewpoint for a scene 310. When a head-mounted display is used as the display destination, a pair of display viewpoints 312a and 312b is set at a distance D1 that takes into account the actual distance between the eyes, and each image is generated in the field of view indicated by the dashed lines. By displaying these image pairs on the head-mounted display at positions corresponding to the user's left and right eyes, the scene 310 can be viewed stereoscopically.

[0088] The distance D1 between the display viewpoints 312a and 312b set in this case is generally called the inter-pupilary distance (IPD), and is, for example, approximately 60 mm for an adult. However, since IPD varies from person to person, it is often possible to set it as a variable parameter for the head-mounted display to achieve optimal stereoscopic vision. Typically, an image pair is generated based on the IPD setting. On the other hand, as shown in the figure, the normal display viewpoints 312a and 312b have a large overlap in the field of view of the scene 310. In other words, from the perspective of using them as learning images, the image pair generated using this setting is redundant and inefficient. Therefore, the pseudo viewpoint generator 72 sets the IPD value to a significantly wider value, such as 1 m.

[0089] The figure shows how display viewpoints 314a and 314b, which are spaced farther apart than the original display viewpoints 312a and 312b, are set by setting the IPD value to D2 (>D1). By generating images based on this setting, as indicated by the dashed-dotted lines, information on a wider range of the scene 310 can be obtained by processing frames at each time point, thereby enabling highly accurate 3D scene information to be generated in a short time. Note that the display viewpoints 314a and 314b set here differ from the actual display viewpoints 312a and 312b. Therefore, as described above, the image correction unit 92 of the client terminal 10 generates display images representing the scene as seen from the actual display viewpoints 312a and 312b using the 3D scene information. This configuration can also be achieved simply by changing the IPD setting value. Therefore, the application execution unit 74 simply performs normal processing, making it easily applicable to conventional content that does not support machine learning.

[0090] According to the above-described display correction mode, the content server 20 generates training images in parallel with the generation of display images, and generates 3D scene information for the scene at each time step. The client terminal 10 sequentially obtains the latest 3D scene information from the content server 20 and uses it to correct and draw the display image. This allows images to be displayed that follow the movement of the viewpoint while accurately representing changes in color due to changes in viewpoint, which cannot be obtained from transmitted images alone. Furthermore, because the client terminal 10 can generate display images with a light load, the content server 20 has greater freedom in the viewpoint from which images are generated, allowing it to collect training images more efficiently.

[0091] 15 shows an overview of the processing flow for a mode in which 3D scene information is used to distribute replay images. This mode is realized in two periods: a main image output phase 320 and a replay image distribution phase 322. During the main image output phase 320, in which the main image of the content is being output, such as during game play, an image processing device, such as the content server 20, collects learning images (S30) and performs machine learning to generate 3D scene information 324 representing the scene for each time step (S32).

[0092] As explained above, the learning images collected in S30 may be drawn by the image processing device itself generating a pseudo viewpoint. On the other hand, in a case where the content server 20 receives multiple display viewpoints, generates main images in parallel, and distributes them to each client terminal 10, as in a multiplayer game, the learning images may be the display images of those images. The following explanation focuses on this case. However, even in this case, the content server 20 may set additional viewpoints to increase the number of learning images.

[0093] The replay image distribution phase 322 is initiated when a user requests distribution at any time, such as after the end of game play. Note that users requesting replay image distribution are not limited to users who performed operations in the main image output phase 320, such as game players. In the replay image distribution phase 322, the content server 20 generates replay images using the saved 3D scene information 324 and outputs them to the client terminal 10 that requested distribution (S36). By updating the 3D scene information at each time step and inputting the time to generate images, they can be displayed as moving images. Furthermore, replay images can be displayed from various positions and directions in response to user operations that change the viewpoint.

[0094] In this embodiment, the larger the display world, the more biased the display viewpoints become in the main image output phase 320. As a result, 3D scene information 324 is generated with high accuracy in areas where the density of display viewpoints is high, but the accuracy of the 3D scene information 324 is low in areas where the density is low. Furthermore, 3D scene information 324 cannot be generated in areas where no display viewpoints exist, and therefore replay images cannot be displayed. Therefore, in the main image output phase 320, the content server 20 generates a heat map indicating the density of display viewpoints (S34). Then, in the replay image distribution phase 322, the content server 20 displays the heat map together with the replay images, allowing the user to refer to the heat map as guidance when operating the viewpoint (S38).

[0095] 16 shows the functional block configuration of the client terminal 10 and content server 20 that realizes the distribution of replay videos. Note that blocks having the same functions as the functional blocks shown in FIG. 6 are assigned the same reference numerals, and descriptions thereof will be omitted where appropriate. Also, while only one client terminal 10 is shown in the example of the figure, at least during the main image output phase, the client terminals 10 of all users participating in the content are connected to the content server 20 and perform the same functions.

[0096] The client terminal 10 includes an input information acquisition unit 50 that acquires input information such as user operations, an image data acquisition unit 52 that acquires image data from the content server 20, and an output unit 54 that outputs display image data. The input information acquisition unit 50 acquires the content of user operations from the input device 14 at any time. The input information acquisition unit 50 also accepts an operation to request the distribution of replay images in a replay image distribution phase 322. The input information acquisition unit 50 also acquires information on the display viewpoint for the main image or replay images from the input device 14 or head-mounted display at any time or at predetermined time intervals. The input information acquisition unit 50 supplies the acquired information to the content server 20 as appropriate.

[0097] The image data acquisition unit 52 acquires display image data from the content server 20. Here, the display image data may include main image data, replay image data, and heat map data. The output unit 54 outputs the display image acquired by the image data acquisition unit 52 to the display device 16 for display.

[0098] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, an application execution unit 74 that executes applications such as electronic games, a 3D scene information generation unit 76 that generates data for 3D scene information, a 3D scene information storage unit 78 that stores the generated data for the 3D scene information, a replay image generation unit 100 that generates replay images, an image data transmission unit 82 that transmits data for display images to the client terminal 10, and a restriction information storage unit 102 that stores restriction information related to the distribution of replay images.

[0099] The input information acquisition unit 70 acquires information on the content of user operations and display viewpoints from the client terminal 10 at any time or at predetermined time intervals, and supplies the information to the application execution unit 74. In the main image output phase, the application execution unit 74 processes an application for content such as an electronic game based on the content of user operations. The application execution unit 74 includes an additional viewpoint setting unit 104, a main image generation unit 84, and a heat map generation unit 106.

[0100] The additional viewpoint setting unit 104 sets an additional viewpoint from which a main image should be generated, independent of the display viewpoint transmitted from the client terminal 10. The additional viewpoint is similar to a pseudo viewpoint in the sense that it is not used for display in the main image output phase, but differs from a pseudo viewpoint in that it takes into account the entire display world and determines a viewpoint that is deemed necessary for optimally generating replay images in accordance with the content. For example, in a role-playing game, the additional viewpoint setting unit 104 sets an additional viewpoint at a location where an event is likely to occur, thereby ensuring the accuracy of 3D scene information representing that location.

[0101] In this way, the additional viewpoint setting unit 104 may predict events that may occur in the display world and set an additional viewpoint accordingly, or may supplement a viewpoint in a location where it is difficult to position a display viewpoint in the main image output phase, taking into account the geographical situation in the display world. The additional viewpoint setting unit 104 may further add a viewpoint where a display viewpoint cannot occur, such as a viewpoint that follows a virtual user present in the display world from behind, a viewpoint that views the virtual user diagonally from above, or a viewpoint that overlooks the display world.

[0102] In this way, the additional viewpoint setting unit 104 may set a fixed additional viewpoint in the display world and use it like a fixed camera, or may set it to move according to the situation or the movement of the virtual user. The additional viewpoint setting unit 104 may also set the additional viewpoint according to a program that defines the application, or may receive the additional viewpoint setting from the user as an initial setting for the main image output phase. In either case, by setting additional viewpoints based on various criteria within the processing capabilities of the content server 20, the accuracy of 3D scene information can be increased and the quality of replay images can be improved. Furthermore, the user can reconfirm the situation that occurred in the display world from a position or orientation that was not visible during the main image output phase.

[0103] The main image generation unit 84 generates frames of a main image corresponding to the display viewpoint transmitted from the client terminal 10 at a predetermined rate. The main image generation unit 84 also generates images of the display world viewed from the viewpoint added by the additional viewpoint setting unit 104 at a predetermined rate. In the main image output phase, the heat map generation unit 106 generates a heat map that represents the distribution of densities of the display viewpoints and the additionally set viewpoints on the surface of the display world. For example, the heat map generation unit 106 colors the map, which provides a bird's-eye view of the display world, to distinguish between areas with a high density of display viewpoints, areas with a medium density, areas with a low density, and areas with no display viewpoint.

[0104] The higher the density of display viewpoints, the more diverse the learning images obtained, and therefore the more accurately 3D scene information is obtained, and the higher the quality of the replay images is thought to be. Conversely, if there are no display viewpoints, or if there are so few that they are considered nonexistent, even if the viewpoint is aligned with that location in the replay image distribution phase, the replay image cannot be displayed because no 3D scene information has been generated. Therefore, by generating a heat map in the main image output phase and making it possible to refer to it when manipulating the viewpoint of the replay image, the user can easily set an appropriate viewpoint.

[0105] In the main image output phase, the 3D scene information generation unit 76 uses the images generated by the application execution unit 74 as learning images and generates 3D scene information representing the scene at each time step through machine learning as described above. In this case, the 3D scene information generation unit 76 may extract only the area of ​​the image generated by the main image generation unit 84 that is necessary for generating a replay image and use it for machine learning. The 3D scene information generation unit 76 may also limit the area of ​​the displayed world for generating 3D scene information based on the heat map generated by the heat map generation unit 106. That is, the 3D scene information generation unit 76 may target locations where the density of display viewpoints and additional viewpoints is higher than a threshold value as targets for generating 3D scene information.

[0106] The 3D scene information storage unit 78 stores the 3D scene information generated by the 3D scene information generation unit 76. The 3D scene information storage unit 78 stores data of the 3D scene information generated at each time step in association with the time axis in the main image output phase. In the replay image distribution phase, when a user requests distribution of a replay image, the replay image generation unit 100 generates a replay image by the above-described volume rendering using the 3D scene information stored in the 3D scene information storage unit 78. At this time, the replay image generation unit 100 acquires a display viewpoint from the input information acquisition unit 70 and generates a replay image while changing the viewpoint accordingly.

[0107] Here, the replay image generator 100 may restrict at least one of the distribution timing and the display viewpoint of the replay image based on the restriction information stored in the restriction information storage unit 102. For example, the replay image generator 100 does not generate a replay image if a predetermined time has not elapsed since the end of the main image output phase. This prevents adverse effects, such as the content being made known early and reducing the desire to purchase the application. Furthermore, the replay image generator 100 does not generate a corresponding replay image when the display viewpoint is operated to a position or direction where display as a replay image is undesirable. In this case, the replay image generator 100 may generate a display image indicating that the display viewpoint exceeds the restriction.

[0108] As an initial process when an application is executed, replay image generation unit 100 reads the above-described restriction information from a setting file that defines the application, and stores the information in restriction information storage unit 102. In the main image output phase, image data transmission unit 82 transmits main image data generated by main image generation unit 84 to client terminal 10 at a predetermined rate. In the replay image distribution phase, image data transmission unit 82 also transmits replay image data generated by replay image generation unit 100 to client terminal 10 in response to a distribution request.

[0109] Here, the image data transmission unit 82 may limit the distribution destinations of replay images using 3D scene information based on the restriction information stored in the restriction information storage unit 102. For example, the image data transmission unit 82 may transmit replay images using 3D scene information only to the client terminals 10 of users who participated in the main image output phase. The image data transmission unit 82 may transmit general replay videos that do not use 3D scene information to the client terminals 10 of other users. In this case, in the main image output phase, replay images are generated from a predetermined display viewpoint and stored in a storage unit (not shown). This type of configuration also makes it possible to prevent detailed content from being easily made public.

[0110] FIG. 17 schematically illustrates a sequence of images generated in the main image output phase of this embodiment. The diagram illustrates the relationship between the viewpoints recognized or generated by the content server 20 and the destinations of each frame generated thereby, with the horizontal axis representing time. In this case, the content server 20 acquires display viewpoints (e.g., display viewpoints 330a, 330b, 330c) from each of the multiple client terminals 10a, 10b, 10c, etc. The content server 20 then generates display image frames (e.g., frames 332a, 332b, 332c) at a predetermined rate corresponding to the display viewpoints and transmits them to each of the client terminals 10a, 10b, 10c, etc. As a result, the client terminals 10a, 10b, 10c, etc. display images that represent the common display world as seen, for example, from the position or orientation of a virtual user.

[0111] The content server 20 further generates learning images (e.g., learning images 336) at a predetermined rate so as to correspond to the viewpoints (e.g., viewpoint 334) additionally set by the additional viewpoint setting unit 104 and indicated by black circles. In the example shown in the figure, there is a slight time difference between the timing of recognizing the multiple display viewpoints and the timing of generating the additional viewpoint to be set, but in reality, these may be simultaneous or may be independent of each other. In addition, the additional viewpoint setting unit 104 may actually add multiple viewpoints.

[0112] The 3D scene information generation unit 76 performs machine learning using all frames of the display image to be transmitted to the client terminal 10 and images corresponding to the additional viewpoints as learning images. For example, in the case of an MMO (Massively Multiplayer Online) game with 100 or more players, 100 or more learning images can be collected per frame. This allows for efficient collection of learning images, improves the accuracy of the 3D scene information representing the scene at each time step, and ultimately makes it easier to maintain the quality of replay images even when the viewpoint changes.

[0113] 18 illustrates an example of a screen that the additional viewpoint setting unit 104 of the content server 20 displays to accept the setting of an additional viewpoint by the user. In this example, the additional viewpoint reception screen 340 has a configuration in which a map of an overhead view of the displayed world is used as a base image, and a camera icon 344 and a message 342 prompting the user to set an additional viewpoint are superimposed on the base image. The user of the client terminal 10 places the icon 344 in a desired position and orientation by, for example, moving it via the input device 14. In response to this, the additional viewpoint setting unit 104 sets the additional viewpoint at a corresponding position and direction in the three-dimensional space of the displayed world.

[0114] The additional viewpoint reception screen 340 further displays a prohibited area 346 for viewpoint setting. The additional viewpoint setting unit 104 controls the display so that the user cannot place an icon 344 in the prohibited area 346. This prevents the viewpoint from being set in an inappropriate location, which could result in the viewpoint being displayed in a replay image or the unnecessary generation of learning images. The position and shape of the prohibited area 346 in the display world are set in advance in an application settings file, etc. Note that the illustrated example shows a reception screen for setting a fixed additional viewpoint, but the type of additional viewpoint accepted by the user is not limited. For example, an additional viewpoint may be set behind the virtual user in the display world. In such a case, the additional viewpoint setting unit 104 may display the options for the viewpoint type using text or the like, allowing the user to select and input.

[0115] FIG. 19 illustrates a heat map generated by the heat map generator 106 of the content server 20. In this example, the heat map 350 uses a map of the display world viewed from above as a base image, and superimposes areas where display viewpoints are distributed (e.g., areas 352a and 352b) with color intensities corresponding to the density levels. In practice, the density levels may be represented by different colors, such as red, yellow, and blue. As illustrated, if there is a bias in the display world among the areas where display viewpoints, and thus virtual users, exist, many of these areas will be unsuitable for generating 3D scene information. Therefore, the heat map generator 106, for example, renders areas where the density of display viewpoints is below a threshold colorless, preventing viewpoints from being set in the replay image distribution phase.

[0116] Areas with a high density of display viewpoints allow for high-precision generation of 3D scene information and are likely to be popular content. Therefore, by targeting such locations as display viewpoints, users viewing replay images can easily enjoy high-quality replay images of popular scenes even in a large display world. The heat map generator 106 may update the heat map at a predetermined rate in response to changes in the distribution of display viewpoints.

[0117] In this case, when the replay image is distributed, the heat map is distributed as a video in synchronization with the replay image, allowing the user to appropriately determine the display viewpoint to respond to changes in the density distribution. This aspect is suitable for content in which the virtual user's range of movement in the displayed world is wide and the density distribution is prone to change. On the other hand, in content in which the virtual user's range of movement is narrow, the heat map generator 106 may integrate the heat map for each time step and distribute a still image of the final heat map obtained.

[0118] 20 illustrates an example of a display screen of a replay image displayed on the display device 16 during the replay image distribution phase. Conventionally, distributed images of games and the like are typically viewed via a browser on a video viewing platform, as a video with a defined display viewpoint. Due to the unique nature of this embodiment, which requires viewpoint manipulation of the replay image, it is difficult to implement on such a general platform.

[0119] Therefore, preferably, a unique platform is provided that provides a UI (User Interface) on a browser for manipulating the viewpoint, allowing users to enjoy replay images using general-purpose devices such as personal computers, tablet devices, and mobile phones. In this case, the content server 20 transmits data setting the replay images, heat maps, and UI to the client terminal 10 using a markup language such as HTML. The client terminal 10 generates a replay image display screen using the browser and displays it on the display device 16. Viewpoint operation information is transmitted from the client terminal 10 to the content server 20 as needed, and corresponding data is transmitted from the content server 20 to the client terminal 10.

[0120] In the illustrated example, the replay image display screen 360 includes a replay image field 362, a heat map field 364, a candidate viewpoint field 366, and a viewpoint control UI 368. The replay image field 362 displays the replay image currently being streamed. The user can change the viewpoint for the displayed scene by operating the viewpoint control UI 368. In this example, the viewpoint control UI 368 is a directional key that can instruct movement of the viewpoint in four directions. For example, when the upward arrow is designated, the viewpoint moves forward. When the rightward arrow is designated, the viewpoint rotates to the right.

[0121] However, the shape and configuration of the viewpoint manipulation UI 368 are not limited to this. For example, the viewpoint position and the line of sight direction may be independently manipulated, or an object located at the center of the field of view may be fixed and the elevation / depression angle, azimuth angle, or distance relative to the fixed object may be changed. Furthermore, the viewpoint manipulation UI 368 is not limited to a GUI (Graphical User Interface), and may be configured such that options representing the type of viewpoint, such as a viewpoint that follows from behind the main object or a viewpoint that overlooks the entire object, are displayed by characters or the like, allowing the user to select and input the option.

[0122] A heat map is displayed in the heat map field 364. As described above, the heat map represents the density distribution of display viewpoints in the main image output phase and serves as an indicator of the quality of replay images using 3D scene information. Therefore, the position of the viewpoint can also be specified on the displayed heat map. When the user designates a point on the heat map using a cursor or touch operation (not shown), the viewpoint of the replay image displayed in the replay image field 362 is moved to the designated position.

[0123] The heat map allows the user to intuitively identify areas where 3D scene information is unavailable or where the accuracy of the 3D scene information is low. Therefore, by setting the viewpoint in a high-density area, it becomes easier to view a lively scene in high image quality. Note that the operation accepted using the heat map is not limited to specifying the viewpoint position, but may also specify the line of sight. In this case, for example, a camera icon or arrow may be superimposed on the heat map, and the line of sight may be specified by changing its orientation.

[0124] Furthermore, it is believed that areas with a high density of display viewpoints generate high-quality 3D scene information regardless of the direction. Therefore, when the viewpoint is set in an area with the highest density, the line of sight may be allowed to change in all directions, while the range of movement of the line of sight may be limited in other areas. When the viewpoint position or line of sight direction is manipulated using the viewpoint manipulation UI 368, an arrow or the like superimposed on the heat map may be linked to the manipulation. This allows intuitive understanding of the relationship between the currently displayed replay image and the viewpoint in the displayed world. Furthermore, when the viewpoint position or line of sight direction exceeds the limited range due to viewpoint manipulation, a concealment object may be superimposed on the corresponding area in the field of view of the currently displayed replay image.

[0125] In particular, when the displayed world is vast, the heat map may also be configured to accept operations such as zooming in and out and moving the display range. The candidate viewpoint field 366 displays thumbnails of replay images from viewpoints selected by the content server 20 according to predetermined criteria as so-called "recommended" images. For example, an area with the highest density in the heat map may be selected, and the candidate viewpoint field 366 may display thumbnails of replay images viewed from several of those viewpoints. Alternatively, replay images in which the virtual user himself or a specified player is within the field of view in the displayed world may be displayed. The heat map may also indicate the position and direction of the viewpoint from which the replay image thumbnails displayed in the candidate viewpoint field 366 are viewed.

[0126] When the user selects one of the thumbnail images using a cursor or touch operation (not shown), the display viewpoint switches, and the replay image that was displayed as a thumbnail is displayed in the replay image field 362. Note that when a viewpoint position is specified on the heat map or a thumbnail image is selected in the candidate viewpoint field 366, there is a possibility that the viewpoint will be discontinuously displaced from the replay image that was previously displayed in the replay image field 362.

[0127] Here, the content server 20 may move the viewpoint by generating a trajectory that smoothly connects the original viewpoint to the new viewpoint, and display a replay image that shows the movement process. For example, the content server 20 may first move the viewpoint into the sky and then give it a movement that makes it appear as if it is descending from there to the new viewpoint. This type of presentation creates a unique enjoyment that can only be achieved with replay images, improving the quality of the viewing experience.

[0128] According to the above-described aspect of distributing replay videos, the content server 20 collects frames of the main image to be transmitted to multiple client terminals 10 in the main image output phase and frames of images corresponding to additionally set viewpoints as learning images, and generates 3D scene information for each time step. This allows for the distribution of replay images that can be viewed from any viewpoint. The content server 20 also generates a heat map representing the density distribution of display viewpoints relative to the main image in parallel with the learning process. The density of the display viewpoints is linked to the accuracy of the 3D scene information and the level of excitement in the scene. Therefore, by displaying the heat map simultaneously with the replay image, it can be used as a basis for viewpoint manipulation for the replay image, making it easy to view exciting scenes in high image quality, even in a vast display world.

[0129] The content server 20 also provides a platform that allows users to watch replay videos and control the viewpoint on a general browser. The displayed screen displays a UI for controlling the viewpoint, as well as a heat map and thumbnail images of recommended viewpoints. This allows users to watch replay videos while easily controlling the viewpoint on a general-purpose device, even in environments where a specific device such as a game console is not available.

[0130] 4. Restrictions on Display Viewpoints by Applications In the above-described modes for viewing saved scenes and replay videos, display from any viewpoint is basically enabled by learning the main image of the content and generating 3D scene information. However, setting an additional viewpoint different from the original display viewpoint outside the application execution unit in order to acquire learning images, or enabling free viewpoint movement using the generated 3D scene information, carries the risk of exposing the displayed world beyond the visible range originally intended by the content.

[0131] For example, in a replay image of a role-playing game, if a user selects a viewpoint that overlooks the displayed world, they may see the place they should reach in the future, which may dampen their interest and reduce their desire to purchase the application.Furthermore, depending on the content and the state of image creation, there are likely to be many viewpoints that content developers do not want, such as the viewpoint of an enemy character or a viewpoint close to a background object.

[0132] Therefore, in this embodiment, restrictions are intentionally imposed on one or both of the viewpoint setting for generating learning images and the viewpoint setting for displaying images using 3D scene information. The content server 20 reads the restriction information set by the developer for each piece of content from the application, uses it to set the viewpoint, or adds it to the 3D scene information as metadata. This embodiment can be combined with the scene saving embodiment and the replay video distribution embodiment described above. Therefore, as with those embodiments, this description will be based on the main image output phase and the viewing phase of an arbitrary viewpoint image using 3D scene information.

[0133] FIG. 21 shows a functional block configuration of the content server 20 in a mode in which the display viewpoint is limited by an application. Note that blocks having the same functions as the functional blocks shown in FIG. 6 are assigned the same reference numerals, and descriptions thereof will be omitted where appropriate. Furthermore, the client terminal 10 is similar to the client terminal 10 shown in FIG. 5 and FIG. 16 , and is therefore not shown. The functional blocks shown in this figure can be combined with either the content server 20 shown in FIG. 5 that enables users to save scenes, or the content server 20 shown in FIG. 16 that enables distribution of replay videos. Furthermore, as described above, at least some of the functions shown in the figure may be performed by the client terminal 10, and it is not intended that the processing entity be limited to the content server 20.

[0134] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, an additional viewpoint setting unit 110 that generates a viewpoint for generating learning images, an application execution unit 74 that executes applications such as electronic games, a 3D scene information generation unit 76 that generates data of 3D scene information, a 3D scene information storage unit 78 that stores the data of the generated 3D scene information, an arbitrary viewpoint image generation unit 114 that generates an image from an arbitrary viewpoint using the 3D scene information, and an image data transmission unit 82 that transmits data of a display image to the client terminal 10. Note that the functional blocks other than the application execution unit 74 can also be collectively referred to as a system unit, because they are responsible for peripheral processing required for the system side of the content server 20, i.e., the application execution unit 74, to execute an application.

[0135] First, in the main image output phase, the application execution unit 74 processes a content application, such as an electronic game, based on the details of a user operation. Here, the application execution unit 74 includes a main image generation unit 84 that generates a main image frame, as well as a viewpoint restriction information storage unit 112 that stores viewpoint restriction information that is set during application development and associated with the application program. The viewpoint restriction information is information that imposes restrictions on at least one of the viewpoint set when generating a learning image in the main image output phase and the display viewpoint operated in the arbitrary viewpoint image output mode. The restriction may be imposed on either the viewpoint position or the line of sight direction, or both.

[0136] For example, during the content development stage, the content server 20 provides a viewpoint restriction setting screen to a developer's terminal (not shown), and the developer inputs restriction information into the setting screen. The setting screen displays possible restriction options, allowing the developer to select an option or simply input numerical values ​​as needed, thereby reducing the effort required to set restriction information. This allows the developer to easily set detailed settings, such as "only permitting line-of-sight in all directions from viewpoints within a radius of 1 to 3 meters from the virtual player." The viewpoint's movable range is not limited to a fixed area in the displayed world, but may also be an area that moves or changes shape depending on the situation. In other words, the restriction information may specify a fixed area in the displayed world, or it may specify changes in the viewpoint restriction range.

[0137] The input information acquisition unit 70 acquires the content of user operations and information on display viewpoints from the client terminal 10 at any time or at predetermined time intervals. The additional viewpoint setting unit 110 has a function similar to that of the pseudo viewpoint generation unit 72 shown in Fig. 5 or the additional viewpoint setting unit 104 shown in Fig. 16, and sets a viewpoint for generating learning images. In other words, the viewpoint set by the additional viewpoint setting unit 110 may be based on the display viewpoint transmitted from the client terminal 10, or may be based on the content of the content, such as the configuration of the display world.

[0138] When setting the additional viewpoint, the additional viewpoint setting unit 110 reads viewpoint restriction information from the viewpoint restriction information storage unit 112 of the application execution unit 74, and sets the viewpoint within the permitted range. Alternatively, the additional viewpoint setting unit 110 may inquire of the application execution unit 74 via an API about whether or not the additional viewpoint can be set for each generated viewpoint. The additional viewpoint setting unit 110 supplies information about the additional viewpoint set through these steps to the application execution unit 74.

[0139] In the main image output phase, the main image generation unit 84 generates, at a predetermined rate, an image corresponding to the display viewpoint transmitted from the client terminal 10 and an image corresponding to the viewpoint additionally set by the additional viewpoint setting unit 110. As described above, the additional viewpoint setting unit 110 generates additional viewpoint information in the same format as the viewpoint information supplied by the input information acquisition unit 70 and supplies this to the application execution unit 74, so that the application execution unit 74 can generate learning images through normal processing without distinguishing between a true display viewpoint and an additional viewpoint.

[0140] The 3D scene information generation unit 76 generates 3D scene information of the scene to be saved by machine learning as described above, using the image generated by the application execution unit 74 as a learning image. The 3D scene information storage unit 78 stores the 3D scene information generated by the 3D scene information generation unit 76. In the viewing phase of the arbitrary viewpoint image, the arbitrary viewpoint image generation unit 114 generates an arbitrary viewpoint image by volume rendering as described above, using the 3D scene information stored in the 3D scene information storage unit 78.

[0141] Here, the arbitrary viewpoint image generation unit 114 acquires a display viewpoint from the input information acquisition unit 70 and generates an arbitrary viewpoint image while changing the viewpoint accordingly. When generating an image, the arbitrary viewpoint image generation unit 114 reads viewpoint restriction information from the viewpoint restriction information storage unit 112 of the application execution unit 74 and generates an image by limiting the viewpoint to within the permitted range. Alternatively, the arbitrary viewpoint image generation unit 114 may inquire of the application execution unit 74 via an API about whether or not each display viewpoint can be set.

[0142] In the range where the setting of an additional viewpoint is not permitted during the main image output phase, it is believed that there are insufficient learning images and the accuracy of the 3D scene information is not high. Therefore, if the generation of display images using that range as a viewpoint is not permitted when generating arbitrary viewpoint images, problems such as a sudden decrease in the quality of images that newly enter the field of view due to viewpoint manipulation can be avoided. Conversely, even if the restriction imposed on the display viewpoint is lifted by some kind of tampering, if the setting of an additional viewpoint used to generate learning images is not permitted and detailed 3D scene information is not generated, the state of that area will not be visible in detail.

[0143] In this way, by imposing restrictions in the viewpoint restriction information on both the viewpoint set for generating learning images and the display viewpoint operated when generating arbitrary viewpoint images, the risk of the displayed world being displayed with an angle of view not desired by the content developer can be reduced. However, as described above, this is not intended to limit the present embodiment, and restrictions may be imposed on only one of the viewpoints. Note that, during the viewing phase of the arbitrary viewpoint image, the arbitrary viewpoint image generation unit 114 may stop movement of the display viewpoint when the display viewpoint transmitted from the client terminal 10 reaches the boundary of the restricted range. Alternatively, the arbitrary viewpoint image generation unit 114 may conceal the area of ​​the image that newly enters the field of view when the restricted range is exceeded by, for example, superimposing a concealing object.

[0144] In the main image output phase, the image data transmission unit 82 transmits data of the main image generated by the main image generation unit 84 to the client terminal 10 at a predetermined rate. In the arbitrary viewpoint image viewing phase, the image data transmission unit 82 also transmits data of the arbitrary viewpoint image generated by the arbitrary viewpoint image generation unit 114 to the client terminal 10.

[0145] The 3D scene information generation unit 76 may read the viewpoint restriction information from the viewpoint restriction information storage unit 112 and store it in the 3D scene information storage unit 78 as metadata for the generated 3D scene information. Fig. 22 shows an example of the data structure of the display 3D scene information in this aspect. The display 3D scene information data 370 includes an identification information field 372, a viewpoint restriction information field 374, and a 3D scene information field 376. The identification information field 372 stores various information for identifying the 3D scene information, such as the identification number of the 3D scene information, the identification information of the original content, and the identification information of the user who requested the generation.

[0146] The viewpoint restriction information field 374 stores the viewpoint restriction information read by the 3D scene information generation unit 76 from the viewpoint restriction information storage unit 112. The 3D scene information field 376 stores the main body of the 3D scene information generated by the 3D scene information generation unit 76. In this case, the arbitrary viewpoint image generation unit 114 first refers to the identification information field 372 to identify the 3D scene information corresponding to the user request and reads it from the 3D scene information storage unit 78. The arbitrary viewpoint image generation unit 114 further reads the viewpoint restriction information from the viewpoint restriction information field 374 to confirm whether the display viewpoint is appropriate, and if it is within the restricted range, generates a display image using the 3D scene information stored in the 3D scene information field 376.

[0147] By associating the viewpoint restriction information with the 3D scene information, the arbitrary viewpoint image generation unit 114 can generate an arbitrary viewpoint image after appropriately restricting the viewpoint, even in an environment without the application execution unit 74. Alternatively, even in a mode in which the display 3D scene information data 370 itself is transmitted to the client terminal 10 or another content server 20, or stored on a recording medium and distributed, the viewpoint restriction desired by the developer of the original content is observed by the arbitrary viewpoint image generation unit 114 provided in the device used to display the arbitrary viewpoint image.

[0148] According to the above-described aspect, viewpoint restriction information is set in advance during content development, taking into account the content of the content, etc. This prevents an image from being displayed unintentionally from a field of view that the content developer does not want when setting a viewpoint for a learning image or generating a display image of an arbitrary viewpoint using 3D scene information obtained by learning outside the application execution unit 74. Furthermore, by adding restriction information to the 3D scene information, restrictions can be imposed on the viewpoint during display, regardless of the environment in which the image is displayed using the 3D scene information.

[0149] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention.

[0150] As described above, the present invention can be used in various information processing devices such as content servers, game devices, head-mounted displays, display devices, mobile terminals, and personal computers, as well as image display systems including any of these.

[0151] 1 Image processing system, 10 Client terminal, 14 Input device, 16 Display device, 20 Content server, 50 Input information acquisition unit, 52 Image data acquisition unit, 54 Output unit, 70 Input information acquisition unit, 72 Pseudo viewpoint generation unit, 74 Application execution unit, 76 3D scene information generation unit, 78 3D scene information storage unit, 80 Standby image generation unit, 81 Saved scene image generation unit, 82 Image data transmission unit, 84 Main image generation unit, 86 3D scene information data transmission unit, 88 3D scene information data acquisition unit, 90 3D scene information storage unit, 92 Image correction unit, 100 Replay image generation unit, 102 Restriction information storage unit, 104 Additional viewpoint setting unit, 106 Heat map generation unit, 110 Additional viewpoint setting unit, 112 Viewpoint restriction information storage unit, 114 Arbitrary viewpoint image generation unit, 122 CPU, 124 GPU, 126 Main memory

Claims

1. An image processing device comprising: an application execution unit that executes an application program and generates, at a predetermined rate, frames of display images that represent a three-dimensional display world in which the situation changes in response to user operations; and a system unit that causes the application execution unit to generate learning images that represent the display world and are different from the display images, and that performs a process of generating 3D scene information that represents three-dimensional information of the display world through machine learning using the learning images as training data and using the information for display, and that restricts the viewpoint set for the display world in the process based on viewpoint restriction information associated with the application program.

2. The image processing device described in claim 1, characterized in that the system unit is provided with an additional viewpoint setting unit that sets a viewpoint within the range of restrictions indicated by the viewpoint restriction information and supplies it to the application execution unit, thereby generating the learning image.

3. The image processing device according to claim 1 or 2, characterized in that the system unit is provided with an arbitrary viewpoint image generation unit that uses the 3D scene information to generate a display image that represents how the display world appears as seen from an arbitrary viewpoint within the range of restrictions indicated by the viewpoint restriction information.

4. The image processing device according to claim 1 or 2, characterized in that the system unit is provided with an arbitrary viewpoint image generation unit that uses the 3D scene information to generate a display image showing the display world as seen from an arbitrary viewpoint, and that conceals an area of ​​the display image that newly comes into view when the viewpoint exceeds the restricted range indicated by the viewpoint restriction information, by superimposing a concealment object.

5. The image processing device according to claim 1 or 2, characterized in that the system unit is provided with a 3D scene information generation unit that generates the 3D scene information by performing the machine learning, and associates the viewpoint restriction information as metadata and stores it in a memory unit.

6. An image processing device as described in claim 1 or 2, characterized in that the system unit changes the viewpoint limit range set for the display world depending on the situation based on the viewpoint limit information that specifies the change in the viewpoint limit range.

7. An image processing device comprising: a 3D scene information storage unit that stores 3D scene information consisting of a neural network that represents three-dimensional information of a displayed world, and viewpoint restriction information that corresponds to the 3D scene information; and an arbitrary viewpoint image generation unit that reads out the 3D scene information and the viewpoint restriction information from the 3D scene information storage unit and generates a display image that represents how the displayed world appears as seen from an arbitrary viewpoint within the range of restrictions indicated by the viewpoint restriction information, by volume rendering using the 3D scene information.

8. An image processing method comprising: a step in which an application execution unit executes an application program and generates, at a predetermined rate, frames of display images that represent a three-dimensional display world in which a situation changes in response to user operations; and a step in which a system unit causes the application execution unit to generate learning images that represent the display world and are different from the display images, and performs a process in which 3D scene information that represents three-dimensional information about the display world is generated by machine learning using the learning images as training data and used for display, and in which a viewpoint set for the display world is restricted based on viewpoint restriction information associated with the application program.

9. A computer program that causes a computer to realize the following functions: a function for executing an application program and generating display image frames at a predetermined rate that represent a three-dimensional display world in which the situation changes in response to user operations; a function for causing a function for generating display image frames to generate learning images that represent the display world and are different from the display images, and a function for generating 3D scene information that represents three-dimensional information of the display world through machine learning using the learning images as training data, and a function for using the information for display in the process to restrict a viewpoint set for the display world based on viewpoint restriction information associated with the application program.

10. A data structure for 3D scene information for display, characterized in that it corresponds to data of 3D scene information consisting of a neural network that represents three-dimensional information of the displayed world, and viewpoint restriction information that indicates restriction information imposed on an arbitrary viewpoint when the data is read from a storage device together with the 3D scene information by an image processing device and a display image that represents the appearance of the displayed world as seen from the arbitrary viewpoint is generated by volume rendering using the 3D scene information.

Citation Information

Patent Citations

  • Image rendering method and apparatus

    JP2022151746A

  • Artificial Intelligence (AI) controlled camera perspective generator and AI broadcaster

    JP2022545128A