Content processing device and content processing method

The content processing device addresses the challenge of applying machine learning to dynamic user-operated content by generating and storing 3D scene information using a main image generating unit and a 3D scene information generating unit, enabling high-quality, versatile 3D scene representation and trophy creation.

WO2025094514A1PCT designated stage expired Publication Date: 2025-05-08SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/032292
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-09-10
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing technologies face challenges in applying machine learning to content where the display world situation changes dynamically based on user operations, such as determining the timing for acquiring learning images and utilizing learned information effectively.

Method used

A content processing device equipped with a main image generating unit, a trophy presentation determining unit, a 3D scene information generating unit, and a 3D scene information storage unit, which generates and stores 3D scene information using machine learning, with the main image representing the state seen from multiple viewpoints used as teacher data.

Benefits of technology

Enables easy application of machine learning to content with dynamic display world changes, allowing for the generation of high-quality 3D scene information that can be viewed from various directions and even used to create real trophies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024032292_08052025_PF_FP_ABST
    Figure JP2024032292_08052025_PF_FP_ABST
Patent Text Reader

Abstract

In the present invention, in a game play phase 210, a content server: determines that the time when the circumstances of a displayed world corresponding to a user operation satisfies a predetermined condition is the timing at which to present a trophy (S10); collects a training image representing a scene being displayed at that point in time (S12); and generates 3D scene information 220 by machine learning (S14). In a trophy appreciation phase 212, the content server generates and outputs a display image representing the trophy from an arbitrary viewpoint by using the 3D scene information 220 upon request from a user (S16). An order for a real-life replica of the trophy is also accepted (S18).
Need to check novelty before this filing date? Find Prior Art

Description

Content processing device and content processing method

[0001] The present invention relates to a content processing device and a content processing method for processing content such as electronic games.

[0002] With the recent expansion of communication networks and advances in image processing technology, it has become possible to enjoy a wide variety of electronic content regardless of the viewing environment. For example, in the field of electronic games, a system has become widespread in which a server collects information related to the status of each client device, such as the content of user operations and location information, and distributes image data that reflects this information as needed, allowing multiple players to participate in the same game regardless of location.

[0003] Meanwhile, in recent years, advances in machine learning technologies such as deep learning have made it possible to acquire various types of information from images. For example, NeRF (Neural Radiance Fields) is a method for representing three-dimensional space using a neural network. NeRF is a method for representing the volume density and radiance of an object in three-dimensional space as a five-dimensional function consisting of position coordinates and direction using a neural network. For example, if an NeRF representation is obtained based on images of an object captured from multiple directions, it is possible to represent the appearance of the object as seen from any viewpoint using volume rendering (see, for example, Non-Patent Document 1).

[0004] Ben Mildenhall and five others, "NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis," Communications of the ACM, January 2022, Vol. 65, No. 1, pp. 99-106

[0005] While image processing using machine learning as described above makes it possible to obtain highly flexible images from limited information, it requires learning using appropriate and sufficient images, which limits the scope of application. For example, in the case of content in which the scene to be displayed changes in real time in response to user operations, there are issues such as when to acquire learning images for a scene that changes from moment to moment and how to use the learned information, making it difficult to implement.

[0006] The present invention was made in consideration of these problems, and its purpose is to easily apply machine learning to content in which the state of the displayed world changes in response to user operations, thereby realizing new functions.

[0007] To solve the above problems, one aspect of the present invention relates to a content processing device, which includes: a main image generation unit that generates main image frames at a predetermined rate, each representing a three-dimensional display world, the state of which changes in response to user operations on content being executed, as viewed from a display viewpoint; a trophy presentation determination unit that determines the timing for presenting a trophy when the state of the display world satisfies a predetermined condition due to user operations; a 3D scene information generation unit that generates 3D scene information representing three-dimensional information of a scene, using machine learning to generate a trophy for at least a portion of the scene being displayed at the presentation timing, and main images representing the trophy as viewed from multiple viewpoints as training data; and a 3D scene information storage unit that stores the 3D scene information in association with the user performing the user operations.

[0008] Another aspect of the present invention relates to a content processing method, which includes the steps of: generating, at a predetermined rate, frames of a main image showing a three-dimensional display world, the state of which changes in response to a user operation on content being executed, as seen from a display viewpoint; determining a time when the state of the display world satisfies a predetermined condition due to the user operation as the timing for presenting a trophy; generating 3D scene information showing three-dimensional information of the scene by machine learning using, as training data, main images showing the trophy as seen from a plurality of viewpoints, with at least a part of the scene being displayed at the presentation timing; and storing the 3D scene information in a storage device in association with a user performing the user operation.

[0009] Any combination of the above components, and any transformation of the present invention into a method, device, system, computer program, data structure, recording medium, etc., are also valid aspects of the present invention.

[0010] According to the present invention, machine learning can be easily applied to content in which the state of the displayed world can change in response to user operations, thereby realizing new functions.

[0011] FIG. 1 is a diagram showing an example of the configuration of a content processing system to which the present embodiment can be applied. FIG. 2 is a diagram showing a schematic diagram of the relationship between a display world and a display image in an electronic game assumed to be processed in the present embodiment. FIG. 3 is a diagram showing an overview of the processing flow in the present embodiment. FIG. 4 is a diagram showing the internal circuit configuration of a client terminal in the present embodiment. FIG. 5 is a diagram showing the configuration of functional blocks of a client terminal and a content server in the present embodiment. FIG. 6 is a diagram showing a schematic diagram of a sequence of images generated in a game play phase in the present embodiment. FIG. 7 is a diagram showing a schematic diagram of an example of an image displayed when it is decided to present a trophy in the present embodiment. FIG. 8 is a diagram showing a schematic diagram of an example of the processing flow performed after the presentation of a trophy in the present embodiment.

[0012] 1 shows an example of the configuration of a content processing system to which this embodiment can be applied. The content processing system 1 includes client terminals 10a, 10b, and 10c that display images of electronic games and the like in response to user operations, and a content server 20 that provides image data used for display. Input devices 14a, 14b, and 14c for user operations and display devices 16a, 16b, and 16c for displaying images are connected to the client terminals 10a, 10b, and 10c, respectively. Communication can be established between the client terminals 10a, 10b, and 10c and the content server 20 via a network 8 such as a WAN (World Area Network) or a LAN (Local Area Network).

[0013] The client terminals 10a, 10b, and 10c may be connected to the display devices 16a, 16b, and 16c and the input devices 14a, 14b, and 14c either wired or wirelessly. Alternatively, two or more of these devices may be integrated. For example, in the figure, the client terminal 10b is connected to a head-mounted display, which is the display device 16b. The head-mounted display can change the field of view of the displayed image by the movement of the user wearing it on their head, so it also functions as the input device 14b.

[0014] The client terminal 10c is a mobile terminal, tablet terminal, or the like, and is integrally configured with a display device 16c and an input device 14c, which is a touchpad covering the display device 16c's screen. Thus, the external shape and connection form of the illustrated devices are not limited. The number of client terminals 10a, 10b, and 10c and content servers 20 connected to the network 8 is also not limited. Hereinafter, the client terminals 10a, 10b, and 10c will be collectively referred to as client terminals 10, the input devices 14a, 14b, and 14c as input device 14, and the display devices 16a, 16b, and 16c as display device 16.

[0015] The input device 14 is a general input device such as a controller, keyboard, mouse, touchpad, or joystick, and receives user operations and supplies the operations to the client terminal 10. The input device 14 may also be various sensors such as a motion sensor or camera provided in a head-mounted display, mobile terminal, or tablet terminal, and may supply the sensor data to the client terminal 10. The display device 16 may be a general display such as a liquid crystal display, plasma display, organic EL display, wearable display, or projector, and displays an image output from the client terminal 10.

[0016] The content server 20 provides content data that accompanies image display to the client terminal 10. In this embodiment, the content server 20 basically generates video and audio data for the content and immediately transmits the data to the client terminal 10 to realize streaming.

[0017] In this case, the content server 20 may sequentially acquire information on user operations on the input device 14 or sensor data acquired by various sensors from the client terminal 10 and reflect the information in images and sounds. This allows multiple users to participate in the same game or communicate in the virtual world. However, the configuration of the image display system is not limited to that shown in the figure. For example, the image generation entity is not limited to the content server 20, but may be the client terminal 10 itself, or both may work together.

[0018] 2 is a schematic diagram showing the relationship between the display world and the displayed image in an electronic game assumed to be the processing target in this embodiment. As mentioned above, the main processing may be performed by either the content server 20 or the client terminal 10, or by both working together, and therefore, the processing will be described here as being performed by a "content processing device" without making a distinction between the two. In this embodiment, for example, an electronic game is assumed in which users adventure and compete as players in a virtual world defined by a three-dimensional space.

[0019] In the illustrated example, a virtual three-dimensional space in which an enemy character 202 and the like exist is the display world 200. A user can move around the display world and fight the enemy character 202 by performing user operations via the input device 14. The content processing device generates a display image 204 at a predetermined rate while appropriately changing the display world in response to user operations, and outputs the display image 204 to the display device 16. At this time, the content processing device acquires the position of a viewpoint and the direction of line of sight that the user can operate, and determines the field of view of the display image 204 accordingly. Hereinafter, the state of objects in the display world 200 that exist within the field of view of the display image 204 will be referred to as a "scene." The position of the viewpoint and the direction of line of sight relative to a scene may also be collectively referred to simply as a "viewpoint."

[0020] In the case of a multiplayer game, the content processing device acquires user operation details from multiple users in parallel and reflects them in the display world 200. Note that the display image 204 shown in the figure is a so-called first-person perspective image in which the display world 200 is viewed from the perspective of a character controlled by a user, but this is not intended to limit the present embodiment. For example, a so-called third-person perspective image in which the character controlled by the user is included in the field of view may also be used. In this embodiment, in such an electronic game, 3D scene information representing the scene being displayed is generated by machine learning when the situation of the game or the display world satisfies a predetermined condition, such as when a user achieves a significant achievement.

[0021] By saving the 3D scene information as a memento, the user can view the moment when they achieved success in the game from various angles at any time later. The 3D scene information can also be used to create a real object from the scene using a 3D printer or the like. Hereinafter, the object for which 3D scene information is generated will be referred to as a "trophy," and the process of generating and saving 3D scene information associated with a user will be referred to as "presenting a trophy." The timing of presenting a trophy is specified within the game program. For example, the content processing device presents a trophy at a scene that gives the user a sense of accomplishment, such as when a predetermined number of enemies are defeated in the displayed world, when a lap time in a car race falls below a predetermined time, or when a set stage is cleared.

[0022] The content processing device collects training images representing the scene viewed from various directions to create 3D scene information for the scene displayed at the time the decision to present a trophy is made. When applying NeRF to machine learning, the content processing device first inputs the viewpoint information, i.e., the virtual viewpoint position and line of sight direction, determined when generating the training images. Using the corresponding training images as training data, the content processing device performs regression using a multilayer perceptron (MLP) to obtain data representing the three-dimensional information of the scene. This data is a neural network that inputs five-dimensional parameters consisting of position coordinates (x, y, z) and a direction vector d(θ, φ) in three-dimensional space, and outputs volume density σ and color information c (RGB) of the three primary colors.

[0023] In this embodiment, the neural network data is referred to as "3D scene information." However, any technology that can estimate three-dimensional information from multiple two-dimensional images can be introduced, not just NeRF, and the representation format of the 3D scene information is not limited. To obtain accurate 3D scene information, it is desirable to collect images of the target scene viewed from as many viewpoints as possible as learning images. Therefore, once it has been determined that a trophy will be presented, the content processing device intensively generates images from a variety of viewpoints by setting multiple viewpoints for generating learning images for the scene.

[0024] The content processing device uses the stored 3D scene information to generate a separate display image, allowing the user to view the trophy from any direction. By using the 3D scene information, it is possible to display the scene as seen from any viewpoint with high quality, with a relatively low load. When NeRF is applied, the content processing device generates a ray r that passes through the pixels of the viewscreen from the display viewpoint, and performs volume rendering by integrating color along that direction to determine the pixel value C(r) of the display image as follows:

[0025]

[0026] where t n, t f are the proximal and distal ends of the ray r, respectively, and T(t) is the cumulative transmittance in the direction of the ray, which can be expressed as follows:

[0027]

[0028] Regarding NeRF, in addition to the basic technique disclosed in, for example, Non-Patent Document 1, various improved techniques have been proposed, and any of these may be adopted in this embodiment. Therefore, detailed explanations will be omitted here. The stored 3D scene information represents information such as the shape, position, and orientation of an object in a target scene. Therefore, the content processing device may use the stored 3D scene information to provide an interface for a service that creates a physical version of the object.

[0029] 3 shows an overview of the processing flow in this embodiment. This embodiment is realized in two periods: a game play phase 210 and a trophy appreciation phase 212. The game play phase 210 is the period during which the user plays the game. During this period, the content processing device, for example, the content server 20, determines the timing of the presentation of a trophy based on the content of the game (S10).

[0030] In response, the content server 20 treats the currently displayed scene and the objects contained therein as trophies and collects training images depicting the trophies from multiple viewpoints (S12).The content server 20 then performs machine learning using the training images as training data to generate 3D scene information 220 depicting the trophies (S14).The trophy appreciation phase 212 begins when the user requests to appreciate the trophies at any time, such as during or after game play is interrupted.During this period, the content processing device, for example, the content server 20, uses the saved 3D scene information 220 to generate trophy images and output them for display (S16).

[0031] Alternatively, the content server 20 may receive an order from the user to have the presented trophy made into an actual product (S18). For example, the content server 20 may display an order screen for the actual product along with the image of the trophy displayed in S16. When the user places an order for the actual product, the content server 20 receives the order, converts the 3D scene information 220 of the trophy into mesh data, and then places an appropriate order to have the trophy made into an actual product using a 3D printer or the like. As a result, the actual presented trophy is sent to the user at a later date.

[0032] 4 shows the internal circuit configuration of the client terminal 10. The client terminal 10 includes a CPU (Central Processing Unit) 122, a GPU (Graphics Processing Unit) 124, and a main memory 126. These components are interconnected via a bus 130. An input / output interface 128 is also connected to the bus 130. Connected to the input / output interface 128 are a communication unit 132 including a peripheral device interface such as a USB or a network interface for a wired or wireless LAN, a storage unit 134 such as a hard disk drive or nonvolatile memory, an output unit 136 that outputs data to the display device 16, an input unit 138 that inputs data from the input device 14, and a recording medium drive unit 140 that drives a removable recording medium such as a magnetic disk, optical disk, or semiconductor memory.

[0033] The CPU 122 executes an operating system stored in the storage unit 134 to control the entire client terminal 10. The CPU 122 also executes various programs read from a removable recording medium and loaded into the main memory 126, or downloaded via the communication unit 132. The GPU 124 has the functions of a geometry engine and a rendering processor, performs drawing processing in accordance with drawing commands from the CPU 122, and stores display images in a frame buffer (not shown). The GPU 124 then converts the display images stored in the frame buffer into video signals and outputs them to the output unit 136. The main memory 126 is composed of RAM (Random Access Memory) and stores programs and data required for processing. The content server 20 may also have a similar internal circuit configuration.

[0034] FIG. 5 shows the functional block configuration of the client terminal 10 and the content server 20 in this embodiment. The illustrated functional blocks can be realized in hardware terms using the CPU, GPU, various memories, and other components shown in FIG. 4 , and in software terms using programs loaded into memory from a recording medium or the like, which perform various functions such as data input, data storage, image processing, and communication. Therefore, those skilled in the art will understand that these functional blocks can be realized in various forms using only hardware, only software, or a combination thereof, and are not limited to any one of these. Also, while the content server 20 is primarily responsible for image processing in this diagram, at least some of this processing may be performed by the client terminal 10.

[0035] The client terminal 10 includes an input information acquisition unit 50 that acquires input information such as user operations, an image data acquisition unit 52 that acquires image data from the content server 20, and an output unit 54 that outputs display image data. The input information acquisition unit 50 acquires the content of user operations from the input device 14 at any time. User operations include selecting or starting a game, and inputting commands for the game being played. The input information acquisition unit 50 also accepts operations related to appreciating a presented trophy and operations to order the actual trophy.

[0036] The input information acquisition unit 50 also acquires information about the display viewpoint from the input device 14 or the head-mounted display at any time or at predetermined time intervals. Technology for detecting the position and posture of the head of a user wearing a head-mounted display and acquiring information about the display viewpoint based on this detection is well known, and this technology may also be applied to this embodiment. The display viewpoint here includes a display viewpoint for viewing game images during play, as well as a display viewpoint for viewing a presented trophy. The input information acquisition unit 50 supplies the acquired information to the content server 20 as appropriate.

[0037] The image data acquisition unit 52 acquires display image data from the content server 20. Here, the display image data may include data of an image of the game being played, data of an image notifying the presentation of a trophy, data of a display image of the presented trophy, and data of an image accepting an order for an actual trophy. The output unit 54 outputs the display image acquired by the image data acquisition unit 52 to the display device 16 for display.

[0038] The content server 20 includes an input information acquisition unit 70 that acquires input information from the client terminal 10, an application execution unit 72 that executes the game application, a trophy presentation processing unit 74 that performs processing related to the presentation of trophies, a 3D scene information generation unit 76 that generates 3D scene information data, a 3D scene information storage unit 78 that stores the generated 3D scene information data, a trophy image generation unit 80 that generates trophy images, a real item order processing unit 82 that processes orders for real trophies, and an image data transmission unit 84 that transmits display image data to the client terminal 10.

[0039] The input information acquisition unit 70 acquires the content of user operations and information on display viewpoints from the client terminal 10 at any time or at predetermined time intervals. The input information acquisition unit 70 supplies the acquired information to the application execution unit 72. The application execution unit 72 processes the game application based on the content of the user operations. The application execution unit 72 includes a game image generation unit 86 and a trophy presentation determination unit 88.

[0040] During the game play phase, the game image generation unit 86 generates game images of a field of view corresponding to a display viewpoint at a predetermined rate. The trophy presentation determination unit 88 determines whether to present a trophy as the game progresses. Specifically, the trophy presentation determination unit 88 stores preset trophy presentation conditions in an internal memory and constantly checks the content of play and other factors against the presentation conditions to detect situations in which a trophy should be presented. Once it has been determined that a trophy should be presented, the game image generation unit 86 generates images representing the current scene as viewed from various viewpoints as learning images. At this time, the application execution unit 72 may pause the game.

[0041] The trophy presentation processing unit 74 performs various processes related to the presentation of trophies. Specifically, the trophy presentation processing unit 74 includes a presentation determination detection unit 90, a viewpoint generation unit 92, and a trophy presentation image generation unit 94. The presentation determination detection unit 90 detects that the trophy presentation determination unit 88 of the application execution unit 72 has decided to present a trophy. For example, the presentation determination detection unit 90 receives a notification from the trophy presentation determination unit 88 that a decision has been made to present a trophy. Alternatively, the presentation determination detection unit 90 may periodically inquire of the trophy presentation determination unit 88 as to whether a decision has been made to present a trophy.

[0042] When a decision to present a trophy is detected, the viewpoint generation unit 92 generates a plurality of viewpoints for acquiring learning images representing the scene of the display world currently being displayed. The viewpoint generation unit 92 supplies the generated viewpoints to the application execution unit 72 in the same format as the normal display viewpoints. As a result, the game image generation unit 86 generates images representing the scene as seen from the generated viewpoints as learning images.

[0043] The trophy presentation image generation unit 94 generates an image notifying the user that a trophy has been presented. This image may represent the scene displayed when the presentation was confirmed. The trophy presentation image generation unit 94 may then accept a user operation to select an object or area to be used as a trophy from the displayed image. This allows the user to create a trophy tailored to their preferences by excluding the background or leaving only specific objects. In this case, the trophy presentation image generation unit 94 supplies information indicating the object or area designated as a trophy to the 3D scene information generation unit 76.

[0044] In the illustrated example, it is assumed that the application execution unit 72 basically generates game images based on viewpoint information supplied from the input information acquisition unit 70. In this case, the viewpoint generation unit 92 generates and supplies viewpoint information in the same format as the viewpoint information supplied from the input information acquisition unit 70, so that the application execution unit 72 can generate learning images through normal processing without distinguishing between a true display viewpoint and a viewpoint generated by the viewpoint generation unit 92. As a result, this embodiment can be easily applied even to conventional content that does not support machine learning.

[0045] However, this embodiment is not limited to this. An API (Application Programming Interface) having a viewpoint generation function may be prepared and specified in the application program, so that the application execution unit 72 is provided with the viewpoint generation unit 92. Similarly, at least a portion of the processing of the trophy presentation image generation unit 94 may be performed by the application execution unit 72 via the API. If the game was paused due to the presentation of a trophy, the application execution unit 72 resumes the game once all learning images have been generated. Alternatively, the application execution unit 72 may resume the game once the user closes the image on the client terminal 10 notifying them of the presentation of a trophy.

[0046] The 3D scene information generation unit 76 acquires the training image generated by the application execution unit 72 and generates 3D scene information for the trophy using the machine learning described above. When a user selects an object or area to be used as a trophy in the trophy presentation scene, the 3D scene information generation unit 76 extracts only the selected object or area from the training image and uses it for machine learning. Alternatively, the 3D scene information generation unit 76 may extract objects or areas selected by itself based on predetermined criteria from the training image.

[0047] For example, when a trophy is presented to a virtual user for defeating an enemy character, the 3D scene information generation unit 76 may treat the enemy character at the moment of defeat as a trophy and extract its image from the training image using well-known techniques such as object recognition. In the case of a third-person perspective game image, the 3D scene information generation unit 76 may treat both the virtual user and the enemy character as trophies and extract them from the training image. The 3D scene information storage unit 78 stores 3D scene information of the trophy generated by the 3D scene information generation unit 76. The 3D scene information storage unit 78 stores the 3D scene information in association with the identification information of the recipient user and information describing the scene. This allows appropriate 3D scene information to be read out from the information when the user requests to view the trophy.

[0048] During the trophy appreciation phase, when a user requests appreciation of an acquired trophy, the trophy image generation unit 80 uses the 3D scene information stored in the 3D scene information storage unit 78 to generate a display image representing the trophy by the volume rendering described above. At this time, the trophy image generation unit 80 acquires a display viewpoint from the input information acquisition unit 70 and generates a display image while changing the viewpoint of the trophy accordingly. During the game play phase, the image data transmission unit 84 transmits data on game images generated by the game image generation unit 86 and data on images notifying the presentation of a trophy to the client terminal 10. During the trophy appreciation phase, the image data transmission unit 84 also transmits data on the trophy image generated by the trophy image generation unit 80 to the client terminal 10.

[0049] The real-item order processing unit 82 receives an order from a user to have a real trophy made, and places the order using the 3D scene information stored in the 3D scene information storage unit 78. The order destination may be another server (not shown) that provides a service for generating a three-dimensional object from mesh data using a 3D printer or the like. The screen for receiving the order is displayed, for example, by the trophy image generation unit 80, which adds the order information to the display image of the trophy.

[0050] In response to an order, the real-item order processor 82 acquires three-dimensional polygon data from a neural network representing 3D scene information and converts it into general mesh data. In this process, the real-item order processor 82 acquires the coordinates of a point cloud on the object's surface, for example, using the original 3D scene information to generate rays and perform rendering calculations in the same way as when generating a display image. Then, based on the coordinates of the point cloud, mesh data of the object is generated using a method such as the marching cubes method (see, for example, Non-Patent Document 1). The real-item order processor 82 matches the mesh data of the trophy with the name, address, and other information of the user who placed the order, and places an order for the trophy to be made into a real object.

[0051] The actual product order processing unit 82 may only convert the 3D scene information of the trophy into mesh data and send the converted data to the client terminal 10. This allows the client terminal 10 to directly order an actual trophy, or the user of the client terminal 10 to generate an actual trophy using a 3D printer or the like. The mesh data converted from the 3D scene information can also be used when displaying an image of the trophy.

[0052] In this case, the trophy image generation unit 80 reads the 3D scene information from the 3D scene information storage unit 78, converts it into mesh data, and uses it to generate a display image. The trophy image generation unit 80 may store the mesh data itself as trophy data in a storage unit (not shown). On the other hand, mesh data allows images to be displayed from any viewpoint using a general viewer. Therefore, the trophy image generation unit 80 may transmit the trophy mesh data to the client terminal 10, allowing the client terminal 10 to generate a display image. This allows users to view their acquired trophies from various viewpoints, even in offline environments.

[0053] 6 is a schematic diagram showing a sequence of images generated during the gameplay phase. The diagram shows the relationship between the viewpoints recognized or generated by the content server 20 and the destinations of each frame generated thereby, with the horizontal axis representing time. The content server 20 basically generates frames of display images (e.g., frame 232) at a predetermined rate corresponding to the display viewpoints (e.g., display viewpoint 230) indicated by white circles, and transmits them to the client terminal 10.

[0054] During this process, when it is decided at time t1 to present a trophy, the trophy presentation processing unit 74 begins processing related to the trophy presentation. Specifically, the content server 20 generates a viewpoint indicated by a black circle (e.g., viewpoint 234) and generates a corresponding training image (e.g., training image 236). As shown in the figure, the rate at which training images are generated may be higher than the rate at which the display frames are generated, depending on the processing power of the content server 20.

[0055] As an example, if the viewpoint generation unit 92 prepares 300 viewpoints and the game image generation unit 86 processes them sequentially, 300 training images can be generated in a few seconds. The content server 20 pauses the game while generating the training images, generates an image (e.g., image 238) indicated by a shaded area to notify the client terminal 10 that a trophy will be presented, and transmits the image to the client terminal 10. The image notifying the presentation is displayed until time t2, when the content server 20 has finished generating a predetermined number of training images. The content server 20 generates 3D scene information for the trophy based on the training images generated up to time t2 and stores the information in the 3D scene information storage unit 78. The content server 20 resumes the game at time t2, generates display image frames at a predetermined rate to correspond to the latest display viewpoint, and transmits the frames to the client terminal 10.

[0056] In the illustrated example, an image notifying the presentation of a trophy is displayed while the training images are being generated, and the training images are not sent to the client terminal 10 but are used only for machine learning. Alternatively, during the period in which the training images are being generated, at least some of the training images may be sent to the client terminal 10 and displayed. The viewpoints generated to display the training images are generated so that the target scene is viewed from various directions. Therefore, by generating the viewpoints in an appropriate order and sequentially using the training images obtained thereby as display images, a dynamic video can be displayed that appears as if it were shot while moving around the target scene. Such an image may be displayed as part of the image notifying the presentation of a trophy.

[0057] Furthermore, images generated for display as game images may be used as learning images. For example, in the case of a multiplayer game, each player may view the same scene from a different viewpoint. Therefore, the content server 20 may store the display image data transmitted to each client terminal 10 for a predetermined period of time, and when it is decided to award a trophy to one of the users, extract a display image representing the target scene and use it for machine learning. Therefore, the images used as learning images may be images taken before the time it was decided to award a trophy.

[0058] 7 shows a schematic example of an image displayed when it has been decided that a trophy will be presented. Immediately after it has been decided that a trophy will be presented, the trophy presentation image generation unit 94 generates and displays an image 240 informing the player of this. In the figure, a message 242 indicating the content of the notification is superimposed on the game image at that time. Meanwhile, the game image generation unit 86 generates a learning image corresponding to the viewpoint generated by the viewpoint generation unit 92.

[0059] The bottom of the figure shows a schematic representation of multiple viewpoints (e.g., viewpoint 246) set for a scene 244 in a three-dimensional image world, and the screen surface (e.g., screen surface 248) of the image set for each viewpoint. However, since the viewpoint generation unit 92 actually generates learning images from all directions, more viewpoints may be set, for example, by evenly distributing viewpoints on the surface of a hemisphere centered on the scene. The game image generation unit 86 renders images on the screen surface corresponding to each viewpoint using ray tracing or the like. The image data transmission unit 84 may extract images of a consecutive series of viewpoints from the images generated in this manner, transmit them to the client terminal 10, and display them sequentially.

[0060] In the figure, the arrows connecting the viewpoints indicate that a series of viewpoint images that circle the periphery of the scene 244 are extracted. By displaying the images in the order indicated by the arrows, the user can enjoy special images that depict the target scene from viewpoints that circle the image world, even during the generation of the learning images. Note that the display of such learning images may also be part of the image 240 announcing the presentation of a trophy.

[0061] 8 shows a schematic example of the processing flow after a trophy is presented. First, the trophy presentation image generation unit 94 generates and displays an image 250 prompting the user to select the object to be presented as a trophy, such as after the trophy presentation notification has been sent. In the figure, a message 252 prompting the user to select the object to be presented and a cursor 254 as a selection tool are superimposed on the game image at the time of presentation. When the user selects the desired object, the 3D scene information generation unit 76 uses existing technology, such as object recognition, to extract an image of the selected object from the training image and use it for training.

[0062] As a result, 3D scene information 256 for the selected object 258 is generated and stored in the 3D scene information storage unit 78. Note that the 3D scene information 256 is actually neural network data, as described above. In trophy appreciation mode, the trophy image generation unit 80 generates a display image 260 representing a trophy in response to a user request, and transmits it to the client terminal 10 for display. While looking at the display image 260, the user can view the trophy from any viewpoint by, for example, rotating the trophy as indicated by the arrow.

[0063] At this time, the trophy image generator 80 continues to generate a display image from a display viewpoint corresponding to the operation using the 3D scene information. Alternatively, as described above, the trophy image generator 80 may convert the 3D scene information formed by the neural network into mesh data and use this to generate a display image. Alternatively, the image data transmitter 84 may transmit the mesh data to the client terminal 10. In this case, the client terminal 10 itself can generate and display the display image using a general viewer.

[0064] The real-item order processor 82 also places an order for the creation of a real trophy in response to a user operation to order a real trophy. At this time, the real-item order processor 82 converts the 3D scene information from the neural network into mesh data, which is used as model data at the time of ordering, enabling the real item to be created by standard means. As a result, a real trophy 262 created using a 3D printer or the like is delivered to the user.

[0065] According to the present embodiment described above, the content server 20 detects the arrival of a timing set within the game application and presents the scene displayed at that time as a trophy. At this time, the content server 20 performs machine learning by intensively generating training images, and generates 3D scene information representing the scene. This allows the user to view the state of the object at the moment when it was decided to present the trophy from various angles at any time thereafter.

[0066] The content server 20 also converts the 3D scene information of the trophy, which is made up of a neural network, into mesh data so that an actual trophy can be created. The content server 20 may also transmit the mesh data itself to the client terminal 10. This allows the user to obtain an actual trophy created using a 3D printer or the like, or to view an image of the trophy offline from any viewpoint.

[0067] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention.

[0068] As described above, the present invention can be used in various information processing devices such as content processing devices, content servers, game devices, mobile terminals, and personal computers, as well as content service providing systems including any of these.

[0069] 1 Image display system, 10 Client terminal, 14 Input device, 16 Display device, 20 Content server, 50 Input information acquisition unit, 52 Image data acquisition unit, 54 Output unit, 70 Input information acquisition unit, 72 Application execution unit, 74 Trophy presentation processing unit, 76 3D scene information generation unit, 78 3D scene information storage unit, 80 Trophy image generation unit, 82 Physical order processing unit, 84 Image data transmission unit, 86 Game image generation unit, 88 Trophy presentation determination unit, 90 Presentation determination detection unit, 92 Viewpoint generation unit, Trophy presentation image generation unit, 122 CPU, 124 GPU, 126 Main memory.

Claims

1. A content processing device comprising: a main image generation unit which generates, at a predetermined rate, frames of a main image which shows a three-dimensional display world, the situation of which changes in response to user operations on content being executed, as seen from a display viewpoint; a trophy presentation determination unit which determines the time when the situation of the display world satisfies predetermined conditions due to the user operations as the timing for presenting a trophy; a 3D scene information generation unit which generates 3D scene information which shows three-dimensional information of a scene by machine learning using as training data the main images which show the trophy as seen from a plurality of viewpoints, and a 3D scene information storage unit which stores the 3D scene information in association with a user performing the user operations.

2. The content processing device described in claim 1, further comprising a viewpoint generation unit that generates a viewpoint for generating an image to be used in the machine learning independently of the display viewpoint at the timing of presenting the trophy, wherein the main image generation unit generates a frame of the main image corresponding to the added viewpoint.

3. A content processing device as described in claim 1 or 2, further comprising a trophy image generation unit that uses the 3D scene information to generate a display image showing the scene designated as the trophy as viewed from an arbitrary viewpoint.

4. A content processing device according to claim 1 or 2, further comprising an image data transmission unit that converts the 3D scene information into mesh data and transmits the mesh data to a client terminal used by the user.

5. A content processing device as described in claim 1 or 2, further comprising an actual item order processing unit that receives an order for an actual item of the 3D scene information from the user and performs ordering processing for the actual item using mesh data obtained by converting the 3D scene information.

6. A content processing device as described in claim 1 or 2, further comprising a trophy presentation image generation unit that generates an image that accepts selection of an area or object to be the trophy from among images of the scene being displayed at the presentation timing and presents the image to the user, wherein the 3D scene information generation unit extracts an image of the area or object selected by the user from the image to be used for machine learning, and then performs machine learning.

7. A content processing device as described in claim 1 or 2, characterized in that the 3D scene information generation unit uses an area extracted from the image used for machine learning based on a predetermined criterion for machine learning.

8. A content processing device according to claim 2, characterized in that the main image generation unit generates images to be used in the machine learning at a frame rate higher than that of the displayed image.

9. A content processing device as described in claim 2, further comprising an image data transmission unit which, at the presentation timing, sequentially outputs, as display images, frames of the main image corresponding to the added viewpoint generated by the main image generation unit, which correspond to a predetermined series of viewpoints.

10. A content processing device according to claim 1 or 2, characterized in that the 3D scene information generating unit generates a neural network using NeRF (Neural Radiance Fields) as the 3D scene information.

11. A content processing method comprising the steps of: generating, at a predetermined rate, frames of a main image showing a three-dimensional display world, the situation of which changes in response to user operations on running content, as seen from a display viewpoint; determining the time when the situation of the display world satisfies predetermined conditions due to the user operations as the timing for presenting a trophy; generating 3D scene information showing three-dimensional information of the scene by machine learning using as training data the main image, which shows the trophy as seen from a plurality of viewpoints, at least a part of the scene being displayed at the time of presentation; and storing the 3D scene information in a storage device in association with the user performing the user operations.

12. A computer program that causes a computer to realize the following functions: a function for generating, at a predetermined rate, frames of a main image showing a three-dimensional display world, the situation of which changes in response to user operations on running content, as seen from a display viewpoint; a function for determining the timing for presenting a trophy when the situation of the display world satisfies predetermined conditions due to the user operations; a function for generating 3D scene information showing three-dimensional information of the scene by machine learning using as training data the main image, which shows the trophy as seen from multiple viewpoints, and a function for storing the 3D scene information in a storage device in association with the user performing the user operations.

Citation Information

Patent Citations

  • Device and method of selecting object for 3D printing

    JP2016207203A

  • Image display method and image display device

    JP2017139725A

  • Information processing device and moving image editing method

    JP2021062173A

  • Bringing achievements to an offline world

    US20120277004A1