Information processing device and method, and information processing system
The system addresses high processing loads and costs in 3D content distribution by using interpolation processing to generate images with different viewpoints, reducing server requirements and maintaining image quality.
Patent Information
- Application Number
- PCT/JP2025/016821
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-08
- Publication Date
- 2025-11-27
AI Technical Summary
Existing 3D content distribution systems face high processing loads and costs due to the need for multiple servers to ensure real-time performance in multi-user, multi-viewpoint simultaneous connections, which is challenging to manage efficiently.
An information processing system that employs interpolation processing to generate images with different viewpoints or times based on initial rendered images, reducing the number of required 3D rendering processes and optimizing server load while maintaining image quality.
This approach reduces processing load on servers, decreases the number of necessary rendering servers, and maintains high-quality video distribution, thereby controlling costs and ensuring smooth user experiences in 3D content delivery.
Smart Images

Figure JP2025016821_27112025_PF_FP_ABST
Abstract
Description
Information processing device and method, and information processing system
[0001] The present technology relates to an information processing device and method, and an information processing system, and more particularly to an information processing device and method, and an information processing system that are capable of reducing the processing load on a server while suppressing costs.
[0002] 2. Description of the Related Art Conventionally, techniques have been proposed for providing content such as video in a three-dimensional (3D) virtual space (hereinafter also referred to as 3D content).
[0003] In a service that provides 3D content, for example, a rendering process called 3D rendering is performed on a server that distributes the 3D content, and the resulting video data is transmitted to a client via a network.
[0004] In addition, there are services that provide 3D content, such as a multi-user, multi-viewpoint simultaneous connection service in which multiple users can view the 3D content and each user can freely change their position and line of sight (orientation) in the virtual space.
[0005] In use cases such as multi-user, multi-viewpoint simultaneous connection services, responses to user input need to be reflected in the video immediately, but the load of 3D rendering processing is very high, making it difficult to ensure real-time performance.
[0006] Furthermore, as a technology related to 3D content, a technology has been proposed in which 3D rendering of an image to be presented to one user is performed by multiple servers, thereby reducing the processing load on the client side (see, for example, Patent Document 1).
[0007] Japanese Patent Application Laid-Open No. 2020-21394
[0008] In order to ensure real-time performance in services such as multi-user multi-view simultaneous connections, one possible method is to prepare multiple servers that perform 3D rendering processing and have these servers execute the processing in parallel, thereby distributing the processing and shortening the processing time. However, while this method can reduce the processing load per server, it requires preparing many servers, which increases costs.
[0009] The present technology has been developed in light of these circumstances, and is intended to reduce the processing load on the server while keeping costs down.
[0010] An information processing device according to a first aspect of the present technology includes an interpolation processing unit that generates, based on a first image generated by a rendering process and having a viewpoint position at a predetermined position in space, a second image by interpolation processing, the second image having a viewpoint position at a different position or time from that of the first image.
[0011] An information processing method according to a first aspect of the present technology includes an information processing device generating, by interpolation processing, a second image having a viewpoint position at a different location or time from that of a first image, based on a first image generated by rendering processing and having a viewpoint position at a predetermined position in space.
[0012] In a first aspect of the present technology, an information processing device generates, based on a first image generated by a rendering process and having a predetermined position in space as a viewpoint position, a second image by an interpolation process, having a viewpoint position at a different location or time from that of the first image.
[0013] An information processing system according to a second aspect of the present technology is an information processing system having a rendering server and a distribution server, wherein the distribution server comprises: a first transmitting unit that transmits to the rendering server a request to execute a rendering process to generate a first image having a viewpoint position that is a predetermined position in a space; a first receiving unit that receives the first image transmitted from the rendering server; and an interpolation processing unit that generates, based on the first image, a second image having a viewpoint position at a different position or time from that of the first image by interpolation processing; and the rendering server comprises: a second receiving unit that receives the execution request, an image generating unit that performs the rendering process in accordance with the execution request to generate the first image, and a second transmitting unit that transmits the first image to the distribution server.
[0014] In a second aspect of the present technology, in an information processing system having a rendering server and a distribution server, the distribution server transmits to the rendering server a request to execute a rendering process to generate a first video having a viewpoint position at a predetermined position in a space, receives the first video transmitted from the rendering server, and generates a second video having a viewpoint position or time different from that of the first video by an interpolation process based on the first video. Also, the rendering server receives the execution request, performs the rendering process in accordance with the execution request, generates the first video, and transmits the first video to the distribution server.
[0015] 1 is a diagram illustrating an example of the configuration of an information processing system. FIG. 1 is a diagram illustrating a distribution video. FIG. 2 is a diagram illustrating a 3D rendering process. FIG. 2 is a diagram illustrating a configuration of a server group. FIG. 3 is a diagram illustrating a configuration of a server group. FIG. 4 is a diagram illustrating an example of a distribution video. FIG. 4 is a diagram illustrating an example of the configuration of a rendering server. FIG. 5 is a diagram illustrating an example of the configuration of a distribution server. FIG. 6 is a diagram illustrating an example of the configuration of a client. FIG. 7 is a flowchart illustrating a distribution process. FIG. 8 is a flowchart illustrating a data generation process. FIG. 9 is a flowchart illustrating a content playback process. FIG. 10 is a diagram illustrating constraints for determining an actual rendering point. FIG. 11 is a diagram illustrating constraints for determining an actual rendering point. FIG. 12 is a diagram illustrating an example of determining an actual rendering point. FIG. 13 is a diagram illustrating an example of determining an actual rendering point. FIG. 14 is a diagram illustrating an example of determining an actual rendering point. FIG. 15 is a diagram illustrating an example of determining an actual rendering point. FIG. 10 is a diagram for explaining determination of an actual rendering point; FIG. 11 is a diagram for explaining determination of a rendering position; FIG. 12 is a diagram for explaining determination of a rendering position; FIG. 13 is a diagram for explaining determination of a rendering position; FIG. 14 is a diagram for explaining determination of a rendering position; FIG. 15 is a diagram for explaining determination of a rendering position; FIG. 16 is a diagram illustrating an example of the configuration of a computer.
[0016] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0017] First Embodiment Configuration Example of Information Processing System The present technology relates to a content distribution system that distributes, as 3D content, images and the like in a three-dimensional virtual space such as a metaverse in which multiple people can participate simultaneously.
[0018] The 3D content may consist of only video, or may consist of video and accompanying audio. In the following, an example of the 3D content delivered will be described, in which a user, more specifically, a user avatar, can freely move within a three-dimensional virtual space and freely change their line of sight (face direction).
[0019] FIG. 1 is a diagram showing an example of the configuration of an embodiment of an information processing system to which the present technology is applied.
[0020] The information processing system 11 shown in FIG. 1 is a content distribution system that generates and distributes 3D content in real time.
[0021] The information processing system 11 includes rendering servers 21-1 to 21-N, a distribution server 22, and clients 23-1 to 23-M.
[0022] In response to a request from the distribution server 22, the rendering servers 21-1 to 21-N generate an image for a specific time (frame) using a 3D rendering process (hereinafter simply referred to as the rendering process) with a specific position in the three-dimensional virtual space as the viewpoint position, and supply the resulting generated image data to the distribution server 22.
[0023] Hereinafter, when there is no need to particularly distinguish between the rendering servers 21-1 to 21-N, they will be simply referred to as the rendering servers 21. Note that when 3D content consists of video and audio, the rendering server 21 also performs audio rendering processing, but in the following, explanations of the audio rendering processing, etc. will be omitted as appropriate.
[0024] The distribution server 22 instructs each rendering server 21 to generate generated video data, and also aggregates the generated video data supplied from the rendering servers 21 to generate distribution video data for presenting 3D content.
[0025] For example, when generating distribution video data, a video is generated that has a different viewpoint position or time (playback time) in the three-dimensional virtual space from the video based on the generated video data, and video data that displays the generated videos (frames) at each time is used as distribution video data. Note that the generated video data may also be used as distribution video data as is. Furthermore, distribution video data is generated for each client (user) connected to the distribution server 22. In particular, distribution video data generated for a client is video data of a video whose viewpoint is the position in the three-dimensional virtual space of the user operating the client.
[0026] The distribution server 22 supplies the generated distribution video data to the clients 23-1 to 23-M.
[0027] Although an example in which one distribution server 22 is provided in the information processing system 11 will be described here, a plurality of distribution servers 22 may be provided in the information processing system 11 .
[0028] The clients 23-1 to 23-M are information processing devices operated by users who view 3D content. Hereinafter, when there is no need to particularly distinguish between the clients 23-1 to 23-M, they will also be simply referred to as clients 23.
[0029] For example, the client 23 may be any of various terminals such as a smartphone, a personal computer, a head mounted display (HMD), or a game device.
[0030] The user can freely change the position and orientation of the user, more specifically, the avatar representing the user, in the three-dimensional virtual space by inputting operations to the client 23. The client 23 generates client attribute information according to the user's input operations and transmits it to the distribution server 22.
[0031] For example, the client attribute information includes device information about the client 23, such as the video frame rate that the client 23 can support when playing 3D content, and 6DoF (Degrees of Freedom) information, which is position and direction information that indicates the position and orientation of the user (avatar) within the three-dimensional virtual space.
[0032] The distribution server 22 performs multiplexing and interpolation processing according to the characteristics of the client 23 , that is, the client attribute information, based on the generated video data and the client attribute information, and generates distribution video data for each client 23 .
[0033] Multiplexing here refers to a process in which, for example, a video (hereinafter also referred to as generated video) at a specific time (playback time) based on generated video data is treated as one frame, and multiple frames (generated videos) are arranged in a specific order to generate multiple frames of distribution video data.
[0034] The interpolation process is a process of generating, based on a plurality of video data, distribution video data for playing back video at a desired line of sight (field of view) and time (playback time) with a desired viewpoint position in a three-dimensional virtual space. The interpolation process may use not only the generated video data but also distribution video data from past times.
[0035] Through multiplexing and interpolation processing by the distribution server 22, distribution video data is obtained for playing video (hereinafter also referred to as distribution video) with the direction of the user's face as the line of sight as seen from the user's (avatar's) position in the three-dimensional virtual space.
[0036] The generation of the video to be distributed will be further described with reference to FIGS.
[0037] For example, as shown in Figure 2, when user U11 is wearing an HMD as client 23, 6DoF information of user U11 is generated by input from movements of user U11, such as movement of user U11 or changes in facial direction, or by operation input from user U11.
[0038] Then, client attribute information including the generated 6DoF information is transmitted from the client 23 to the distribution server 22. Note that although only one user U11 is illustrated here, in reality, multiple users (clients 23) are connected to the distribution server 22, and these multiple users participate in events, etc., that take place in a three-dimensional virtual space.
[0039] The distribution server 22 instructs each of the rendering servers 21 to generate generated video data in accordance with the client attribute information received from each client 23. In this case, the distribution server 22 instructs the generation of generated video data with a predetermined position in the three-dimensional virtual space as the viewpoint position at a predetermined time (playback time) of the 3D content.
[0040] In response to an instruction from the distribution server 22, each rendering server 21 generates generated video data at a different viewpoint position and time.
[0041] The distribution server 22 generates distribution video data for each client 23 based on the generated video data generated by each rendering server 21 and the client attribute information received from the client 23, and transmits (distributes) the data to the client 23. As a result, the client 23 displays a two-dimensional video (2D image) corresponding to the position and orientation of the user U11 in the three-dimensional virtual space as the distribution video P11.
[0042] In the rendering server 21, as shown in FIG. 3, for example, prepared scene data is used to generate generated video data through 3D rendering processing.
[0043] The scene data includes three-dimensional object data, which is information about the video and audio of one or more objects placed in a three-dimensional virtual space, and scene description information, which indicates the placement position and orientation of the objects at each time.
[0044] For example, the three-dimensional object data includes video object data and audio object data for each object.
[0045] Video object data is data for displaying a video (image) of an object, and is made up of, for example, model data and texture data that indicate the three-dimensional shape of the object, etc. Audio object data is audio data for playing the sound of the object.
[0046] The scene description information is information that indicates which object is located at which position in the three-dimensional virtual space and in which direction it faces at each time (playback time) of the 3D content. In other words, the scene description information is information that describes what kind of scene each scene in the 3D content is.
[0047] The rendering server 21 generates generated video data by 3D rendering processing using such scene data.
[0048] In the 3D rendering process, as shown in the center of the figure, each object is arranged in a three-dimensional virtual space according to the scene description information, and a two-dimensional (2D) image of the three-dimensional virtual space as seen from a desired viewpoint position SP11 is generated. The viewpoint position SP11 may be the position of a predetermined user, or may be a position different from the user's position, i.e., an arbitrary position where no user is present. More specifically, the 3D rendering process also generates audio data for reproducing sounds to be heard at the desired viewpoint position SP11.
[0049] The generated image generated by the 3D rendering process may be a panoramic image (omnidirectional image) that is an image (image) in all directions (up, down, left, and right) as seen from viewpoint position SP11, i.e., a 360-degree image, or may be a 2D cropped image obtained by cropping a portion of the panoramic image. For example, when a cropped image is generated as the generated image, a region of the field of view when facing a predetermined direction is cropped from the panoramic image and used as the generated image.
[0050] In the following description, it is assumed that the generated image generated by the 3D rendering process is a panoramic image. However, even in such a case, it is not always necessary to generate an omnidirectional image depending on the viewpoint position, such as when the viewpoint position is near the edge of the three-dimensional virtual space. Therefore, generation of a portion of the panoramic image may be omitted depending on the viewpoint position, etc. In other words, a panoramic image may be generated in which the image of the three-dimensional virtual space is not displayed in a portion of the image, such as an image that is essentially 280 degrees.
[0051] The distribution server 22 generates a distribution video, which is a 2D video, by interpolation processing, multiplexing, etc. based on the multiple generated videos generated by the rendering server 21. As a result, the client 23 displays a distribution video P21 with a field of view determined by the user's position and orientation in the three-dimensional virtual space.
[0052] In addition, the 3D rendering process may use not only scene data but also client attribute information and video data captured in real time.
[0053] For example, if the 3D content is live footage, it is conceivable to record a live performance of a performer in real space and use the resulting video and audio data for 3D rendering processing. In this way, photorealistic performers can be placed in a three-dimensional virtual space, and a live performance in a three-dimensional virtual space such as a metaverse space can be distributed as 3D content.
[0054] Furthermore, when client attribute information is used in the 3D rendering process, the client attribute information may include video object data and audio object data of the user (avatar).
[0055] <About the Present Technology> The information processing system 11 is a server-client system made up of a server and a client 23 connected via a network.
[0056] A plurality of clients 23 (users) are connected to the server, and the server generates distribution video to be presented to the users and provides the distribution video to the clients 23 .
[0057] Here, the server is composed of one or more servers (a server group) connected via a network. The server group realizes a rendering function that performs 3D rendering processing and a distribution image generation function that generates distribution images of different positions and times in a three-dimensional virtual space by interpolating the generated images obtained by the 3D rendering processing.
[0058] In this case, it is possible that one distribution server 22 realizes both the rendering function and the distribution video generation function, as shown in Fig. 4. In the example of Fig. 4, the distribution server 22 also functions as the rendering server 21 shown in Fig. 1.
[0059] However, basically, an information processing system 11 to which the present technology is applied is provided with a server group including a distribution server 22 and a plurality of rendering servers 21, as shown in Fig. 5. Here, the distribution server 22 has a distribution video generation function, and each rendering server 21 has a rendering function.
[0060] As described above, the distribution server 22 instructs the rendering server 21 to generate a generated image at a predetermined position at a predetermined time. At this time, the position and time (rendering cycle) of the generated image are optimized depending on the temporal and spatial position and time at which the generated image should be generated to facilitate interpolation processing, the area in the three-dimensional virtual space where users (avatars) are gathering (concentrating), and so on.
[0061] Furthermore, generated video for a position or time that is likely to be referenced when generating video for distribution, that is, that is frequently used in interpolation processing, may be generated at high speed by a rendering server 21 that is close to the distribution server 22 or that has a short transmission time. This allows the distribution server 22 to obtain the generated video more quickly and use it in interpolation processing.
[0062] Note that some or all of the processing performed by distribution server 22 may be performed by client 23. In other words, client 23 may realize a distribution video generation function.
[0063] A possible use case of the information processing system 11 is, for example, distribution of live content in the metaverse space.
[0064] In this use case, for example, a live music concert by an artist is held in a three-dimensional virtual space realized by a group of servers.
[0065] Artists perform in real space (the real world), and their performances are captured in real time by cameras and microphones. Based on the captured video and audio data, photorealistic images of the artists are placed in a three-dimensional virtual space (metaverse space), realizing a live music performance in the three-dimensional virtual space. Multiple users can participate in the live music performance in the three-dimensional virtual space via a network, and each user can share time and space with other users.
[0066] The user's viewpoint or an avatar that includes the user's own viewpoint is placed in the 3D virtual space, and the user can freely change the position of their viewpoint or avatar with, for example, 6DoF. In other words, the user (avatar) can freely move within the 3D virtual space that is the music concert venue, and can change the viewport by facing in any direction.
[0067] In such a case, the video shown in Fig. 6, for example, is displayed as the distributed video on the client 23. In this example, an artist A11 is performing a musical piece on a stage in a three-dimensional virtual space, and multiple avatars, including a specific user's avatar OD11, are positioned around the stage as audience members of the live music concert.
[0068] In this live music content, the artist's movements and voice, as well as the user's movements and operations, are reflected in the three-dimensional virtual space in real time, allowing for interaction between the artist and the user, or between users themselves.
[0069] Incidentally, in use cases such as the above-mentioned music concert, where multiple users participate as avatars in an event in a three-dimensional virtual space realized by a group of servers via a network, and can move to any position and look in any direction, multiple users can simultaneously connect to the three-dimensional virtual space from multiple viewpoints.
[0070] In a multi-user, multi-viewpoint simultaneous connection, it is necessary to generate video distribution in real time according to the user's viewpoint position, and reducing the load of the 3D rendering process in this case becomes an issue. In other words, assuming that one rendering server 21 performs 3D rendering process for multiple users, the issue from the viewpoint of economic rationality is how to increase the number of users that one rendering server 21 can handle.
[0071] If the processing load of the 3D rendering process can be reduced, the number of simultaneously connected users per rendering server 21 can be increased, thereby suppressing an increase in the number of rendering servers 21, which is costly.
[0072] One way to reduce the processing load of the 3D rendering process is to reduce the resolution of the 2D generated image that is generated as a result of the 3D rendering process. However, simply lowering the resolution of the generated image will have a significant impact on the image quality of the distributed image. In other words, the quality of the distributed image will decrease.
[0073] Therefore, in the present technology, the processing load of the 3D rendering process is reduced (lowered) by reducing the number of two-dimensional generated images generated by the 3D rendering process in the entire information processing system 11.
[0074] This technology reduces the processing load by lengthening the cycle in which 3D rendering processing is performed (hereinafter also referred to as the rendering cycle), that is, by reducing the number of 3D rendering processing operations per second, and by reducing the spatial positions at which 3D rendering processing is performed.
[0075] For example, lengthening the rendering cycle means thinning out the times on the time axis at which 3D rendering processing is performed, and such temporal thinning reduces the processing load of the 3D rendering processing.
[0076] Even when the rendering cycle is lengthened, motion blur in the generated video can be prevented by appropriately setting the exposure time (shutter speed) of the generated video separately from the rendering cycle, i.e., the frame rate of the generated video. The exposure time here refers to the length of the period of scene data used when generating one frame (one time slot) of generated video through 3D rendering processing. In other words, one frame of generated video is generated using data for the period set as the exposure time out of the data for the entire period of the 3D content that makes up the scene data.
[0077] Furthermore, reducing the positions at which generated images are generated by 3D rendering processing in a three-dimensional virtual space means thinning out the positions that are the target of simultaneous 3D rendering processing, i.e., spatial thinning. By spatial thinning out, the processing load can be reduced by reducing the positions that are subject to 3D rendering processing at the same time.
[0078] The above-described temporal and spatial thinning can reduce the number of 3D rendering processes required per second. However, if left as is, depending on the user's viewpoint, the distributed video may not be displayed appropriately or the frame rate of the distributed video may be reduced, making it difficult for the user to comfortably view 3D content. In other words, the quality of the user's viewing experience may be degraded.
[0079] Therefore, in this technology, the degradation of the quality of the user's viewing experience is suppressed by combining the above-mentioned temporal and spatial thinning with interpolation processing based on the generated video.
[0080] In the interpolation process, frames (generated images) generated by the 3D rendering process are used, and frames (distributed images) for the thinned-out times and positions are generated by an interpolation process that is different from the 3D rendering process.
[0081] In this case, it is assumed that the processing load when generating a frame (video to be distributed) through interpolation processing is lower than the processing load when generating the same frame through 3D rendering processing. In other words, interpolation processing is a process that generates video to be distributed with a lower processing load than 3D rendering processing. A lower processing load can be said to mean a smaller amount of processing required to generate video, so performing interpolation processing not only reduces the processing load but also shortens the processing time required to complete generation of the video to be distributed.
[0082] Another feature of the present technology is how to specifically perform temporal and spatial thinning, that is, how to determine the positions and times (playback times) at which thinning will be performed, and this feature will be described later.
[0083] To summarize the above, this technology uses roughly the following algorithm to reduce the processing load on the server while keeping costs down, while also maintaining the quality of the distributed video presented to the user.
[0084] That is, first, in the distribution server 22, the spatial position and time (playback time) of the generated video generated by the 3D rendering process, i.e., the frame (hereinafter also referred to as the base frame), are determined based on the user's (avatar's) position, orientation (viewing information), movement, movement speed, etc. At this time, an appropriate position and time are determined so that high-quality video can be generated for distribution in the subsequent interpolation process.
[0085] Next, a base frame (generated image) at the determined position and time is generated by a 3D rendering process.
[0086] Finally, distribution video (interpolated frame) is generated by interpolation processing based on the base frame. Note that not only the base frame but also other distribution video (interpolated frame) may be used for the distribution video, or other base frames may be generated by interpolation processing based on the base frame. Furthermore, while interpolation processing generally uses two or more frames (base frames or interpolated frames), distribution video may also be generated using only one frame.
[0087] <Configuration Example of Rendering Server> FIG. 7 is a diagram showing a configuration example of the rendering server 21. As shown in FIG.
[0088] The rendering server 21 includes a control unit 71 , a receiving unit 72 , a scene data storage unit 73 , a video generating unit 74 , and a transmitting unit 75 .
[0089] The control unit 71 controls the overall operation of the rendering server 21. The receiving unit 72 communicates with the distribution server 22 via the network.
[0090] For example, the receiving unit 72 receives scene data transmitted from the distribution server 22 and supplies it to the scene data storage unit 73, or receives rendering parameters transmitted from the distribution server 22 and supplies it to the image generating unit 74. More specifically, the distribution server 22 transmits to the rendering server 21 an execution request including the rendering parameters to generate a generated image, that is, to request the execution of a 3D rendering process.
[0091] Scene data is data for constructing a three-dimensional virtual space including objects such as avatars that are the subjects of the 3D content, and includes the above-mentioned three-dimensional object data and scene description information.
[0092] The rendering parameters are various parameters (information) used in the 3D rendering process. In other words, the rendering parameters are parameters for generating generated image data by the 3D rendering process. For example, the rendering parameters include time information, viewpoint information, image generation cycle information, and exposure time information.
[0093] The time information is information indicating the time (playback time) of the 3D content at which a generated video is to be generated, that is, information indicating the time at which a generated video is to be generated.
[0094] The viewpoint information is information indicating the viewpoint position of the 3D content from which the generated image is to be generated, i.e., information indicating the position (viewpoint position) in the three-dimensional virtual space that is the target for generating the generated image. Note that, when a clipped image obtained by clipping a portion of the panoramic image is generated as the generated image, the viewpoint information may also include information indicating the region from which the clipping will occur (viewing area), i.e., the line of sight direction.
[0095] The video generation cycle information indicates the rendering cycle, which is the cycle for generating the generated video (video generation cycle), i.e., the frame rate of the generated video. The exposure time information indicates the exposure time (shutter speed) of the generated video.
[0096] The scene data storage unit 73 stores (holds) the scene data supplied from the receiving unit 72 and supplies the held scene data to the video generating unit 74 as appropriate.
[0097] The video generation unit 74 generates generated video data (generated video) by performing 3D rendering processing based on the rendering parameters supplied from the receiving unit 72 and the scene data supplied from the scene data storage unit 73, and supplies the generated video data to the transmitting unit 75.
[0098] The transmitting unit 75 transmits the generated video data supplied from the video generating unit 74 to the distribution server 22 via the network.
[0099] <Configuration Example of Distribution Server> FIG. 8 is a diagram showing a configuration example of the distribution server 22. As shown in FIG.
[0100] The distribution server 22 includes a control unit 101 , a receiving unit 102 , a rendering parameter generating unit 103 , a transmitting unit 104 , a scene data storage unit 105 , a receiving unit 106 , a video storage unit 107 , an interpolation processing unit 108 , and a transmitting unit 109 .
[0101] The control unit 101 controls the overall operation of the distribution server 22. The receiving unit 102 receives client attribute information transmitted from each client 23 via the network and supplies it to the rendering parameter generating unit 103.
[0102] The rendering parameter generating unit 103 functions as a determining unit that determines the position and time in the three-dimensional virtual space at which a generated image is generated by 3D rendering processing.
[0103] That is, the rendering parameter generation unit 103 determines the position and time at which 3D rendering processing is to be performed based on the client attribute information of each client 23 (user) supplied from the receiving unit 102 and previously generated rendering parameters (time information and viewpoint information). In other words, the rendering parameter generation unit 103 thins out the positions and times that are the targets of the 3D rendering processing. The rendering parameter generation unit 103 also determines which rendering server 21 is to generate a generated image for which position and time, i.e., which generated image (base frame) is to be generated.
[0104] The rendering parameter generation unit 103 generates rendering parameters for each rendering server 21 based on the determination results of the position and time for performing the 3D rendering process and the client attribute information, and supplies the generated rendering parameters to the transmission unit 104. More specifically, an execution request including the rendering parameters is generated and supplied to the transmission unit 104.
[0105] In the following, the four-dimensional space formed by the three spatial axes of the three-dimensional virtual space and the time axis indicating the playback time of the 3D content will also be referred to as the rendering space, and points in the rendering space will also be referred to as the rendering points. Furthermore, among the rendering points in the rendering space, rendering points determined by the position and time at which the 3D rendering process is actually performed will also be referred to as the actual rendering points.
[0106] Therefore, it can be said that the rendering parameter generating unit 103 determines the actual rendering point at a future time based on the client attribute information and the actual rendering point at a past time.
[0107] The transmitting unit 104 transmits the rendering parameters (execution request) supplied from the rendering parameter generating unit 103 and the scene data supplied from the scene data storage unit 105 to the rendering server 21 via the network.
[0108] The scene data storage unit 105 stores (holds) scene data of 3D content in advance, and supplies the held scene data to the transmission unit 104 as appropriate.
[0109] The receiving unit 106 receives the generated video data transmitted from each rendering server 21 and supplies it to the video storage unit 107. The video storage unit 107 stores the generated video data supplied from the receiving unit 106 and supplies it to the interpolation processing unit 108 as appropriate.
[0110] The interpolation processing unit 108 references the rendering parameters generated by the rendering parameter generation unit 103 and the client attribute information of each client 23 as necessary, and performs multiplexing and interpolation processing based on the generated video data stored in the video storage unit 107, thereby generating distribution video data for each client 23 and supplying it to the transmission unit 109. The interpolation processing unit 108 generates frames (distribution video) of rendering points that are not actual rendering points in the rendering space.
[0111] The transmitting unit 109 transmits the distribution video data for each client 23 supplied from the interpolation processing unit 108 to each client 23 via the network.
[0112] <Configuration Example of Client> FIG. 9 is a diagram showing a configuration example of the client 23. As shown in FIG.
[0113] The client 23 is connected to an input unit 141 that is operated by the user and a display 142 that is a display unit that displays the distributed video.
[0114] For example, the input unit 141 may be composed of a mouse, a keyboard, a controller, a touch panel superimposed on the display 142, or the like, and supplies signals according to user operations to the client 23. Alternatively, the input unit 141 may be composed of various sensors that detect user movements, and may output the detection results of the user movements to the client 23 as user input.
[0115] The input unit 141 and the display 142 may be provided in the client 23 .
[0116] The client 23 includes a control unit 151 , a user input acquisition unit 152 , a client attribute information generation unit 153 , a transmission unit 154 , a reception unit 155 , and a video display control unit 156 .
[0117] The control unit 151 controls the overall operation of the client 23. The user input acquisition unit 152 supplies a signal corresponding to the user input supplied from the input unit 141 to the client attribute information generation unit 153. For example, the user input acquisition unit 152 acquires information (signals) relating to the movement of the user (avatar) and supplies the information (signals) to the client attribute information generation unit 153.
[0118] The client attribute information generation unit 153 identifies the overall movement of the user (avatar) and the movement of each part of the user based on the signal supplied from the user input acquisition unit 152, and generates client attribute information based on the identification results and information previously stored in the client 23. The client attribute information generation unit 153 supplies the generated client attribute information to the transmission unit 154.
[0119] For example, the client attribute information includes 6DoF information obtained from the results of identifying the user's movements, and device information previously stored in the client 23. In addition, the client attribute information may include information indicating the speed and direction of movement of the user (avatar), and the speed and direction of movement of each part of the user.
[0120] The transmitting unit 154 transmits the client attribute information supplied from the client attribute information generating unit 153 to the distribution server 22 via the network. The receiving unit 155 receives the distribution video data transmitted from the distribution server 22 and supplies it to the video display control unit 156.
[0121] The video display control unit 156 supplies the distribution video data supplied from the receiving unit 155 to the display 142, and causes the display 142 to display the distribution video.
[0122] <Description of Distribution Processing> The operation of the information processing system 11 will be described.
[0123] When each client 23 plays back 3D content, the distribution server 22 transmits scene data of the 3D content to each rendering server 21 in advance at an appropriate timing. That is, the transmission unit 104 of the distribution server 22 reads out scene data from the scene data storage unit 105 and transmits it to the rendering server 21. Furthermore, the reception unit 72 of the rendering server 21 receives the scene data transmitted from the distribution server 22 and supplies it to the scene data storage unit 73 for storage (holding).
[0124] When the scene data is supplied to each rendering server 21, the distribution server 22 then performs distribution processing at an arbitrary timing to distribute the 3D content.
[0125] The distribution process performed by the distribution server 22 will be described below with reference to the flowchart of FIG.
[0126] In step S11, the receiving unit 102 receives the client attribute information transmitted from each client 23 and supplies it to the rendering parameter generating unit 103. More specifically, the timing at which the client attribute information is transmitted differs for each client 23.
[0127] In step S12, the rendering parameter generating unit 103 determines an actual rendering point at a future time based on the client attribute information supplied from the receiving unit 102 and the result of determining the actual rendering point in the past.
[0128] For example, in step S12, the actual rendering point at a future time is determined based on at least one of the following: information about each user in the three-dimensional virtual space, the degree of density of multiple users in the three-dimensional virtual space, the position and time of previously generated images (results of determining the past actual rendering point), and, as appropriate, information about each rendering server 21 obtained from the rendering server 21, etc.
[0129] Here, the information about the user includes, for example, the user's position in the three-dimensional virtual space, the direction of the user's face (direction of line of sight), movement (speed and direction of movement), movement of each part of the user's body (speed of movement), etc. Furthermore, the information about the rendering server 21 includes, for example, the processing capacity of the rendering server 21, the processing load of the rendering server 21 at the current point in time (current time), the number of rendering servers 21 connected to the distribution server 22, etc.
[0130] As an example, actual rendering points are determined so that they are spaced as evenly (uniformly) as possible in the rendering space.
[0131] In other words, the positions and times at which the 3D rendering process is performed are determined so that the generated images are generated at as equal intervals (uniform) as possible in both time and space. That is, from a spatial perspective, the positions at which the generated images are generated are determined so that the multiple positions at which the generated images are generated are arranged at approximately equal intervals (approximately equal intervals) in the three-dimensional virtual space. Furthermore, the times at which the generated images are generated, i.e., the playback times of the generated images, are determined so that the generated images are generated at approximately equal intervals in the time direction as well.
[0132] If the actual rendering points are arranged at approximately equal intervals, i.e., if the generated images (base frames) that form the basis of the distributed images are generated at approximately equal intervals in both time and space, it will be possible to obtain uniform distributed images at any playback time and at any position in the three-dimensional virtual space, thereby maintaining high quality (image quality) of the distributed images.
[0133] Furthermore, by appropriately determining the actual rendering point based on client attribute information, etc., it is possible to reduce the 3D rendering resources required for generating 3D content (realizing an application). Furthermore, since a service for simultaneous connection of multiple users and multiple viewpoints can be provided with fewer rendering servers 21 (computer resources), it is possible to reduce the financial costs associated with operating and using the rendering servers 21.
[0134] In step S13 , the rendering parameter generating unit 103 generates a request for execution of 3D rendering processing for each rendering server 21 according to the result of determining the actual rendering point in step S12 , and supplies the request to the transmitting unit 104 .
[0135] For example, the rendering parameter generation unit 103 determines which rendering server 21 should be in charge of (execute) the 3D rendering process at which actual rendering point, based on at least one of the rendering parameters generated in the past, the allocation result (distribution result) of past base frames (generated images) to the rendering servers 21, the real space distance from the distribution server 22 to the rendering servers 21, the processing capacity and current processing load of each rendering server 21, and the number of rendering servers 21 connected to the distribution server 22. In other words, the allocation of base frames (generated images) to the rendering servers 21 is determined.
[0136] The rendering parameter generating unit 103 generates rendering parameters for the rendering server 21 in charge of each actual rendering point.
[0137] For example, based on the result of determining the actual rendering point this time, the rendering parameter generation unit 103 generates time information indicating the time (playback time) indicated by the actual rendering point and viewpoint information indicating the position indicated by the actual rendering point.
[0138] Furthermore, for example, the rendering parameter generation unit 103 generates video generation period information indicating the frame rate of the generated video based on the result of determining the actual rendering point this time, more specifically, the placement interval in the time axis direction (time direction) of the actual rendering points with different playback times that are handled by one rendering server 21 in the rendering space.
[0139] Furthermore, for example, the rendering parameter generation unit 103 determines the exposure time of the generated image based on device information about the client 23 (compatible frame rate and image refresh rate) contained in the client attribute information of each client 23, and generates exposure time information indicating the result of the determination.
[0140] The rendering parameter generation unit 103 generates rendering parameters including the time information, viewpoint information, video generation cycle information, and exposure time information obtained in this manner, and supplies an execution request including the rendering parameters to the transmission unit 104. The rendering parameter generation unit 103 also stores client attribute information of each client 23 in the execution request as necessary. This client attribute information is used as appropriate, for example, when placing a user's avatar in a three-dimensional virtual space.
[0141] In step S14, the transmitting unit 104 transmits the execution request supplied from the rendering parameter generating unit 103 to each rendering server 21. That is, the transmitting unit 104 transmits, to each of one or more rendering servers 21, a request to execute the 3D rendering process for generating generated images (base frames) at mutually different positions or times.
[0142] When an execution request is sent, each rendering server 21 performs 3D rendering processing, and the resulting generated video data (base frame) is sent to the distribution server 22 .
[0143] In step S15, the receiving unit 106 receives the generated video data transmitted from each rendering server 21 and supplies it to the video storage unit 107 for storage.
[0144] In step S16 , the interpolation processing unit 108 generates distribution video data for each client 23 by performing interpolation processing based on the generated video data, and supplies the generated distribution video data to the transmission unit 109 .
[0145] For example, the interpolation processing unit 108 generates distribution video data by appropriately referring to the client attribute information of the destination client 23 and the rendering parameters of each rendering server 21, and performing multiplexing and interpolation processing based on multiple generated video data.
[0146] More specifically, as an example, the interpolation processing unit 108 generates, through interpolation processing, a panoramic image in which the viewpoint is the position of the user (avatar) in the three-dimensional virtual space indicated by the 6DoF information included in the client attribute information.
[0147] The interpolation process may be any process that requires less processing (processing load) than 3D rendering, such as AI interpolation using a trained AI (Artificial Intelligence) model, interpolation using motion prediction based on motion vectors, or other frame upconversion processes (frame interpolation).
[0148] In addition, the interpolation processing unit 108 cuts out the area of the user's field of view, which is determined by the viewpoint position and line of sight indicated by the 6DoF information, from the obtained panoramic image, and uses the resulting two-dimensional cut-out image as one frame of distribution image.
[0149] Furthermore, the interpolation processing unit 108 appropriately arranges the distributed video for each playback time in an appropriate order, i.e., performs multiplexing, etc., to generate distributed video data for displaying a specified number of frames of distributed video at a frame rate indicated by the device information included in the client attribute information.
[0150] Because interpolation processing imposes a lower processing load than 3D rendering processing, it is possible to generate images in a shorter processing time and reduce the processing load on the rendering server 21. In other words, it is possible to achieve more efficient generation of images to be distributed. Therefore, by combining the determination of actual rendering points, i.e., thinning out rendering points, with interpolation processing, it is possible to shorten the response time to user input. In other words, it is possible to reflect the user's movements and the like in the distributed images with little delay, thereby improving the quality of the user's 3D content viewing experience.
[0151] In step S17, the transmitting unit 109 transmits the distribution video data for each client 23 supplied from the interpolation processing unit 108 to each client 23. As a result, the 3D content is distributed to each client 23.
[0152] If the 3D content also includes audio, the interpolation processing unit 108 or the like generates audio data constituting the 3D content by performing interpolation processing based on the audio data received from the rendering server 21. The transmitting unit 109 then transmits the audio data to the client 23 together with the distribution video data.
[0153] In step S18, the control unit 101 determines whether or not to end the process of distributing the 3D content. For example, in step S18, it is determined that the process is to end when distribution of the 3D content has been completed up to the last frame.
[0154] If it is determined in step S18 that the process is not yet finished, the process then returns to step S11, and the above-described process is repeated.
[0155] On the other hand, if it is determined in step S18 that the process is to be ended, the control unit 101 stops the operation of each unit of the distribution server 22, and the distribution process ends.
[0156] In this way, the distribution server 22 determines the actual rendering point based on the client attribute information, etc., and requests the rendering server 21 to perform 3D rendering processing according to the determination result. In addition, the distribution server 22 generates distribution video data for each client 23 by interpolation processing based on the generated video data, etc.
[0157] In this way, it is possible to reduce the processing load on the rendering server 21 and the distribution server 22 while keeping costs down, and also to maintain the quality of the distributed video presented to the user.
[0158] <Description of Data Generation Process> When the distribution process described with reference to Fig. 10 is started in the distribution server 22, the rendering server 21 performs the data generation process shown in Fig. 11. Hereinafter, the data generation process by the rendering server 21 will be described with reference to the flowchart in Fig. 11.
[0159] In step S 41 , the receiving unit 72 receives the execution request transmitted from the distribution server 22 and supplies it to the video generating unit 74 .
[0160] In step S42 , the video generating unit 74 performs 3D rendering processing based on the rendering parameters included in the execution request supplied from the receiving unit 72 and the scene data stored in the scene data storage unit 73 .
[0161] For example, the video generation unit 74 identifies a target period to be subjected to the 3D rendering process based on the time information and exposure time information included in the rendering parameters. For example, the target period is a period whose start time is the time indicated by the time information (playback time) and whose length is indicated by the exposure time information.
[0162] The video generation unit 74 places images of each object in a three-dimensional virtual space based on information about the target period from the information about the entire period of the scene description information included in the scene data and the video object data in the three-dimensional object data included in the scene data.
[0163] For example, the image generation unit 74 places an image of a performer (artist) at a predetermined position in the three-dimensional virtual space based on image data of the performer (artist) at a live performance or other event captured in real space. Furthermore, the image generation unit 74 places an image of an avatar corresponding to a user at a position in the three-dimensional virtual space indicated by the 6DoF information based on the client attribute information included in the execution request, more specifically, the 6DoF information included in the client attribute information. The image object data, i.e., the image data (texture data) and model data, used to obtain an image of each user's avatar may be stored in the rendering server 21 by some means or may be included in the client attribute information.
[0164] The video generation unit 74 generates a generated video as a base frame, based on the results of arranging objects, performers, and avatars in the three-dimensional virtual space and the viewpoint information included in the rendering parameters, with the viewpoint position being the position indicated by the viewpoint information in the three-dimensional virtual space at the time indicated by the time information.
[0165] The video generator 74 generates a panoramic video at one or more times indicated by the time information as a generated video (base frame) at a video generation period (rendering period) indicated by the video generation period information included in the rendering parameters, thereby obtaining generated video data at the frame rate indicated by the video generation period.
[0166] The video generation unit 74 executes the above process as a 3D rendering process, and supplies the resulting generated video data to the transmission unit 75 .
[0167] In step S43 , the transmitting unit 75 transmits the generated video data supplied from the video generating unit 74 to the distribution server 22 .
[0168] In addition, if the 3D content also has audio, the control unit 71 or the like also generates audio data for the target period that constitutes the 3D content, and the transmission unit 75 transmits the audio data to the distribution server 22 together with the generated video data.
[0169] In step S44, the control unit 71 determines whether or not to end the process of generating generated video data. For example, in step S44, it is determined that the process is to be ended if an instruction to end the process is received from the distribution server 22.
[0170] If it is determined in step S44 that the process is not yet finished, the process then returns to step S41, and the above-described process is repeated.
[0171] On the other hand, if it is determined in step S44 that the process is to be ended, the control unit 71 stops the operation of each unit of the rendering server 21, and the data generation process ends.
[0172] In this way, the rendering server 21 performs 3D rendering processing based on the rendering parameters in response to a request from the distribution server 22, and transmits the resulting generated video data to the distribution server 22. By performing 3D rendering processing in accordance with the rendering parameters in this way, generated video data can be generated with a small processing load.
[0173] <Explanation of Content Playback Processing> When the distribution processing described with reference to Fig. 10 is started in the distribution server 22, the client 23 performs the content playback processing shown in Fig. 12. The content playback processing by the client 23 will be described below with reference to the flowchart in Fig. 12.
[0174] In step S71, the client attribute information generating unit 153 determines whether or not there is a user input.
[0175] For example, the user inputs the user's movement by operating the input unit 141 or by moving while wearing a sensor serving as the input unit 141, and a signal indicating the input is acquired as user input by the user input acquisition unit 152. Upon acquiring the user input, the user input acquisition unit 152 supplies a signal corresponding to the user input to the client attribute information generation unit 153.
[0176] When a signal corresponding to a user input is supplied from the user input acquisition unit 152, the client attribute information generation unit 153 determines in step S71 that a user input has been made.
[0177] If it is determined in step S71 that there is user input, that is, if the user moves and the position of the user (avatar) in the three-dimensional virtual space changes, the client attribute information generation unit 153 generates client attribute information in step S72.
[0178] That is, the client attribute information generation unit 153 generates 6DoF information by identifying the movement of the user (avatar) based on the signal supplied from the user input acquisition unit 152, and generates client attribute information including the 6DoF information and device information. The client attribute information generation unit 153 supplies the obtained client attribute information to the transmission unit 154.
[0179] In step S73, the transmitting unit 154 transmits the client attribute information supplied from the client attribute information generating unit 153 to the distribution server 22, and then the process proceeds to step S74.
[0180] If it is determined in step S71 that there is no user input, the processes of steps S72 and S73 are not performed, and the process then proceeds to step S74.
[0181] If the processing of step S73 has been performed or if it is determined in step S71 that there has been no user input, in step S74 the receiving unit 155 receives the distribution video data transmitted from the distribution server 22 and supplies it to the video display control unit 156.
[0182] In step S75, the video display control unit 156 supplies the distributed video data supplied from the receiving unit 155 to the display 142, causing the display 142 to display the distributed video. This results in the 3D content being played back. If the 3D content also has audio, the receiving unit 155 also receives audio data of the 3D content. The control unit 151 or the like then supplies the audio data to a speaker (not shown), and the audio of the 3D content is also played back by the speaker.
[0183] In step S76, the control unit 151 determines whether or not to end the process of playing back the 3D content. For example, in step S76, it is determined that the process should end if the playback of the 3D content has been completed up to the last frame.
[0184] If it is determined in step S76 that the process is not yet to be ended, the process then returns to step S71, and the above-described process is repeated.
[0185] On the other hand, if it is determined in step S76 that the process is to be ended, the control unit 151 stops the operation of each unit of the client 23, and the content playback process ends.
[0186] In this way, the client 23 receives the distribution video data from the distribution server 22 and displays (plays) the distribution video, thereby allowing the user to view 3D content.
[0187] <Regarding Determination of Actual Rendering Point> Here, a specific example of determination of the actual rendering point will be described.
[0188] The constraints (prerequisites) for determining the position and time for performing 3D rendering processing in a three-dimensional virtual space, that is, the actual rendering point, will be described below with examples.
[0189] For example, as shown in FIG. 13, consider a case where five avatars, avatar A, avatar B, avatar C, avatar D, and avatar E, are located at different positions in a three-dimensional virtual space and can freely change their viewpoints, that is, their positions and orientations.
[0190] In FIG. 13, the letters "A" to "E" represent avatars A to E, respectively, and the arrows drawn on each avatar indicate the direction of the avatar.
[0191] In this example, the rendering server 21 needs to perform 3D rendering simultaneously (in parallel) so as to obtain a total of five images with the positions of the avatars as the viewpoints.
[0192] Now, suppose there is a constraint that one rendering server 21 can generate one generated image at a time. Also, suppose that one rendering server 21 generates an image (distributed image) with the position of one avatar as the viewpoint position, without performing interpolation processing in the distribution server 22.
[0193] In such a case, for example, as shown in Figure 14, when there are two rendering servers 21 installed in the information processing system 11, the maximum number of distribution videos that can be generated simultaneously is two, and it is not possible to generate distribution videos for five people.
[0194] As a method for obtaining images from the viewpoint positions of the five avatars, it is conceivable to operate two rendering servers 21 in a time-sharing manner to generate distribution images for the five avatars.
[0195] However, in this method, the rendering server 21 operates in a time-sharing manner, so while it is generating distribution video from the viewpoint of a certain avatar, i.e., distribution video for a certain user, it cannot generate distribution video for other users (avatars).
[0196] In other words, the generation cycle of the distributed video for each user becomes longer. In other words, frames of the distributed video are thinned out. This impairs the smoothness of the movement of subjects such as objects in the distributed video, and the user cannot enjoy a comfortable viewing experience.
[0197] Therefore, in this technology, interpolation processing is performed as a subsequent process after the 3D rendering process, making it possible to generate distribution video for each user without impairing the user's viewing experience, i.e., without degrading the quality of the distribution video.
[0198] At this time, a feature of this technology is that the position and time (period) of the image (generated image) generated by the 3D rendering process, i.e., the position of the actual rendering point in the rendering space, is determined taking into account the interpolation process performed in a later stage.
[0199] In other words, the method for determining the position and time (period) of the generated image generated by the 3D rendering process is a method for determining which position and time the image will be generated by the 3D rendering process, and which position and time the image will be generated by the interpolation process.
[0200] First, in this technology, a rendering space is assumed as described above. That is, as shown in Fig. 15 , for example, a four-dimensional space with a time axis (t) and three spatial axes (x, y, z) as axes is defined as the rendering space.
[0201] 15, for ease of viewing, the three spatial axes are combined into one, and the rendering space, which is originally a four-dimensional space, is depicted as a two-dimensional space consisting of one time axis and one spatial axis. In the other figures described below, the rendering space will also be depicted as a two-dimensional space as appropriate, similar to the case of FIG.
[0202] In the example of FIG. 15, the horizontal axis represents the time axis, and the vertical axis represents the space axis.
[0203] When 3D content is free viewpoint content with 6DoF (six degrees of freedom), the number of axes in space (spatial axes) is essentially six.
[0204] However, if the image generated by the 3D rendering process is a 360-degree panoramic image that includes all viewing directions (field of view), there is no need to consider the three axes that represent the viewing direction, namely, the yaw, pitch, and roll axes. Therefore, it is sufficient to consider only the x, y, and z directions as spatial axes that describe positions in the three-dimensional virtual space, and the rendering space can be a four-dimensional space consisting of three spatial axes and one time axis.
[0205] In the rendering parameter generation unit 103 of the distribution server 22, basically, as shown in Figure 15, multiple actual rendering points are arranged (placed at equal distances) from each other as equally spaced (equally spaced) as possible within the rendering space.
[0206] In this example, each circle in the figure represents one actual rendering point, and the actual rendering points are arranged at equal intervals in the rendering space. For example, at a given time t1, four actual rendering points are arranged, including actual rendering point RP11, which is determined by time t1 and position P31 in the three-dimensional virtual space.
[0207] For example, if multiple users move at different speeds in the real space, the difference in their speeds will be reflected in the movements of their avatars in the three-dimensional virtual space. In this case, the actual rendering points are determined so that they are arranged as evenly (uniformly) as possible, taking into account the speeds of the avatars. However, the number of rendering points that can be subject to 3D rendering processing at the same time, i.e., the number of actual rendering points at the same time, is limited by the number of rendering servers 21.
[0208] In reality, 3D rendering processing cannot be performed unlimitedly at any position and time. The number of 3D rendering processing that can be performed per unit time and per unit space is limited by constraints such as the number of rendering servers 21, and it is necessary to determine the position and time at which the 3D rendering processing is performed, i.e., the actual rendering point.
[0209] The generated video obtained by the 3D rendering process is supplied (transmitted) to the distribution server 22. The distribution server 22 then performs interpolation processing such as frame up-conversion, which generates a video (distributed video) at a time different from a frame (generated video or distributed video) at a predetermined time. This generates a new frame (distributed video) on the time axis, allowing for a distributed video with a higher frame rate.
[0210] Furthermore, the distribution server 22 performs interpolation processing such as AI interpolation based on an image (generated image or distributed image) viewed from a predetermined position in the three-dimensional virtual space and an image (generated image or distributed image) viewed from a position slightly away from the predetermined position. In this case, the interpolation processing uses a position between two points as the viewpoint position and generates a distributed image viewed from that viewpoint position.
[0211] As described above, this technology is characterized by optimally thinning out actual rendering points, in other words, determining actual rendering points for efficient interpolation processing, and performing interpolation processing at rendering points that are not actual rendering points on the time axis or space axis, in order to maintain the quality of the distributed video while reducing the rendering processing load.
[0212] A more specific example of determining the actual rendering point will now be described.
[0213] This technology anticipates both cases where the user, or more specifically, parts of the user's avatar's body (such as hands), are not included in the distributed video; that is, where only the user's viewpoint moves in the three-dimensional virtual space, and where part of the avatar is reflected in the distributed video.
[0214] 16 and 17, a specific example of determining the actual rendering point when part of an avatar is reflected in the distributed video will be described. Note that if part of the avatar is not included in the distributed video, part of the avatar will not be reflected in the distributed video regardless of the user's movement, so the speed of the user's hand or other movement does not affect the generation cycle (frame rate) of the distributed video.
[0215] 16 and 17 show a rendering space drawn as a two-dimensional space for convenience, with the horizontal axis representing the time axis and the vertical axis representing the space axis. In addition, in Fig. 16 and Fig. 17, the scale position of the time axis indicates, for example, the playback time of one frame of the distributed video.
[0216] The examples shown in Figures 16 and 17 focus on one user (avatar) in a three-dimensional virtual space. In particular, Figures 16 and 17 show an example in which the avatar itself does not move within the three-dimensional virtual space, but moves its arms and legs in place. One such example would be an avatar dancing or the like while remaining in a certain place without changing its position (viewpoint position).
[0217] When an avatar appears in the distributed video, the rendering cycle is affected by the user's (avatar's) body movements, i.e., the speed of movements of specific body parts such as the hands. Generally, the smaller the difference between the video used in the interpolation process for generated video and the distributed video obtained by the interpolation process, the higher the accuracy of the interpolation process and the higher the quality of the distributed video.
[0218] Considering these characteristics, it is conceivable that the rendering parameter generation unit 103 determines the actual rendering point based on the speed of movement of the user's (avatar's) body parts specified by the client attribute information. In other words, it is conceivable to generate image generation period information indicating the image generation period for generating a generated image based on the speed of movement of the user's body parts.
[0219] In particular, in this case, it is conceivable to shorten the image generation cycle of the generated image by the 3D rendering process if the avatar's body movements are fast, and lengthen the image generation cycle if the avatar's body movements are slow. This makes it easier to generate high-quality (high-image-quality) video for distribution while reducing the processing load of the 3D rendering process.
[0220] FIG. 16 shows an example in which the movement (speed of movement) of a body part of an avatar is fast.
[0221] In this case, as shown in the upper part of Fig. 16, a generated image is generated every two divisions along the time axis. Here, the positions of black circles, such as actual rendering point RP21 and actual rendering point RP22, represent actual rendering points, that is, rendering points where generated images are generated. Also, here, the position of the avatar in the three-dimensional virtual space is set as the position that is the target of the 3D rendering process, and the generated image at the actual rendering point is used as the distributed image as is.
[0222] As shown in the lower part of FIG. 16, at scale positions (rendering points) between actual rendering points in the time axis direction, i.e., at the positions indicated by white circles, distribution video is generated by interpolation processing based on the generated video.
[0223] In this example, for example, an interpolation process is performed based on the generated video of actual rendering point RP21 and the generated video of actual rendering point RP22 to generate a distribution video of rendering point RP23, which is located at a time in the future relative to those actual rendering points.
[0224] In this way, when the movement of an avatar's body parts is fast, the image generation cycle (frame rate) of the generated image, i.e., the cycle of the 3D rendering process, is made relatively short. In other words, the interval between actual rendering points is made narrow. When the avatar moves quickly, performing the 3D rendering process at shorter time intervals makes it easier to perform interpolation processing for rendering points between actual rendering points.
[0225] On the other hand, if the movement of the avatar's body parts is slow, a generated image is generated every three divisions along the time axis, as shown in the upper part of Fig. 17. Here, the positions of black circles, such as actual rendering point RP31 and actual rendering point RP32, represent actual rendering points.
[0226] Also, as shown in the lower part of Figure 17, at each of the two scale positions (rendering points) between the actual rendering points in the time axis direction, i.e., the positions of the white circles, distribution images are generated by interpolation processing based on the generated images.
[0227] In this example, for example, distribution video of rendering point RP33 and rendering point RP34 located between actual rendering point RP31 and actual rendering point RP32 is generated by interpolation processing.
[0228] Note that the movement of an avatar is not constant, and the movement of each part may speed up or slow down, so the interval between actual rendering points along the time axis (image generation period) can be determined dynamically according to the speed of the movement.
[0229] FIG. 18 shows an example in which a plurality of users (avatars) exist in a three-dimensional virtual space.
[0230] 18 shows a rendering space drawn as a two-dimensional space for convenience, with the horizontal axis representing the time axis and the vertical axis representing the space axis. In addition, in FIG. 18, the scale position of the time axis indicates, for example, the playback time of one frame of the distributed video.
[0231] In the example shown in Figure 18, four avatars, avatar 1 to avatar 4, exist apart from each other in a three-dimensional virtual space, and each avatar moves its limbs and other parts of its body to dance, etc., without changing its position (viewpoint position).
[0232] In FIG. 18, for example, the positions on the space axis marked with the letters "Avatar 1" to "Avatar 4" indicate the positions of avatars 1 to 4 in the three-dimensional virtual space.
[0233] For example, if video is generated with the position of each avatar in the three-dimensional virtual space as the actual rendering point at all times, since the four avatars are located far apart, the position of each avatar at each playback time is set as the actual rendering point, as shown on the left side of the figure. Note that in Figure 18, the positions of the black circles represent the actual rendering points, that is, the rendering points where the generated video is generated.
[0234] In contrast, when combining interpolation and thinning of actual rendering points according to the speed of movement of each avatar's body parts, the interval between actual rendering points will differ depending on the speed of movement of the avatar's hands and other parts, as shown on the right side of the figure. In other words, the image generation cycle of the generated image changes depending on the speed of movement of the avatar's body parts.
[0235] In the drawing, the positions of the white circles on the right side indicate the positions of rendering points where distribution video is generated by interpolation processing based on the generated video.
[0236] Therefore, for example, in the example on the right side of the figure, focusing on avatar 1, generated images are generated by 3D rendering processing at actual rendering points RP41 and RP42, and these generated images are used as the distributed images as they are. Also, at rendering point RP43, which is between actual rendering points RP41 and RP42, distributed images are generated by interpolation processing. In this way, for avatar 1, 3D rendering processing is performed every two divisions along the time axis.
[0237] In contrast, for avatar 4 in the example on the right side of the figure, images are generated by 3D rendering at actual rendering points RP51 and RP52, and these generated images are used as-is as the distributed images. Furthermore, images to be distributed are generated by interpolation at rendering points RP53 to RP55 between actual rendering points RP51 and RP52. In this way, 3D rendering is performed for avatar 4 every four ticks along the time axis.
[0238] In this example, it can be seen that for avatar 1, which moves quickly, a generated image is generated in a short image generation cycle, and for avatar 4, which moves slowly, a generated image is generated in a longer image generation cycle than for avatar 1.
[0239] Next, an example in which an avatar moves within a three-dimensional virtual space will be described.
[0240] In particular, this section explains how to determine the rendering period of the image seen by the avatar (user) when focusing on the movement speed of the avatar, that is, the image generation period of the generated image when the position of the avatar is the actual rendering point.
[0241] When the avatar moved its body without moving, as shown in Figures 16 to 18, its position on the spatial axis of the renaming space did not change.
[0242] On the other hand, in this example, since the avatar moves within the three-dimensional virtual space, the position of the avatar on the spatial axis of the rendering space changes.
[0243] However, even when an avatar moves, the shorter the distance in the three-dimensional virtual space between the position where the generated image is generated by the 3D rendering process and the position where the distribution image is desired to be generated by the interpolation process, the higher the accuracy of the interpolation process and the higher quality of the distribution image can be obtained.
[0244] Considering these characteristics, it is conceivable that the rendering parameter generation unit 103 determines the actual rendering point based on the moving speed of the user (avatar) specified by the client attribute information. In other words, it is conceivable to generate image generation period information indicating the image generation period for generating a generated image based on the moving speed of the user.
[0245] In particular, in this case, it is conceivable to shorten the image generation cycle of the images generated by the 3D rendering process if the avatar's movement speed in the three-dimensional virtual space is fast, and lengthen the image generation cycle of the images generated if the avatar's movement speed is slow. This makes it easier to generate high-quality (high-image-quality) images for distribution while reducing the processing load of the 3D rendering process.
[0246] A specific example of determining the actual rendering point depending on the moving speed of an avatar will be described with reference to FIG.
[0247] 19 shows a rendering space drawn as a two-dimensional space for convenience, with the horizontal axis representing the time axis and the vertical axis representing the space axis. This example focuses on one user (avatar) in a three-dimensional virtual space.
[0248] For example, consider the cases where an avatar moves in a three-dimensional virtual space at a fast speed and at a slow speed, i.e., the case where the avatar moves as shown by the line L11 and the case where the avatar moves as shown by the line L12.
[0249] Here, the position of the avatar is set as the viewpoint position, and when the avatar moves a predetermined distance in the three-dimensional virtual space, a generated image is generated by 3D rendering processing.
[0250] In such a case, as shown on the left side of the figure, the actual rendering points are arranged at equal distances (equal intervals) when viewed in the spatial axis direction in the rendering space. Note that in Figure 19, the positions of the black circles represent the actual rendering points, and the positions of the white circles represent the positions of the rendering points where the distribution video is generated by interpolation processing.
[0251] Specifically, in this example, a generated image is generated approximately every two divisions in the spatial axis direction of the rendering space. That is, the positions of the actual rendering points are determined so that each actual rendering point is located approximately every two divisions in the spatial axis direction of the rendering space and is positioned on the division in the time axis direction.
[0252] For example, if the avatar moves as shown on the line L11, five positions on the line L11 are set as actual rendering points, as shown on the left side of the figure. On the other hand, if the avatar moves as shown on the line L12, four positions on the line L12 are set as actual rendering points, as shown on the left side of the figure.
[0253] Comparing these examples, it can be seen that the actual rendering points are arranged at approximately equal intervals in the spatial axis direction, and therefore the faster the avatar's movement speed, the more actual rendering points are provided. In this example, at the actual rendering points, the generated image generated by the 3D rendering process is used as the distributed image as is.
[0254] Furthermore, as shown on the right side of the figure, interpolation processing is performed on some or all of the positions on the time axis scale between adjacent actual rendering points on the line L11 or the line L12.
[0255] For example, when the avatar moves as shown in line L11, the distribution video of the positions (rendering points) between the actual rendering points, such as rendering point RP61 and rendering point RP62, which are on line L11 and on the time axis scale, is generated by interpolation processing.
[0256] In particular, in this example, because the avatar moves at a high speed, it can be seen that generated images are generated by 3D rendering processing or distributed images are generated by interpolation processing at every scale position on the time axis. As a result, when viewed along the time axis, distributed images (frames) exist at every scale position, that is, distributed images exist at every scale position.
[0257] On the other hand, when the avatar moves as shown in line L12, the distribution video of a position (rendering point) between the actual rendering points, such as rendering point RP63, which is on line L12 and on the time axis scale, is generated by interpolation processing.
[0258] In this example, because the avatar's movement speed is slow, it can be seen that, when viewed along the time axis, a generated image is generated by 3D rendering processing or a distributed image is generated by interpolation processing at every three divisions. In other words, when viewed along the time axis, a distributed image exists every three divisions. However, because the avatar's movement speed is slow in this example, the quality of the user's video viewing experience does not deteriorate even if the frame rate of the distributed image is lowered. Note that the distributed image may always be generated at the same frame rate regardless of the avatar's movement speed.
[0259] As described above, when actual rendering points are arranged at approximately equal intervals on the spatial axis and distribution images are generated by interpolation at positions between the actual rendering points, the faster the avatar's movement speed, the more generated images per unit time and the more distributed images are generated by interpolation. In other words, the faster the avatar's movement speed, the shorter the image generation cycle of the distribution images and the execution cycle of the interpolation process become.
[0260] In Figure 19, an example is described in which there is one user (avatar) in a three-dimensional virtual space, but there may be multiple avatars in the three-dimensional virtual space, and these avatars may move at different speeds.
[0261] FIG. 20 shows an example in which a plurality of users (avatars) are present in a three-dimensional virtual space and these avatars move at different speeds.
[0262] 20 shows a rendering space drawn as a two-dimensional space for convenience, with the horizontal axis representing the time axis and the vertical axis representing the space axis. Note that in Fig. 20, the positions of the black circles represent actual rendering points, and the positions of the white circles represent rendering points where video to be distributed is generated by interpolation processing.
[0263] In this example, avatar 1, which moves at a fast speed as indicated by line L21, and avatar 2, which moves at a slow speed as indicated by line L22, exist in a three-dimensional virtual space, and avatar 1 and avatar 2 intersect at a certain point in the three-dimensional virtual space.
[0264] Also in this example, as in the example of Figure 19, for each avatar, actual rendering points are placed approximately every two graduations in the spatial axis direction in the rendering space to generate generated images, and distribution images are generated by interpolation processing at positions between these actual rendering points.
[0265] In such a case, for avatar 1 with a fast moving speed, five positions on the line L21 are set as actual rendering points, and between adjacent actual rendering points, a distributed image is generated by interpolation processing at one position on the line L21.
[0266] Specifically, for actual rendering points such as actual rendering point RP71 and actual rendering point RP72 on line L21, generated images are generated by 3D rendering processing, and the generated images are used as the distributed images. Also, distributed images of positions (rendering points) between actual rendering points on line L21 and on the time axis scale, such as rendering point RP73, are generated by interpolation processing.
[0267] For avatar 1, because the moving speed of avatar 1 is fast, generated images are generated by 3D rendering processing or distributed images are generated by interpolation processing at all scale positions on the time axis.
[0268] On the other hand, for example, for avatar 2, which moves slowly, four positions on line L22, such as actual rendering point RP81 and actual rendering point RP82, are set as actual rendering points, and generated images are generated at these actual rendering points.
[0269] Furthermore, for example, the distribution video of some positions (rendering points) between the actual rendering points, which are on the line L22 and on the time axis scale, is generated by interpolation processing. Here, the interpolation processing is performed for three positions (rendering points) on the line L22, including the rendering point RP83.
[0270] For avatar 2, the movement speed of avatar 2 is slow, so when viewed along the time axis, generated images are generated by 3D rendering processing or distributed images are generated by interpolation processing at positions every three scale marks.
[0271] As described above, in the example of FIG. 20 as well, similarly to the example of FIG. 19, the faster the moving speed of the avatar, the shorter the image generation cycle of the generated image and the execution cycle of the interpolation process become.
[0272] Generalizing the examples described with reference to Figures 16 to 20, the problem of how to generate distribution video for each user comes down to the problem of determining the position and time at which video is generated by 3D rendering processing and the position and time at which video is generated by interpolation processing within a four-dimensional space (rendering space) formed by a spatial axis and a time axis.
[0273] The constraints in this case are to balance the number of rendering points that can be simultaneously 3D rendered at a given time (number of simultaneous renderings) with maintaining the quality of the distributed video. While satisfying this condition, the points on which 3D rendering will be performed (actual rendering points) and the points on which interpolation will be performed (rendering points that are not actual rendering points) are determined.
[0274] Therefore, the actual rendering point may be determined as shown in Fig. 21. Fig. 21 shows a rendering space that is conveniently drawn as a two-dimensional space with the horizontal axis representing the time axis and the vertical axis representing the space axis. In Fig. 21, the positions of the black circles represent the actual rendering points, and the positions of the white circles represent the positions of the rendering points at which the video to be distributed is generated by interpolation processing.
[0275] For example, if there is no restriction on the number of rendering points that can be 3D rendered simultaneously, 3D rendering can be performed at all rendering points in the rendering space, i.e., at all positions and times, as shown on the left side of the figure, to generate the generated image (distributed image).
[0276] On the other hand, if there is a constraint that the number of rendering points that can be 3D rendered simultaneously is two, then the number of actual rendering points at each time point should be two or less, as shown on the right side of the figure. In this case, the intervals between actual rendering points in the time axis direction and the space axis direction should be determined according to the placement position and movement speed of the avatar, the speed of movement of the avatar's body parts, etc.
[0277] For example, at the time enclosed by frame T11 on the right side of the figure, 3D rendering is performed at two actual rendering points RP91 and RP92, and interpolation is performed at the remaining rendering points RP93 and RP94. In this example, the actual rendering points are spaced at appropriate intervals, so the quality of the distributed video does not deteriorate.
[0278] Next, a specific example of determining the actual rendering point taking into consideration the number of rendering servers 21 that can simultaneously perform 3D rendering processing will be described.
[0279] 22, assume that there are five rendering servers 21 in an information processing system 11, and that it is possible to simultaneously perform 3D rendering processing for five locations in a three-dimensional virtual space. Also assume that five users (avatars), User A to User E, exist in the three-dimensional virtual space.
[0280] In such a case, as shown in FIG. 23, one rendering server 21 can be assigned to each user (avatar).
[0281] 23, the horizontal axis represents time, i.e., the time slot corresponding to the time of one frame of 3D content (distributed video), and the vertical axis represents the position in the three-dimensional virtual space where the 3D rendering process is performed. Therefore, in Fig. 23, one rectangle corresponds to one rendering point in the rendering space, and in particular, hatched (diagonally lined) rectangles represent rendering points where the 3D rendering process is performed, i.e., actual rendering points.
[0282] In Fig. 23, the positions marked with the letters "A" to "E" on the spatial axis indicate the positions of the avatars of users A to E in the three-dimensional virtual space. Note that in other figures described below, the positions marked with letters indicating users (avatars) also indicate the positions of the users in the three-dimensional virtual space.
[0283] In this example, one rendering server 21 is assigned to one user (avatar), and the position of the avatar in the three-dimensional virtual space at each time can be set as the position where 3D rendering processing is performed.
[0284] Therefore, there is no need to perform interpolation processing, and the generated video can be used as the video to be distributed as is, so that a highly accurate and smooth video to be distributed can be presented to the user.
[0285] 24, for example, assume that an information processing system 11 has two rendering servers 21, and is capable of simultaneously performing 3D rendering processing on two locations in a three-dimensional virtual space. Five users (avatars), User A to User E, exist in the three-dimensional virtual space. Furthermore, assume that no interpolation processing is performed.
[0286] In such a case, as shown in Figure 25, 3D rendering processing can only be performed at two locations (two users) at the same time, and the generated image cannot be presented as distribution image to the remaining three locations (three users).
[0287] 25, the horizontal axis represents time (time slots), and the vertical axis represents positions in the three-dimensional virtual space that can be the targets of 3D rendering processing. Also, in Fig. 25, hatched (diagonally lined) rectangles represent rendering points where 3D rendering processing is performed, i.e., actual rendering points.
[0288] In this example, for example, at time TL11 (time slot), 3D rendering processing is performed at the positions of user A's avatar and user C's avatar, and a generated image (distributed image) at the viewpoint position of user A and a generated image (distributed image) at the viewpoint position of user C are generated.
[0289] At the next time TL12, 3D rendering processing is performed at the positions of user B's avatar and user D's avatar, generating a generated video for the viewpoint position of user B and a generated video for the viewpoint position of user D. In this case, because 3D rendering processing is not performed at time TL12, for example, at user A's position, the generated video (distribution) generated at time TL11 continues to be displayed to user A.
[0290] In this way, if interpolation is not performed and the number of positions at which 3D rendering can be performed simultaneously is smaller than the total number of users, the update cycle of the video presented to each user will be long, resulting in the presented video having a low frame rate and with unsmooth object movements, significantly reducing the quality (value) of the user's viewing experience.
[0291] Therefore, this technology performs interpolation processing to shorten the update time (update cycle) of the distributed video and suppress the degradation of the user's viewing experience quality. At this time, the actual rendering point is determined so that the distributed video can be obtained with as little degradation as possible through the interpolation processing.
[0292] When determining the actual rendering point, there are two possible methods: assigning a rendering server 21 to a user, i.e., setting the position of the user (avatar) as the target position for the 3D rendering process; or setting a specific position as the target position for the 3D rendering process regardless of the user's position.
[0293] Hereinafter, the method of assigning a rendering server 21 to a user will be referred to as the user assignment method, and the method of setting a specific position as the target position for 3D rendering processing regardless of the user's position will be referred to as the position assignment method.
[0294] The rendering parameter generation unit 103 of the distribution server 22 determines the actual rendering point using a user allocation method or a position allocation method depending on the client attribute information, the results of past actual rendering point determinations, the number and processing capacity of the rendering servers 21, the processing load, etc.
[0295] For example, as shown in FIG. 24, it is assumed that there are two rendering servers 21 and that it is possible to simultaneously perform 3D rendering processing for two locations in a three-dimensional virtual space.
[0296] Also, as shown in Fig. 26, five users (avatars), User A to User E, exist in the three-dimensional virtual space. In Fig. 26, the circles with the letters "A" to "E" indicate the positions of the avatars of User A to User E in the three-dimensional virtual space. Hereinafter, User A's avatar will also be referred to as avatar A.
[0297] When each avatar is located at a position shown in FIG. 26, the rendering parameter generating unit 103 determines the actual rendering point by the user allocation method, as shown in FIG. 27, for example.
[0298] 27, the horizontal axis represents time, i.e., time slots corresponding to the time of one frame of 3D content (distributed video), and the vertical axis represents positions in the three-dimensional virtual space that can be the target of 3D rendering processing. Therefore, in Fig. 27, one rectangle corresponds to one rendering point in the rendering space, and in particular, hatched rectangles represent actual rendering points where 3D rendering processing is performed.
[0299] In addition, in FIG. 27, arrows pointing from a rectangle corresponding to a rendering point to another rectangle indicate the relationship between the reference source and the reference destination during interpolation processing.
[0300] Specifically, the image (generated image or distributed image) at the rendering point represented by the rectangle at the start position of the arrow is used in the interpolation process to generate the distributed image at the rendering point represented by the rectangle at the end position of the arrow. This also applies to Figure 30, which will be described later.
[0301] In the user allocation method, the rendering parameter generation unit 103 determines the actual rendering point for each future time based on the position, orientation, movement direction, and movement speed of each avatar (user) in the three-dimensional virtual space, the determination result of the actual rendering point for a past time, etc. In particular, in the user allocation method, the rendering point determined by the target time and the position of the avatar is set as the actual rendering point.
[0302] 27, at time TL21 (time slot), the positions of avatar A and avatar E are set as the positions to be subjected to the 3D rendering process (positions at which generated images are generated). In other words, the rendering point determined by time TL21 and the position of avatar A and the rendering point determined by time TL21 and the position of avatar E are set as the actual rendering points.
[0303] Therefore, at time TL21, two rendering servers 21 are assigned to user A (avatar A) and user E (avatar E), respectively.
[0304] Furthermore, at the next time TL22, the positions of avatar C and avatar D are set as the positions to be subjected to the 3D rendering process, and the generated image at the position of avatar C and the generated image at the position of avatar D are directly used as the distributed image of user C and the distributed image of user D.
[0305] On the other hand, at time TL22, the distribution videos of the remaining users A, B, and E are generated by interpolation processing.
[0306] For example, the distributed video of user A at time TL22 is generated by interpolation processing based on the distributed video (generated video) of user A at time TL21 and the generated video (distributed video) of user C at time TL22. That is, the distributed video of user A that is close in time, i.e., the frame (base frame) of the immediately preceding distributed video, and the generated video (base frame) of avatar C (user C) that is close in distance to avatar A at the same time are used for the interpolation processing.
[0307] In this way, by using high-quality base frames that are close in time and distance for the interpolation process, it is possible to obtain high-quality video for distribution with little degradation.
[0308] Furthermore, for example, if the number of users (avatars) is large compared to the number of positions at which 3D rendering processing can be performed at the same time, it may be possible to adopt a position allocation method in which the positions to be targeted for 3D rendering processing in the three-dimensional virtual space are fixed positions.
[0309] In the position allocation method, for example, generated images (base frames) are generated for each of a plurality of fixed positions different from the position of the user (avatar), and distribution images for the positions of all users (avatars) in the three-dimensional virtual space are generated by interpolation processing.
[0310] In such a position allocation method, a common base frame (generated image) that serves as input for the interpolation process is first generated by a 3D rendering process.
[0311] At this time, the position for generating the base frame is determined based on the density of avatars and the locations through which many avatars move, so that the base frame can be used to generate (interpolate) the distribution video of as many users (avatars) as possible.
[0312] In the interpolation process, a distribution video is generated from a base frame (generated video) that is close in distance and time to the position of the avatar (user) in the three-dimensional virtual space, with the viewpoint being the position of the avatar. That is, for example, the interpolation process uses not only past distribution video from the same position (avatar position), but also generated video (base frames) from the same or past times in positions surrounding the position.
[0313] As described above, interpolation processing is premised on a lower processing load than 3D rendering processing, and for example, interpolation processing is performed using AI interpolation, motion prediction based on motion vectors, etc. In particular, to ensure real-time performance, it is preferable to perform interpolation processing with a short processing time.
[0314] In the position allocation method described above, the rendering parameter generation unit 103 determines the actual rendering point at a future time based on the position, orientation, movement direction, movement speed of each avatar (user) in the three-dimensional virtual space, and the determination results of the actual rendering point at a past time.
[0315] In this case, it is important to determine the actual rendering point, that is, the position and time (rendering timing) for which a common base frame is to be generated.
[0316] For example, the actual rendering point is determined by taking into consideration whether the candidate rendering point is a point where interpolation processing is easy in terms of time and space, and in which area of the 3D virtual space avatars (users) are concentrated, etc. In other words, the position in the 3D virtual space that is the target of the 3D rendering processing and the cycle of the 3D rendering processing (image generation cycle) are optimized.
[0317] Furthermore, for example, a base frame (generated image) at a position that is likely to be referenced during interpolation processing at another viewpoint position is generated at high speed by a rendering server 21 located at a position close in distance to the distribution server 22, so that it can be immediately used for interpolation processing at another viewpoint position. In this case, the rendering parameter generation unit 103 determines the rendering server 21 that is to be responsible for generating the generated image based on the frequency of use (reference frequency) of the generated image (base frame) for interpolation processing and the distance from the distribution server 22 to the rendering server 21.
[0318] A specific example of a position allocation method will be described with reference to FIGS.
[0319] 28, assume that an information processing system 11 has two rendering servers 21 and is capable of simultaneously performing 3D rendering processing on two locations in a three-dimensional virtual space. Also assume that six users, User A to User F, are connected to the distribution server 22 via clients 23.
[0320] Furthermore, suppose that the avatars of users A to F are located in a three-dimensional virtual space, as shown in Fig. 29. In Fig. 29, the positions marked with letters "A" to "F" indicate the positions of the avatars of users A to F.
[0321] In such a case, the rendering parameter generation unit 103 determines the positions of the circles on which the letters "V1" to "V5" are written using the position allocation method as rendering positions V1 to V5, which are positions to be subjected to the 3D rendering process. Note that, although each rendering position is assumed to be a fixed position that does not change regardless of time, the placement of each rendering position may also be changed depending on time.
[0322] The rendering parameter generation unit 103 determines the rendering point in the rendering space determined by the rendering position and time (playback time) as the actual rendering point so that 3D rendering processing is performed at two of the five rendering positions at each time.
[0323] As a result, 3D rendering and interpolation processing are performed, for example, as shown in Figure 30. In Figure 30, the horizontal axis represents time (time slots) and the vertical axis represents the spatial axis. Also, in Figure 30, rectangles with diagonal hatching represent rendering points where 3D rendering processing is performed, i.e., actual rendering points, and rectangles with horizontal hatching represent rendering points where distribution video is generated by interpolation processing.
[0324] In this example, at time TL31 (time slot), generated images (base frames) are generated by 3D rendering processing for rendering positions V4 and V5. At time TL32, generated images (base frames) are generated by 3D rendering processing for rendering positions V1 and V3.
[0325] In this way, at each time, 3D rendering processing is performed at two of the five rendering positions V1 to V5.
[0326] Furthermore, for example, the distributed video of the viewpoint position of user D (avatar D) at time TL32 is generated by interpolation based on the generated video (base frame) at rendering position V3 at the same time, TL32, and the generated video (base frame) at rendering position V4 at past time TL31. Note that the distributed video may be generated based on one video (generated video or distributed video), such as the distributed video of user F at the time following time TL32. In other words, the distributed video at each time may be generated based on one or more videos.
[0327] As shown in Figure 29, by setting the rendering position to an appropriate position depending on the position of the user (avatar), interpolation processing can be performed using base frames that are close in time and space, resulting in high-quality distributed video.
[0328] As an example, a position allocation method may be considered in which a rendering position is determined so that for all users (avatars) in the three-dimensional virtual space, there is always a rendering position at each time that is within a predetermined distance from the avatar.
[0329] Including the above-described position allocation method and user allocation method, a number of methods (decision algorithms) can be considered for determining which position in the three-dimensional virtual space should be set as the rendering position to be subjected to the 3D rendering process.
[0330] Specific examples of the method include an even placement method in which multiple rendering positions are evenly (uniformly) placed within the three-dimensional virtual space, and an uneven distribution method in which multiple rendering positions are unevenly distributed within the three-dimensional virtual space. Also possible patterns include whether the rendering positions are fixed or movable.
[0331] For example, as shown in FIG. 31, when multiple users (avatars) exist in a three-dimensional virtual space, five rendering positions V1 to V5 are determined by the uniform placement method or the uneven placement method.
[0332] In this case, for example, as shown on the left side of Fig. 32, in the uniform arrangement method, the three-dimensional virtual space is divided equally, and rendering positions V1 to V5 are uniformly arranged based on the division results. At this time, each rendering position is fixed.
[0333] Furthermore, in the uneven distribution method, the placement positions of rendering positions V1 to V5 are determined according to the environment in the three-dimensional virtual space, as shown on the right side of Figure 32, for example, but the uneven distribution method does not necessarily require the rendering positions to be placed evenly.
[0334] The environment in the three-dimensional virtual space here refers to, for example, the position and orientation (line of sight) of each avatar (user), the direction of movement, the speed of movement, and the density of avatars in each area (position) in the three-dimensional virtual space, and the arrangement of rendering positions changes depending on such an environment. For example, more rendering positions are provided in positions where the density of avatars is high.
[0335] In the uneven distribution method, two patterns are possible: the rendering position is a fixed position that does not change regardless of time (fixed position), or the rendering position is a movable position that moves (changes) over time.
[0336] In both the uniform distribution method and the uneven distribution method, the user basically views the distributed video generated by the interpolation process. That is, the generated video (base frame) obtained by the 3D rendering process with the rendering position as the viewpoint position is basically not presented directly to the user but is used for the interpolation process. However, in some cases, the user (avatar) may be located at the same position as or close to the rendering position, and in such cases the generated video (base frame) may be used as the distributed video as is.
[0337] The uniform distribution method and the uneven distribution method will be further explained.
[0338] For example, in the uniform placement method, a three-dimensional virtual space containing many users (avatars) is divided equally, as shown in the upper part of Fig. 33, and rendering positions are evenly placed according to the division results, as shown in the lower part of the figure. Here, the positions of the circles with the letters "A" to "E" written on them are set as rendering positions, and it can be seen that the rendering positions are evenly placed within the three-dimensional virtual space.
[0339] The uniform placement method is the simplest method for determining the rendering position, and does not require optimization of the rendering position according to the time of day.
[0340] For example, the uniform placement method is considered to be effective when there are a sufficiently large number of users in the three-dimensional virtual space, etc. However, with the uniform placement method, there is a possibility that many rendering positions will not be used in the interpolation process and will be wasted if the structure of the three-dimensional virtual space includes places where users do not come or cannot approach, if the number of users is small, or if there is a large imbalance in the distribution of user positions.
[0341] Furthermore, by gridding the 3D virtual space and limiting the positions that the user (avatar) can move to to grid points, the positions where the video to be distributed is generated through interpolation can also be fixed, so such gridding and the uniform placement method can be combined. In this case, it becomes easier to determine the rendering position, the interpolation process, and the video to be referenced during the interpolation process.
[0342] In the uneven distribution method, multiple rendering positions are unevenly distributed according to the environment in the 3D virtual space. In this case, depending on how the rendering positions are determined, it is possible to present the generated image (base frame) at the rendering position as it is to the user as the distributed image.
[0343] In the uneven distribution method, when the number of users in the three-dimensional virtual space is small, it is possible to set the rendering position as the movable position and the position of the user (avatar) as the rendering position, as shown in the upper part of Figure 34, for example.
[0344] In this example, each circle with the letters "A" to "C" written on it represents a rendering position, and the position of each of the three avatars (users) in the three-dimensional virtual space is set as the rendering position. In this case, the position of the avatar and the rendering position are aligned, so the rendering position moves in accordance with the movement of the avatar.
[0345] In this way, the generated image (base frame) at the rendering position can be directly supplied to the user as the distributed image and displayed. In this case, the highest quality image can be obtained as the distributed image.
[0346] Furthermore, for example, if the number of users in the three-dimensional virtual space increases from the state shown in the upper part of Figure 34, the purpose of the 3D rendering process switches from generating video for distribution to generating base frames for interpolation processing.
[0347] 34, each circle marked with a letter "A" to "D" represents a rendering position, and in this example, the number of rendering positions is increased from 3 to 4. The increase in rendering positions may be achieved, for example, by increasing the number of rendering servers 21 used, or by thinning out the actual rendering points in the time direction.
[0348] For example, the rendering position is determined by a predetermined algorithm, such as determining a position where interpolation processing is easy to perform as the rendering position.
[0349] At this time, whether the rendering position is a fixed position or a movable position is determined by the rendering parameter generation unit 103 according to the environment of the three-dimensional virtual space, such as the situation of the 3D content scene, the user's movement, and the degree of crowding.
[0350] For example, if the 3D content is an event held in a three-dimensional virtual space and the time and place where users will gather in the three-dimensional virtual space are known in advance, the rendering position may be located at the location where the users will gather. Furthermore, if the density of users is high, the rendering position may be fixed. Furthermore, a mixture of fixed and movable rendering positions may be used.
[0351] In the example described above, when the rendering parameter generation unit 103 determines the actual rendering point in the rendering space, that is, the time (playback time) and position (rendering position) at which the 3D rendering process is performed, the temporal distance and the spatial distance are treated as equivalent.
[0352] In this case, for example, based on the movement of a user (avatar) in a three-dimensional virtual space, spatial distance can be converted into a time difference, or conversely, time difference can be converted into spatial distance, and the actual rendering point can be determined based on the conversion result.
[0353] As an example, it is assumed that a distribution video can be generated by interpolation processing for the position of an avatar (user) in a three-dimensional virtual space, that is, within a range of 1 m from the user's viewpoint position.
[0354] In such a case, for example, a stream of a user moving at a speed of 1 m / s (30 fps) could be composed of generated video (base frames) at 1 fps and video generated by interpolation at 29 fps. Similarly, a stream of a user moving at a speed of 2 m / s (30 fps) could be composed of generated video (base frames) at 2 fps and video generated by interpolation at 28 fps.
[0355] Furthermore, for example, when temporal distance and spatial distance are treated equally, rendering positions are arranged evenly in the temporal and spatial directions.
[0356] As an example, if the spatial distance is unified, the rendering positions are arranged evenly (at equal distances) within the three-dimensional virtual space, and if the temporal distance is unified, the time at which the 3D rendering process is performed is determined so that the time period is equally spaced.
[0357] Furthermore, for example, if the density of users (avatars) in a three-dimensional virtual space changes over time, it is possible to move the rendering positions according to the density at each time so that the rendering positions are spaced at equal intervals in time and space.
[0358] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.
[0359] FIG. 35 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0360] In the computer, a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .
[0361] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0362] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0363] In a computer configured as described above, the CPU 501 loads a program recorded in the recording unit 508, for example, into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the above-described series of processes.
[0364] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0365] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0366] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0367] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0368] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0369] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0370] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0371] Furthermore, the present technology can also be configured as follows.
[0372] (1) An information processing device comprising: an interpolation processing unit that generates, by interpolation processing, a second video having a viewpoint position or a time different from that of a first video, based on a first video generated by a rendering processing and having a viewpoint position at a predetermined position in a space. (2) The information processing device according to (1), wherein the interpolation processing is a process with a lower processing load than the rendering processing. (3) The information processing device according to (1) or (2), wherein the interpolation processing unit generates the second video having a viewpoint position that is a user's position in the space, and further comprises a first transmission unit that transmits the second video to a client of the user. (4) The information processing device according to any one of (1) to (3), further comprising: a second transmission unit that transmits, to one or more rendering servers, requests to execute the rendering processing to generate the first video at mutually different positions or times; and a reception unit that receives the first video transmitted from the rendering servers. (5) The information processing device according to (4), further comprising a determination unit that determines a position and a time within the space at which the first video is to be generated based on at least one of information about users within the space, a density of the multiple users within the space, positions and times of the first videos generated in the past, and information about the rendering server. (6) The information processing device according to (5), wherein the information about the user is at least one of the user's position, orientation, movement speed, movement direction, and speed of movement of a body part within the space. (7) The information processing device according to (5) or (6), wherein the information about the rendering server is at least one of the number of the rendering servers, the processing capacity of the rendering server, and the current processing load of the rendering server. (8) The information processing device according to any one of (5) to (7), wherein the determination unit determines the positions within the space at which the first video is to be generated such that the multiple positions at which the first video is to be generated are approximately uniformly arranged within the space.(9) The information processing device according to any one of (5) to (8), wherein the determination unit determines the time at which the first video is to be generated so that the first video is generated at approximately equal intervals in the time direction. (10) The information processing device according to any one of (5) to (9), wherein the determination unit sets a predetermined position of the user in the space as the position at which the first video is to be generated. (11) The information processing device according to any one of (5) to (9), wherein the determination unit sets a position different from the position of the user in the space as the position at which the first video is to be generated. (12) The information processing device according to any one of (5) to (9), wherein the position at which the first video is to be generated in the space is a fixed position. (13) The information processing device according to any one of (5) to (11), wherein the position at which the first video is to be generated in the space is a movable position. (14) The information processing device according to any one of (5) to (7), wherein the determination unit determines a position within the space at which the first image is to be generated such that a plurality of positions at which the first image is to be generated are unevenly distributed within the space. (15) The information processing device according to any one of (5) to (14), wherein the determination unit determines a time at which the first image is to be generated such that a cycle at which the first image is generated is shortened when a speed of movement of a body part of the user or a moving speed of the user is fast. (16) The information processing device according to any one of (1) to (15), wherein the first image is a 360-degree view image. (17) The information processing device according to (4), wherein the execution request includes exposure time information indicating an exposure time of the first image. (18) The information processing device according to any one of (5) to (15), wherein the determination unit determines the rendering server to be in charge of generating the first video based on a frequency of use of the first video in the interpolation process and a distance from the information processing device to the rendering server. (19) An information processing method including: an information processing device generating, by interpolation processing, a second video having a viewpoint position different from that of the first video, based on a first video generated by rendering processing and having a viewpoint position at a predetermined position in space.(20) An information processing system having a rendering server and a distribution server, wherein the distribution server comprises: a first transmitting unit that transmits to the rendering server a request to execute a rendering process to generate a first video having a viewpoint position at a predetermined position in a space, a first receiving unit that receives the first video transmitted from the rendering server, and an interpolation processing unit that generates, by interpolation processing, a second video having a viewpoint position or a time different from that of the first video, and the rendering server comprises: a second receiving unit that receives the execution request, an video generating unit that performs the rendering process in response to the execution request and generates the first video, and a second transmitting unit that transmits the first video to the distribution server. (21) The information processing system according to (20), wherein the interpolation processing is a process having a lower processing load than the rendering processing. (22) The information processing system according to (20) or (21), wherein the interpolation processing unit generates the second video with a viewpoint position that is the user's position in the space, and the distribution server further includes a third transmission unit that transmits the second video to a client of the user. (23) The information processing system according to any one of (20) to (22), wherein the distribution server further includes a determination unit that determines a position and time in the space at which to generate the first video, based on at least one of information about the user in the space, a density of the users in the space, positions and times of the first videos generated in the past, and information about the rendering server. (24) The information processing system according to (23), wherein the information about the user is at least one of the user's position, orientation, movement speed, movement direction, and speed of movement of a body part in the space. (25) The information processing system according to (23) or (24), wherein the information about the rendering servers is at least one of the number of the rendering servers, the processing capacity of the rendering servers, and the current processing load of the rendering servers.(26) The information processing system according to any one of (23) to (25), wherein the determination unit determines a position in the space at which the first video is to be generated such that a plurality of positions at which the first video is to be generated are arranged approximately evenly within the space. (27) The information processing system according to any one of (23) to (26), wherein the determination unit determines a time at which the first video is to be generated such that the first video is generated at approximately equal intervals in the time direction. (28) The information processing system according to any one of (23) to (27), wherein the determination unit sets a predetermined position of the user in the space as the position at which the first video is to be generated. (29) The information processing system according to any one of (23) to (27), wherein the determination unit sets a position different from the position of the user in the space as the position at which the first video is to be generated. (30) The information processing system according to any one of (23) to (27), wherein the position within the space at which the first image is generated is a fixed position. (31) The information processing system according to any one of (23) to (29), wherein the position within the space at which the first image is generated is a movable position. (32) The information processing system according to any one of (23) to (25), wherein the determination unit determines the position within the space at which the first image is generated such that a plurality of positions at which the first image is generated are unevenly distributed within the space. (33) The information processing system according to any one of (23) to (32), wherein the determination unit determines the time at which the first image is generated such that the cycle at which the first image is generated is shortened when the speed of movement of a body part of the user or the moving speed of the user is fast. (34) The information processing system according to any one of (20) to (33), wherein the first image is a panoramic image. (35) The information processing system according to any one of (20) to (34), wherein the execution request includes exposure time information indicating an exposure time of the first video. (36) The information processing system according to any one of (23) to (33), wherein the determination unit determines the rendering server to be in charge of generating the first video based on a frequency of use of the first video in the interpolation process and a distance from the distribution server to the rendering server.
[0373] REFERENCE SIGNS LIST 11 information processing system, 21-1 to 21-N, 21 rendering server, 22 distribution server, 23-1 to 23-M, 23 client, 71 control unit, 72 receiving unit, 74 video generation unit, 75 transmitting unit, 101 control unit, 103 rendering parameter generation unit, 104 transmitting unit, 106 receiving unit, 108 interpolation processing unit
Claims
1. An information processing device that includes an interpolation processing unit that generates, based on a first image generated by a rendering process and having a predetermined position in space as a viewpoint position, a second image by interpolation processing, the second image having a different viewpoint position or time from the first image.
2. The information processing device according to claim 1, wherein the interpolation process is a process that has a lower processing load than the rendering process.
3. The information processing device according to claim 1, wherein the interpolation processing unit generates the second image with the user's position in the space as a viewpoint position, and further comprises a first transmission unit that transmits the second image to the user's client.
4. An information processing device as described in claim 1, further comprising: a second transmitting unit that transmits to one or more rendering servers, respectively, requests to execute the rendering process to generate the first image at a different location or time; and a receiving unit that receives the first image transmitted from the rendering server.
5. The information processing device of claim 4, further comprising a determination unit that determines the position and time within the space at which to generate the first image based on at least one of information about users within the space, the density of multiple users within the space, the position and time of the first image generated in the past, and information about the rendering server.
6. The information processing device according to claim 5, wherein the information about the user is at least one of the user's position, orientation, movement speed, movement direction, and speed of movement of a body part in the space.
7. The information processing device according to claim 5, wherein the information relating to the rendering servers is at least one of the number of the rendering servers, the processing capabilities of the rendering servers, and the current processing load of the rendering servers.
8. The information processing device according to claim 5, wherein the determination unit determines the positions within the space at which the first video is to be generated so that a plurality of positions at which the first video is to be generated are arranged approximately uniformly within the space.
9. The information processing device according to claim 5, wherein the determination unit determines the time at which the first video is to be generated so that the first video is generated at approximately equal intervals in the time direction.
10. The information processing device according to claim 5, wherein the determination unit determines a predetermined position of the user in the space as the position at which the first image is generated.
11. The information processing device according to claim 5, wherein the determination unit determines a position different from the position of the user in the space as the position at which the first image is to be generated.
12. The information processing device according to claim 5, wherein the position in the space where the first image is generated is a fixed position.
13. The information processing device according to claim 5, wherein the position in the space where the first image is generated is a movable position.
14. The information processing device according to claim 5, wherein the determination unit determines the positions within the space at which the first image is generated so that a plurality of positions at which the first image is generated are unevenly distributed within the space.
15. The information processing device according to claim 5, wherein the determination unit determines the time at which the first image is generated so that the cycle for generating the first image is shortened when the speed of movement of the user's body parts or the user's moving speed is fast.
16. The information processing device according to claim 1, wherein the first image is a panoramic image.
17. The information processing device according to claim 4, wherein the execution request includes exposure time information indicating an exposure time of the first image.
18. An information processing device according to claim 5, wherein the determination unit determines the rendering server to be responsible for generating the first image based on the frequency of use of the first image in the interpolation process and the distance from the information processing device to the rendering server.
19. An information processing method including: an information processing device generating, by interpolation processing, a second image having a different viewpoint position or time from that of the first image, based on a first image generated by rendering processing and having a viewpoint position at a predetermined position in space.
20. An information processing system having a rendering server and a distribution server, wherein the distribution server comprises: a first transmitting unit that transmits to the rendering server a request to execute a rendering process to generate a first image having a viewpoint position at a predetermined position in space; a first receiving unit that receives the first image transmitted from the rendering server; and an interpolation processing unit that generates, based on the first image, a second image having a viewpoint position or time different from that of the first image by interpolation processing; and the rendering server comprises: a second receiving unit that receives the execution request; an image generating unit that performs the rendering process in accordance with the execution request and generates the first image; and a second transmitting unit that transmits the first image to the distribution server.
Citation Information
Patent Citations
Multi-view video playing method and device and storage medium
CN114666565A
Video processing method and device, electronic equipment and storage medium
CN116801030A
Extraction program, image generation program, extraction method, image generation method, extraction device, and image generation device
JP2022162268A
Expanded field of view re-rendering for VR viewing
JP2022180368A
Predictive server-side rendering of scenes
US20200092599A1