Cloud rendering system and cloud rendering method

WO2026191588A1PCT designated stage Publication Date: 2026-09-17SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/007037
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-11
Filing Date
2026-02-26
Publication Date
2026-09-17

Smart Images

  • Figure JP2026007037_17092026_PF_FP_ABST
    Figure JP2026007037_17092026_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to a cloud rendering system and a cloud rendering method that make it possible to suppress operational costs while providing a highly satisfying viewing experience. This cloud rendering system comprises: a plurality of rendering devices configured to render video content; and a control server that receives, from an edge server, a request signal related to rendering of a region of interest for a display device with respect to the video content, allocates a first rendering device from among the plurality of rendering devices to the display device on the basis of the region of interest, determines whether a first transmission delay of the first rendering device, which corresponds to a change in the region of interest, satisfies a delay time constraint for the display device, and allocates a second rendering device having a second transmission delay smaller than the first transmission delay with respect to the change in the region of interest from among the plurality of rendering devices to the display device on the basis of the result of the determination. The present technology can be applied to a cloud rendering system.
Need to check novelty before this filing date? Find Prior Art

Description

Cloud Rendering System and Cloud Rendering Method

[0001] The present technology relates to a cloud rendering system and a cloud rendering method, and particularly relates to a cloud rendering system and a cloud rendering method capable of suppressing operation costs while providing a highly satisfactory viewing experience.

[0002] Conventionally, services that provide video content in real time from a server on a network to one or more clients are known. Examples of such video content include videos of virtual space live events such as sports matches and music live performances.

[0003] In addition, as a technology related to video content generation, for example, a technology for generating characteristic video by combining and editing video captured by a fixed camera and video captured by a moving camera has been proposed (see, for example, Patent Document 1).

[0004] Japanese Unexamined Patent Application Publication No. 2024-39517

[0005] By the way, in use cases where video content provided by a server on a network can be viewed in a plurality of regions of interest (hereinafter also referred to as ROI (Region of Interest)), it is desired to achieve both providing a highly satisfactory viewing experience and suppressing operation costs.

[0006] For example, generation of multi-view video that simultaneously displays video of a plurality of ROIs (hereinafter also referred to as ROI video) requires video generation processing involving real-time 3D rendering processing. However, such video generation processing imposes an extremely high processing load, and at present, it is difficult to execute the video generation processing with a single rendering server (rendering instance).

[0007] Therefore, one possible approach is to use multiple rendering servers connected via a network to perform 3D rendering and other processes simultaneously, and to have these multiple rendering servers handle the video generation process. However, depending on how the rendering servers are allocated to perform the processing, significant delays may occur in providing content video to clients, potentially reducing user satisfaction with the viewing experience. Furthermore, if there are multiple rendering servers with different processing capabilities and usage costs, inefficient utilization of these servers can lead to increased operational costs.

[0008] This technology was developed in light of these circumstances, and aims to provide a highly satisfying viewing experience while keeping operating costs down.

[0009] One aspect of this technology is a cloud rendering system which comprises a plurality of rendering devices each configured to render a predetermined video content, and a control server configured to receive a request signal from an edge server regarding the rendering of a region of interest of a display device for the predetermined video content, assign a first rendering device from the plurality of rendering devices to the display device based on the region of interest, determine whether a first transmission delay of the first rendering device in response to changes in the region of interest satisfies a delay time constraint for the display device, and, based on the determination result that the first transmission delay does not satisfy the delay time constraint, assign a second rendering device having a second transmission delay smaller than the first transmission delay in response to changes in the region of interest to the display device from the plurality of rendering devices, instead of the first rendering device.

[0010] One aspect of this technology is a cloud rendering method for a cloud rendering system having a plurality of rendering devices, each configured to render a predetermined video content, and a control server, wherein the control server receives a request signal from an edge server regarding the rendering of a region of interest of a display device for the predetermined video content; assigns a first rendering device from the plurality of rendering devices to the display device based on the region of interest; determines whether a first transmission delay of the first rendering device in response to a change in the region of interest satisfies a delay time constraint for the display device; and, based on the determination result that the first transmission delay does not satisfy the delay time constraint, assigns a second rendering device from the plurality of rendering devices to the display device, in place of the first rendering device, the second rendering device having a second transmission delay smaller than the first transmission delay in response to a change in the region of interest.

[0011] One aspect of this technology includes receiving a request signal from an edge server regarding the rendering of a region of interest of a display device for a predetermined video content; assigning a first rendering device to the display device from a plurality of rendering devices based on the region of interest; determining whether a first transmission delay of the first rendering device in response to changes in the region of interest satisfies a delay time constraint for the display device; and, based on the determination that the first transmission delay does not satisfy the delay time constraint, assigning a second rendering device to the display device from the plurality of rendering devices, in place of the first rendering device, the second rendering device having a second transmission delay smaller than the first transmission delay in response to changes in the region of interest.

[0012] This is a diagram showing an example configuration of an information processing system. This is a diagram showing an example configuration of a control server. This is a diagram showing an example configuration of a rendering server. This is a diagram showing an example configuration of an edge server. This is a diagram showing an example configuration of a client. This is a diagram showing an example display screen. This is a diagram explaining latency constraints. This is a diagram explaining transmission latency. This is a diagram explaining C2P latency. This is a diagram showing an example of rendering server assignment. This is a diagram showing an example of rendering server assignment. This is a diagram showing an example of rendering server assignment. This is a diagram explaining how to change the rendering server assignment. This is a flowchart explaining the rendering server determination process. This is a flowchart explaining the assignment process. This is a flowchart explaining the 6DoF assignment process. This is a flowchart explaining the 0DoF assignment process. This is a flowchart explaining the dynamic reallocation process. This is a flowchart explaining the allowable latency determination process. This is a flowchart explaining the video generation process. This is a flowchart explaining the content playback process. This is a diagram showing an example configuration of an information processing system. This is a diagram showing an example configuration of a computer.

[0013] The following describes embodiments to which this technology is applied, with reference to the drawings.

[0014] <First Embodiment> <Example of Information Processing System Configuration> Figure 1 is a diagram showing an example of the configuration of an information processing system to which this technology is applied.

[0015] The information processing system 11 shown in Figure 1 is a content provision system that allows video content to be viewed from multiple perspectives, i.e., multiple ROIs.

[0016] The information processing system 11 includes a cloud computing environment 21, clients 31-1 to 31-3, a performer terminal device 32, and edge servers 41-1 to 41-4. These cloud computing environment 21 to edge servers 41-4 are connected by a network.

[0017] Hereafter, when there is no need to distinguish between clients 31-1 to 31-3, they will simply be referred to as client 31. Similarly, when there is no need to distinguish between edge servers 41-1 to 41-4, they will simply be referred to as edge server 41.

[0018] The cloud computing environment 21 consists of a control server 51 and rendering servers 61-1 to 61-4, which are interconnected via a network, and functions as a cloud rendering system that generates video content to be provided to the client 31. Hereinafter, when there is no need to distinguish between rendering servers 61-1 to 61-4, they will simply be referred to as rendering server 61.

[0019] The information processing system 11 provides a three-dimensional virtual space music live service to a large number of clients 31. Specifically, multiple servers, such as a control server 51 and a rendering server 61 located in the cloud computing environment 21, are used to construct a music live performance that takes place in a three-dimensional virtual space. The video of the music live performance, along with the accompanying audio, is then distributed as video content to the large number of clients 31 via the network.

[0020] In the information processing system 11, a wide variety of devices with diverse performance levels (client devices) are assumed to be clients 31. For example, a device with low processing power, such as a so-called thin client, that cannot perform sufficient real-time 3D rendering processing within the client device itself may also exist as a client 31. Therefore, the generation of video content viewed by client 31 is basically performed not within client 31 (client device), but on an external server, i.e., a server that constitutes the cloud computing environment 21. In the following explanation, an example of video content being a live music performance will be described, but the video content can be anything.

[0021] In the music live venue LV11, constructed within a three-dimensional virtual space, artist avatars exist, and these avatars perform singing and dancing. These avatar performances are filmed by multiple virtual cameras, including virtual cameras 71-1 to 71-3, located within the three-dimensional virtual space. Hereafter, unless there is a need to distinguish between them, the multiple virtual cameras in the three-dimensional virtual space, including virtual cameras 71-1 to 71-3, will simply be referred to as virtual camera 71.

[0022] In a three-dimensional virtual space, images of multiple distinct ROIs (Regions of Interest) are generated by capturing images using numerous virtual cameras 71 positioned and positioned in different directions. An ROI is a region within the three-dimensional virtual space that is targeted for capture by the virtual cameras 71. In other words, the virtual cameras 71 capture the ROI region within the three-dimensional virtual space, resulting in an image with the ROI region as the subject. The ROI changes depending on the position and orientation of the virtual cameras 71 within the three-dimensional virtual space. Hereafter, the image obtained by the virtual cameras 71 will also be referred to as the ROI image.

[0023] In the information processing system 11, video of a predetermined ROI (region of interest) is provided to the client 31 as video content. In other words, the video content is provided as ROI (region of interest) video captured by a virtual camera 71 in a three-dimensional virtual space.

[0024] For example, the space in which a music live performance takes place changes over time, meaning it changes as time passes. However, not all processing required to provide video content needs to be performed on the same server. Specifically, the scene construction process, which determines where and what kind of objects and lighting to place in the three-dimensional virtual space, and the real-time 3D rendering process, which generates the ROI video that client 31 views, do not necessarily need to be performed on the same server.

[0025] In general, real-time 3D rendering (hereinafter simply referred to as 3D rendering) is computationally intensive when generating high-resolution, high-frame-rate video. In addition, in multi-view distribution situations where multiple different ROI videos are distributed to a large number of clients 31, the number of video streams to be handled increases, making the process of generating multiple ROI videos (3D rendering) even more computationally intensive.

[0026] Therefore, the cloud computing environment 21 is configured to include a control server 51 that performs scene construction processing for load balancing, and rendering servers 61-1 to 61-4 that perform 3D rendering processing.

[0027] In the cloud computing environment 21, there are multiple data centers located in different geographical locations, each containing servers of different scales (different processing speeds) or one or more servers. Servers that are physically close to the end clients and directly send and receive data with them are specifically called edge servers. In the example in Figure 1, the cloud computing environment 21 is equipped with a control server 51 and a rendering server 61 as servers. Edge servers 41-1 to 41-4 are also provided as edge servers connected to clients 31, etc.

[0028] While it is possible to perform 3D rendering processing on the edge server 41, considering the case where a single ROI is shared by multiple clients 31, the 3D rendering processing will basically be performed on the rendering server 61.

[0029] The number of rendering servers 61, edge servers 41, and clients 31 provided in the information processing system 11 is not limited to the numbers shown in Figure 1, and any number may be provided.

[0030] The cloud computing environment 21 is configured such that computing resources such as control servers 51 and rendering servers 61 are interconnected via a high-speed network. For example, the cloud computing environment 21 includes rendering servers 61 with different processing capabilities (performance) and rendering servers 61 with different usage costs (costs incurred when using them).

[0031] Each server within the cloud computing environment 21 may be located in a physically distant location from one another. These servers are connected by a high-speed optical network, such as IOWN (Innovative Optical and Wireless Network) (registered trademark) APN (All-Photonics Network). Such high-speed optical networks have the characteristics of low latency and deterministic latency, meaning that the transmission delay between servers is deterministic and does not fluctuate. This characteristic makes it possible to calculate a transmission delay without fluctuations.

[0032] The information processing system 11 can easily determine whether a data transmission path satisfies a predetermined delay time constraint by utilizing the definite low latency characteristic of the high-speed optical network described above. Based on the determination result of whether the delay time constraint is met, the control server 51 determines which rendering server 61 to assign to the client 31.

[0033] In the example shown in Figure 1, the edge server 41 is located outside the cloud computing environment 21, but the edge server 41 may also be included within the cloud computing environment 21.

[0034] The control server 51 receives a connection request from the client 31 and determines which rendering server 61 will be responsible for generating the ROI video to be supplied to the client 31, i.e., which rendering server 61 will be assigned to the client 31.

[0035] The control server 51 notifies the rendering server 61 and the client 31 of the communication partner and establishes communication between the server and the client. The relationship between the rendering server 61 and the client 31 is dynamically changed not only at the time of the initial connection, but also depending on changes in the viewing status of the client 31 and the connection status of other clients 31.

[0036] The control server 51 constructs the appearance (scene) of the three-dimensional virtual space, designed by the performer and constantly changing, in real time, and generates 3D scene data, which is the blueprint of that scene, and supplies it to each rendering server 61. Furthermore, based on the action information supplied by the client 31, the control server 51 reflects the content of the audience's (users') feedback, such as cheers and the movement of penlights, that is, the audience's actions (reactions) to the music live performance, into the scene. In other words, the control server 51 generates 3D scene data that reflects the content of the audience's feedback based on the action information.

[0037] 3D scene data is data used to construct (configure) a three-dimensional virtual space that includes objects such as avatars, which are the subjects of video content. For example, 3D scene data includes three-dimensional object data, which is information about the video and audio of one or more objects placed in the three-dimensional virtual space, and scene description information, which indicates the position and orientation of the objects at each point in time.

[0038] For example, three-dimensional object data includes video object data and audio object data for each object. Video object data is data used to display the image of the object, and consists of, for example, model data and texture data that show the three-dimensional shape of the object. Audio object data is audio data used to play the sound of the object.

[0039] Furthermore, scene description information indicates which objects are located where in the three-dimensional virtual space and which direction they are facing at each point in the video content (music live performance). In other words, scene description information describes what each scene in the video content is like.

[0040] The rendering server 61 is a rendering device that generates ROI images by performing 3D rendering processing. The rendering server 61 receives 3D scene data from the control server 51 and reconstructs the scene according to the 3D scene data. The rendering server 61 also receives ROI information from the client 31 to identify the ROI, performs 3D rendering processing based on the scene reconstruction result and the ROI information, and sends the ROI image obtained as a result of the 3D rendering processing to the client 31. More specifically, the rendering server 61 also performs 3D audio rendering processing based on the scene reconstruction result and the ROI information, and generates audio data for the sound accompanying the ROI image.

[0041] ROI information is information that indicates the ROI (Region of Interest) viewed by client 31. Specifically, for example, ROI information indicates the location and direction in which a user viewing video content on client 31 is looking within the three-dimensional virtual space (music live venue LV11); in other words, it indicates the position and orientation of the virtual camera 71.

[0042] The rendering server 61 also plays a role in uploading action information related to audience (user) actions (reactions) to the live music performance, such as cheering and waving concert lights (penlights), to the control server 51. By reflecting the action information for each user in the scene, the control server 51 can provide users with so-called interactive (two-way) video content. This interactive function allows for a highly immersive experience for users, making them feel as if they are participating in a live music performance.

[0043] The edge server 41 is located at a short distance from the client 31 in terms of network topology, and is a server that exchanges data between the client 31 and the cloud computing environment 21. For example, the edge server 41 receives ROI video data transmitted from the rendering server 61, and transmits the ROI video data to the client 31.

[0044] The edge server 41 performs simple image processing to provide content video to the client 31 as necessary. Specifically, for example, the edge server 41 performs image processing to arrange a plurality of ROI videos received from the rendering server 61 side by side or performs PiP (Picture in Picture) composition to obtain one video, and superimposition processing related to HUD (Head-Up Display), and the like.

[0045] In general, the edge server 41 has GPU (Graphics Processing Unit) resources, and some edge servers 41 are capable of performing 3D rendering processing. However, as described above, there is a case where one ROI video is shared by a plurality of clients 31, that is, a case where one ROI video displayed by a plurality of clients 31 is generated by one server. Therefore, it is not always appropriate to perform 3D rendering processing by the edge server 41 closest to the client 31, and thus 3D rendering processing is basically performed by the rendering server 61 in the cloud computing environment 21.

[0046] The client 31 is a display device (for example, a first display device or a second display device) that receives the ROI video from the edge server 41 and presents the ROI video to a user. More specifically, the client 31 also receives audio obtained by 3D audio rendering processing from the edge server 41, and presents the ROI video and the audio as video content to the user via a display, a speaker, or the like.

[0047] In the information processing system 11, there is no particular limitation on what type of viewing device each client 31 is, and each client 31 may be a device of a different type from other clients 31. Examples of devices constituting the client 31 include a flat panel display fixed at a predetermined position in real space, an HMD (Head Mounted Display) that follows the movement of a user's head, XR (Extended Reality) glasses such as AR (Augmented Reality) glasses, mobile game terminals, and portable video display terminals such as smartphones.

[0048] In the example of Fig. 1, the client 31-1 is a wearable device such as an HMD that can be worn on a user's head. The client 31-2 is constituted by a personal computer or the like, and the client 31-3 is constituted by a smartphone or the like. Each of the clients 31-1 to 31-3 is connected to the cloud computing environment 21 via each of the edge servers 41-1 to 41-3.

[0049] Each client 31 can view only one ROI video, or can display and view a plurality of ROI videos simultaneously. That is, the content viewing mode, such as the number of ROI videos to be displayed, is not fixed.

[0050] As an example, a user can display ROI videos in display formats such as split multi-screen display, in which the entire display screen is divided into a plurality of regions and an ROI video is displayed in each region, and PiP display, in which another ROI video is combined and displayed within one ROI video. In split multi-screen display, a plurality of ROI videos are displayed arranged in a tile pattern. Through video presentation in such display formats as split multi-screen display and PiP display, it is possible to provide users with a viewing experience that cannot be achieved in real space.

[0051] Client 31 has a bidirectional communication function. That is, client 31 not only receives video and audio generated in the cloud computing environment 21, but also uploads action information related to user actions (reactions) such as cheers, stamps, and emotes (expressions of emotion) to the control server 51, and can reflect the user's actions in the three-dimensional virtual space.

[0052] The director's terminal device 32 is an information processing device operated by the director who directs the music live performance, and consists of, for example, a personal computer. The director's terminal device 32 is connected to the cloud computing environment 21 via the edge server 41-4. For example, in response to the director's operations, the director's terminal device 32 transmits scene control information related to the scene, such as the placement of objects and lighting in the music live venue LV11 (three-dimensional virtual space), and the effects to be activated in the music live venue LV11, to the control server 51.

[0053] This section describes how video content is viewed at Client 31.

[0054] In the information processing system 11, users can view ROI images with free viewpoint (6DoF (Degree of Freedom)), 3DoF, or 0DoF, and can also perform multi-viewpoint viewing by displaying multiple different ROI images simultaneously. For example, a user could enlarge one of the multiple ROI images, or switch between displayed ROI images (viewpoint positions) to move between multiple locations.

[0055] 6DoF video refers to video that allows for six degrees of freedom of movement. Users can move their heads left and right (yaw), up and down (pitch), and diagonally (roll), and move forward and backward, up and down, and left and right. In other words, when viewing 6DoF video, users can freely manipulate their viewpoint position and head direction (line of sight) within the three-dimensional virtual space, or in other words, the shooting position and shooting direction of the virtual camera 71.

[0056] 3DoF video is video that allows for three degrees of freedom of movement, allowing the user to move their head left and right (yaw), up and down (pitch), and diagonally (roll). In other words, when viewing 3DoF video, the user can freely manipulate the direction of their own head in the three-dimensional virtual space, or in other words, the shooting direction of the virtual camera 71, but their viewpoint position, or in other words, the shooting position of the virtual camera 71, remains fixed. 0DoF video is fixed viewpoint video in which the user cannot change their viewpoint position or head direction.

[0057] For example, if client 31 is using a device such as an HMD that can view 6DoF (free-viewpoint video) where the viewpoint position and direction of gaze change according to the user's head movements, then low latency is required so that the displayed image immediately follows the head movements in order to reduce so-called VR sickness. For example, if client 31 is an HMD, when the user moves their head (HMD), the virtual camera 71 is operated according to that movement.

[0058] Low latency in 6DoF video means that the M2P latency (Motion-to-Photon latency), which is the delay time from when the user operates the virtual camera 71 until the image of the position and direction after the operation is generated and displayed on the HMD worn by the user, is small.

[0059] When generating ROI images by performing 3D rendering processing on a server in the cloud, i.e., the rendering server 61, rather than rendering within the HMD, M2P latency includes the round-trip data transmission delay time between the client 31 and the rendering server 61. Therefore, the longer the distance between the client 31 and the rendering server 61 in the network topology, the worse the M2P latency becomes. Generally, it is said that data traveling through optical fiber experiences a delay of 500 microseconds per 100 km. Therefore, from the perspective of data transmission delay, it is desirable to perform 3D rendering processing on the rendering server 61, which is located as close as possible to the client 31 in the network topology.

[0060] On the other hand, for 0DoF video, which does not involve user viewpoint control, the ROI video obtained through 3D rendering can be unilaterally transmitted to the client 31. Therefore, M2P latency, which is the response delay to user input, is not an indicator for selecting the optimal rendering server 61 for 0DoF video.

[0061] When viewing 0DoF video, only one-way transmission delay needs to be considered, and the priority of selecting a nearby rendering server 61 is lower compared to 6DoF video. However, if it is necessary to reflect actions from the client 31, such as cheering or the movement of penlights, in the three-dimensional virtual space, the delay time until the user's actions are reflected in the three-dimensional virtual space must be considered.

[0062] Even with fixed-viewpoint video (0DoF video) where the user cannot change their viewpoint or head direction, the music live director or production staff may still move or switch the virtual camera 71. Therefore, even if it is 0DoF video to the user, it is not necessarily fixed-viewpoint camera footage. In this case, if the director's camera operations can be reflected in the rendered video viewed by the director with low latency, the director's intentions can be reflected more accurately. For this reason, when the director moves the virtual camera 71, it is necessary to generate the video on the rendering server 61, which is close to the director, i.e., the director's terminal device 32.

[0063] <Example of Control Server Configuration> Figure 2 shows an example of the configuration of the control server 51. The control server 51 includes a receiving unit 101, a transmitting unit 102, a control unit 103, a scene data storage unit 104, a scene data generation unit 105, and a transmitting unit 106.

[0064] The receiving unit 101 receives various types of information transmitted from the client 31 and the performer terminal device 32 via the network and supplies them to the control unit 103.

[0065] For example, the receiving unit 101 receives an initial connection request from the client 31 requesting to connect to the rendering server 61, viewing mode information regarding the viewing mode on the client 31, and action information. Also, for example, the receiving unit 101 receives scene control information indicating the placement of objects and lighting in the three-dimensional virtual space, the effects to be activated, and various requests from the performer terminal device 32.

[0066] Viewing format information includes, for example, viewing ROI information indicating the ROI to be viewed, display environment information regarding the display environment on client 31, and display configuration information indicating the display configuration of ROI video, such as single-viewpoint viewing or multi-viewpoint viewing.

[0067] The viewing ROI information can be any information that allows the client 31 to identify the ROI of the video being viewed. For example, the viewing ROI information can be ID information indicating a virtual camera 71 that captures the ROI as its subject, or ID information indicating a viewpoint corresponding to the virtual camera 71. The viewing ROI information may also include information indicating the type of video of the ROI being viewed, such as 0DoF video or 6DoF video, i.e., information indicating the degrees of freedom of the ROI video. Furthermore, the viewing ROI information may include ROI information indicating the viewing position and gaze direction (head direction) when viewing 6DoF video or 3DoF video, i.e., the shooting position and shooting direction of the virtual camera 71.

[0068] For example, the display environment information includes information indicating the frame rate of the ROI video to be displayed on client 31, information indicating the resolution of the ROI video, and information indicating whether or not there is interactivity when viewing the ROI video, i.e., whether or not there is action information.

[0069] The transmission unit 102 transmits information supplied from the control unit 103 to the client 31, the performer terminal device 32, and the rendering server 61 via the network. For example, the transmission unit 102 sends connection instructions to the client 31 and the rendering server 61 indicating the connection destination (connection partner) determined in response to the initial connection request, etc. Also, for example, the transmission unit 102 sends responses to the performer in response to various requests from the performer to the performer terminal device 32.

[0070] The control unit 103 controls the operation of the entire control server 51. For example, the control unit 103 determines which rendering server 61 to assign to the client 31, i.e., which is the rendering server 61 that generates the ROI video to be viewed by the client 31, based on viewing format information supplied from the receiving unit 101. Also, for example, the control unit 103 generates connection instructions to the rendering server 61 and the client 31 based on the result of determining the assignment of the rendering server 61, and supplies them to the transmitting unit 102.

[0071] For example, the control unit 103 generates responses to the performer in response to various requests from the performer supplied by the receiving unit 101 and supplies them to the transmitting unit 102. Furthermore, for example, the control unit 103 supplies scene control information, viewing mode information, action information, and the allocation results of the rendering server 61 supplied by the receiving unit 101 to the scene data generation unit 105 and instructs the generation of 3D scene data and rendering parameters.

[0072] Rendering parameters are various parameters (information) used in 3D rendering processing. In other words, rendering parameters are parameters for generating ROI images through 3D rendering processing. For example, rendering parameters include information that can identify the ROI to be generated, such as time information and viewing ROI information, frame rate information, and exposure time information. Note that rendering parameters may not include information that can identify the ROI or frame rate information. In such cases, the rendering server 61 can obtain (receive) ROI information and frame rate information from the client 31 and perform 3D rendering processing.

[0073] Time information indicates the time (playback time) of the video content for which the ROI video will be generated; in other words, it indicates the time for which the ROI video will be generated. Frame rate information indicates the frame rate of the ROI video. Exposure time information indicates the exposure time (shutter speed) of the ROI video.

[0074] The scene data storage unit 104 holds (stores) the 3D scene data supplied from the scene data generation unit 105, and also supplies the stored 3D scene data to the scene data generation unit 105 as appropriate.

[0075] The scene data generation unit 105 constructs a scene based on the 3D scene data for past time periods held in the scene data storage unit 104 and the scene control information supplied from the control unit 103, and generates new 3D scene data for a new time period. Action information is also used as appropriate in the generation of the 3D scene data.

[0076] The scene data generation unit 105 supplies the 3D scene data obtained through the construction process to the transmission unit 106, and also supplies the 3D scene data to the scene data storage unit 104 for storage (updating). In addition, the scene data generation unit 105 generates rendering parameters for each rendering server 61 based on the viewing mode information supplied from the control unit 103 and the allocation results of the rendering servers 61, and supplies them to the transmission unit 106.

[0077] The transmission unit 106 transmits the 3D scene data and rendering parameters supplied from the scene data generation unit 105 to the rendering server 61 via the network.

[0078] <Example of Rendering Server Configuration> Figure 3 shows an example of the configuration of the rendering server 61.

[0079] The rendering server 61 includes a control unit 141, a receiving unit 142, a scene data storage unit 143, a video generation unit 144, and a transmission unit 145.

[0080] The control unit 141 controls the operation of the entire rendering server 61. The receiving unit 142 communicates with the control server 51 and the client 31 via the network. For example, the receiving unit 142 receives 3D scene data from the control server 51 and supplies it to the scene data storage unit 143, or receives rendering parameters from the control server 51 and supplies them to the image generation unit 144.

[0081] The scene data storage unit 143 stores (holds) the 3D scene data supplied from the receiving unit 142 and supplies the stored 3D scene data to the video generation unit 144 as appropriate. The video generation unit 144 generates an ROI image by performing 3D rendering processing based on rendering parameters supplied from the receiving unit 142 and the 3D scene data supplied from the scene data storage unit 143, and supplies it to the transmission unit 145. The transmission unit 145 transmits the ROI image, or more specifically the video data of the ROI image, supplied from the video generation unit 144 to the client 31 via the network.

[0082] <Example of Edge Server Configuration> Figure 4 shows an example of the configuration of edge server 41.

[0083] The edge server 41 includes a receiving unit 171, a control unit 172, a transmitting unit 173, a receiving unit 174, a display image generation unit 175, and a transmitting unit 176.

[0084] The receiving unit 171 receives various types of information from the client 31 via the network and supplies it to the control unit 172. For example, the receiving unit 171 receives initial connection requests, viewing mode information, and action information transmitted from the client 31.

[0085] The control unit 172 controls the operation of the entire edge server 41. For example, the control unit 172 supplies initial connection requests, viewing mode information, and action information supplied from the receiving unit 171 to the transmitting unit 173, and supplies responses to information sent from the client 31 to the control server 51 and rendering server 61, supplied from the receiving unit 174, to the transmitting unit 176. Also, for example, the control unit 172 instructs the display video generation unit 175 to generate video content to be displayed on the client 31.

[0086] The transmitting unit 173 transmits various information, such as initial connection requests, viewing mode information, and action information, supplied from the control unit 172, to the control server 51 and the rendering server 61 via the network. The receiving unit 174 receives various information from the control server 51 and the rendering server 61 via the network and supplies it to the control unit 172 and the display image generation unit 175. For example, the receiving unit 174 receives a connection instruction from the control server 51 to the rendering server 61 for the client 31 and supplies it to the control unit 172, or receives ROI video transmitted from the rendering server 61 and supplies it to the display image generation unit 175.

[0087] The display video generation unit 175 generates a video of the video content to be displayed on the client 31 (hereinafter also referred to as the display video) based on one or more ROI videos supplied from the receiving unit 174, in accordance with instructions from the control unit 172, and supplies it to the transmission unit 176. For example, the display video will display one or more ROI videos. Specifically, for example, if single-viewpoint viewing is performed on the client 31, a display video containing one ROI video is generated. Also, for example, if multi-viewpoint viewing is performed on the client 31, a display video is generated in which multiple ROI videos are arranged in a predetermined display format such as split multi-screen display or picture-in-picture display.

[0088] The transmission unit 176 transmits various information (responses), such as connection instructions, supplied from the control unit 172 to the client 31 via the network. The transmission unit 176 also transmits the display video, or more specifically the video data of the display video, supplied from the display video generation unit 175 to the client 31 via the network.

[0089] <Example of Client Configuration> Figure 5 shows an example of the configuration of client 31.

[0090] The client 31 is connected to an input unit 201 operated by the user and a display unit 202 which displays the image to be shown.

[0091] For example, the input unit 201 may consist of a mouse, keyboard, controller, or a touch panel superimposed on the display 202, and supplies signals to the client 31 in accordance with the user's operation. Alternatively, the input unit 201 may consist of various sensors that detect the user's movements, and the detection results of the user's movements may be output to the client 31 as user input. The input unit 201 and the display 202 may be provided on the client 31.

[0092] For example, the user performs an input operation via the input unit 201 to change the ROI (Region of Interest) in the three-dimensional virtual space (user input). The user input operation to change the ROI can be anything, such as a gesture operation, a button operation, a touch operation, or a combination of multiple input operations.

[0093] The client 31 includes a control unit 211, a user input acquisition unit 212, a generation unit 213, a transmission unit 214, a reception unit 215, and a video display control unit 216.

[0094] The control unit 211 controls the operation of the entire client 31. The user input acquisition unit 212 supplies signals to the generation unit 213 that correspond to the user's input to the client 31, which is supplied from the input unit 201. For example, the user input acquisition unit 212 acquires information (signals) related to the user's movements, that is, information related to changes in ROI in the three-dimensional virtual space, and supplies it to the generation unit 213.

[0095] The generation unit 213 generates various types of information based on signals supplied from the user input acquisition unit 212 and supplies them to the transmission unit 214. For example, the generation unit 213 generates initial connection requests, action information, and viewing mode information in response to user operations on the input unit 201 and supplies them to the transmission unit 214. Also, for example, the generation unit 213 identifies changes in ROI based on signals supplied from the user input acquisition unit 212, generates ROI information corresponding to those changes, and supplies it to the transmission unit 214.

[0096] The transmission unit 214 transmits various types of information supplied from the generation unit 213 to the control server 51 and the rendering server 61 via the network. For example, the transmission unit 214 transmits initial connection requests, action information, and viewing mode information to the control server 51, and ROI information to the rendering server 61.

[0097] The receiving unit 215 receives various types of information transmitted from the control server 51 and the rendering server 61. For example, the receiving unit 215 receives information such as connection instructions transmitted from the control server 51 and supplies it to the control unit 211, or receives video data of the display image transmitted from the edge server 41 and supplies it to the video display control unit 216.

[0098] The video display control unit 216 supplies the video data of the display image received from the receiving unit 215 to the display 202, causing the display image to be displayed on the display 202. More specifically, the receiving unit 215 receives not only the video data of the display image but also the audio data of the sound associated with that display image. This audio data is supplied to a speaker (not shown), and sound is played back based on the audio data.

[0099] Note that the configuration of client 31 shown in Figure 5 is merely an example, and the configuration of client 31 can be any configuration. As mentioned above, client 31 may be an HMD, smartphone, or mobile game terminal in which the input unit 201 and display 202 functions are integrated.

[0100] Figure 6 shows an example of the display image shown on client 31.

[0101] In the example shown in Figure 6, the main image, ROI image P11, is displayed across almost the entire screen. In this example, ROI image P11 is a 6DoF image of the music live venue LV11 as seen from the front of the stage. In other words, ROI image P11 is an image captured by a virtual camera 71 that can be operated by the user.

[0102] Furthermore, in the diagram of ROI video P11, four sub-videos, ROI videos P12 to P15, are displayed in picture-in-picture (PiP) format at the top. In other words, the displayed video as a whole has a multi-viewpoint screen configuration in which the four ROI videos P12 to P15 are displayed in picture-in-picture format on a single ROI video P11.

[0103] For example, ROI video P12 displays a 0DoF ROI video captured by a virtual camera 71 operated by the director, ROI video P13 displays a 0DoF ROI video captured by a virtual camera 71 located behind the stage at the music live venue LV11, ROI video P14 displays a 0DoF ROI video captured by a virtual camera 71 located above the stage at the music live venue LV11, and ROI video P15 displays a 0DoF ROI video captured by a virtual camera 71 located to the left of the stage at the music live venue LV11.

[0104] By viewing a display that simultaneously shows multiple ROI (Region of Interest) images, users can watch the live music performance taking place at the music venue LV11.

[0105] Furthermore, in the displayed video, it is possible for the user to specify a particular sub-video, such as ROI video P12, and have that sub-video displayed as the main video, or to display only one ROI video selected from the main video and sub-videos as the displayed video.

[0106] <Regarding the allocation of rendering servers> The allocation of rendering servers 61 to clients 31, performed by the control server 51 (control unit 103), that is, the determination of the rendering server 61 that will perform the 3D rendering processing of the ROI video viewed by client 31, will be explained.

[0107] When presenting video content, multiple ROI videos (ROI videos) are generated by 3D rendering. While it is necessary to generate multiple ROI videos simultaneously, 3D rendering generally has a high processing load, and a single rendering server 61 (rendering instance) may not be able to generate videos while maintaining the set frame rate. Therefore, the information processing system 11 generates ROI videos to be distributed (supplied) to each client 31 by distributing the processing load across multiple servers. In other words, multiple rendering servers 61 are used to generate multiple ROI videos.

[0108] In this case, for example, for an ROI that can be moved 6DoF by user operation, high-speed response to changes in viewpoint movement and gaze direction (head direction) is important, and a rendering server 61 located close to the client 31 in the network topology is assigned to the client 31. By having the rendering server 61 perform 3D rendering processing, it is possible to provide ROI video (6DoF video) with a short response delay (M2P latency) to user operation.

[0109] Furthermore, for example, ROI video with high frame rates but no user interaction, a rendering server 61 located relatively close to the client 31 in the network topology is assigned. Relatively close here means farther than in the case of 6DoF video as described above, but close enough to present the ROI video within an acceptable response delay range. By having 3D rendering processing performed on such a rendering server 61, the time it takes for the client 31 to receive the ROI video can be shortened. This makes it possible to provide the user with video that takes advantage of the high frame rate characteristics, reflecting scene changes in the three-dimensional virtual space in the video viewed by the client 31 (ROI video) with low latency.

[0110] Furthermore, ROI video with a low frame rate and no user interaction (two-way communication) may have a longer latency than high-frame-rate video, but this latency is less noticeable. Therefore, for low-frame-rate ROI video, a rendering server 61 located far away from client 31 in the network topology can be assigned to client 31. In this case, the usage cost (cost at the time of use) is low, and by using a rendering server 61 located far away from client 31 in the network topology, server usage costs can be reduced, i.e., operating costs can be suppressed.

[0111] In addition, for example, ROI video that is shared with multiple users but does not involve user interaction (two-way communication), i.e., ROI video that is viewed by multiple users, it is conceivable to use a rendering server 61 that is equidistant in the network topology from each user's client 31 that views (shares) the ROI video. By doing so, the ROI video can be provided to each user in a way that does not cause significant differences in the viewing experience for each user, such as differences in latency.

[0112] As described above, in the information processing system 11 that generates multiple ROI videos, such as multi-viewpoint videos, in the cloud computing environment 21, the allocation of rendering servers 61 is determined according to the characteristics of the video content viewing format on the client 31. In this way, it is possible to efficiently utilize the limited resources of the cloud computing environment 21 (rendering servers 61 and control servers 51) while maintaining user satisfaction with the viewing experience, thereby suppressing operating costs. In other words, it is possible to suppress operating costs while providing a highly satisfying viewing experience.

[0113] Further explanation will be provided regarding the allocation of rendering server 61.

[0114] For example, the control server 51 (control unit 103) determines whether a specific type of delay time that occurs when ROI video is generated by a rendering server 61 that is a candidate for assignment satisfies the delay time constraint. Based on this determination, the rendering server 61 to be assigned to the client 31 is determined.

[0115] Referring to Figures 7 to 9, various delay times used to determine whether or not the delay time constraint is met will be explained. Note that in Figures 7 to 9, the same reference numerals are used for parts corresponding to those in Figure 1, and their explanations will be omitted as appropriate.

[0116] The delay times used to determine whether or not the delay time constraint is met are, for example, three types of delay times: transmission delay time, M2P latency, and C2P latency (Click-to-Photon latency), as shown in Figure 7.

[0117] The transmission delay time is the unidirectional delay time between the control server 51 and the client 31, as shown by arrow A11. As will be explained in more detail later, the transmission delay time includes not only the data transmission time but also the processing time in the control server 51 and the rendering server 61.

[0118] M2P latency is the round-trip delay time (response delay time) between client 31 and rendering server 61, as shown by arrow A12. M2P latency is the response delay time from when a user performs an operation on client 31 that changes the ROI without affecting the three-dimensional virtual space, until the image reflecting that operation reaches client 31. As will be explained in more detail later, M2P latency includes not only the data transmission time but also the processing time on rendering server 61.

[0119] C2P latency is the round-trip delay time (response delay time) between the client 31 and the control server 51, as shown by arrow A13. In other words, C2P latency is the response delay time from when a user performs an operation on the client 31 that affects the three-dimensional virtual space, until that operation is reflected in the three-dimensional virtual space and the resulting image of the three-dimensional virtual space reaches the client 31. As will be described in more detail later, C2P latency includes not only the data transmission time but also the processing time in the control server 51 and the rendering server 61.

[0120] For example, the control server 51 (control unit 103) classifies the viewing modes of video content on the client 31 into the following three types of viewing modes.

[0121] (Viewing Mode V1) Viewing without changes in viewpoint or gaze from the client 31 (Viewing 0DoF video) (Viewing Mode V2) Viewing of free-viewpoint video (Viewing 6DoF video) (Viewing Mode V3) Viewing with interactivity (Interactivity where user actions are reflected in the scene)

[0122] Viewing mode V1 is a mode of viewing 0DoF video that does not have interactivity. More specifically, viewing mode V1 includes viewing a single 0DoF video, viewing pre-made 0DoF videos produced by the distribution organizer, and multi-viewpoint viewing of multiple 0DoF videos simultaneously. For example, in multi-viewpoint viewing of multiple 0DoF videos, a collection of images captured from different positions and directions by multiple virtual cameras 71 in a three-dimensional virtual space is viewed. In this case, the client 31 cannot control the position or direction of the virtual cameras 71.

[0123] If the viewing mode of the video content on client 31 is viewing mode V1, the control server 51 determines whether the one-way transmission delay time satisfies the delay time constraint and decides which rendering server 61 to assign to client 31. In this case, even when viewing from multiple viewpoints, if the assignment is made in such a way that the delay time constraint is satisfied, multiple 0DoF videos can be synchronized with sufficient accuracy.

[0124] Furthermore, in the case of multi-viewpoint viewing, in order to synchronize multiple 0DoF video feeds, the display timing should, in principle, be aligned with the 0DoF video feed with the largest transmission delay. Therefore, the upper limit of the transmission delay time of the latest received 0DoF video feed should be defined as a delay time constraint. By doing so, the occurrence of synchronization errors can be suppressed, and a highly satisfying viewing experience can be provided to the user.

[0125] In viewing mode V2, 6DoF video is viewed, allowing the client 31 to control the position and orientation of the virtual camera 71 within the three-dimensional virtual space, i.e., the shooting position and shooting direction. Viewing mode V2 is a viewing mode that does not involve bidirectional communication such as sending and receiving action information. The user can move the virtual camera 71 in the three-dimensional virtual space in each direction of the position coordinates (x, y, z), and rotate the shooting direction of the virtual camera 71 in the yaw, pitch, and roll directions. In other words, since the user can control the virtual camera 71 with six degrees of freedom, video in which the virtual camera 71 (viewpoint and line of sight) can be controlled with six degrees of freedom is also called 6DoF video.

[0126] Possible means of operating the virtual camera 71 on the client side 31 include a keyboard or mouse device, a device held and operated by the user, an HMD (head-mounted display) worn on the user's head to sense the position and orientation of the body (head), and a touch panel. In the classification of viewing modes, although there are no 6 degrees of freedom, viewing 3DoF video captured by a 3-degree-of-freedom virtual camera 71 that can only control the shooting direction, for example, will also be included in viewing mode V2.

[0127] When the viewing mode of the video content on client 31 is viewing mode V2, in addition to the transmission delay time of one-way transmission, M2P latency is also considered as responsiveness to the user's head movements, i.e., the operation of the virtual camera 71. In other words, responsiveness to user operation is also important. Therefore, the control server 51 determines whether the M2P latency, which indicates responsiveness, satisfies the delay time constraint, in addition to the transmission delay time, and decides which rendering server 61 to assign to client 31. In this case, if the assignment is made so that the delay time constraint is satisfied, 6DoF video can be presented with a sufficiently low response delay to changes in ROI.

[0128] Viewing mode V3 is a bidirectional viewing mode in which, for example, actions such as cheers can be sent from client 31 to the music live venue LV11, that is, action information is sent from client 31 to control server 51, and user actions can be reflected in the scene. In this case, it is necessary to select rendering server 61 so as to satisfy the constraints on response delay time (responsiveness) from when a user action occurs until that action is reflected in the scene.

[0129] Therefore, when the viewing mode of the video content on client 31 is viewing mode V3, the control server 51 determines whether the C2P latency, which indicates responsiveness, satisfies the delay time constraint, and decides which rendering server 61 to assign to client 31. This suppresses a decrease in the response speed to user actions and provides users with a highly satisfying viewing experience.

[0130] For example, the criteria for determining whether the delay time constraint is met, i.e., the allowable delay time, will be different for each of the viewing modes V1 to V3.

[0131] In the aforementioned viewing mode V1, since the transmission is unidirectional from the control server 51 to the client 31, there are basically no constraints on the response delay time, and only the transmission delay time needs to be considered. Furthermore, in viewing mode V2, the viewpoint position and head direction, i.e., ROI, can be freely changed within the three-dimensional virtual space, but this does not affect the three-dimensional virtual space (scene).

[0132] The videos viewed using viewing modes V1 and V3 include pre-produced videos created by the distribution organizer. Such videos are also called director's cuts and are videos from the perspective the organizer wants viewers to see (what they consider to be the best perspective for the production). Client 31 cannot change the content of these videos (such as the direction of the viewer's gaze).

[0133] As described above, for each viewing mode, indicators suitable for the characteristics of that viewing mode, namely the transmission delay time and response delay time defined for that viewing mode, are used to determine which rendering server 61 to assign to the client 31.

[0134] Since different factors influence the user's viewing experience depending on the viewing method, assigning the rendering server 61 to a server using metrics corresponding to the viewing method on the client 31 can improve responsiveness to user operations and enhance the synchronization accuracy of multiple ROI videos. This allows for a more satisfying viewing experience for the user.

[0135] Note that the videos in viewing modes V1 to V3 are not mutually exclusive and may be viewed simultaneously. For example, in multi-viewpoint viewing, multiple ROI videos, including 6DoF videos, may be viewed at the same time.

[0136] The calculation methods for each delay time used to determine whether or not the delay time constraints are met will be explained. First, referring to Figure 8, the calculation of the one-way transmission delay time until the ROI video reaches client 31 will be explained.

[0137] Figure 8 shows the transmission time between each device and the processing time at each device. Specifically, Figure 8 shows the scene generation time Tc(n), transmission time Tc(n,m), 3D rendering processing time Tr(n), transmission time Te(n,m), and transmission time Tp(m) as transmission and processing times.

[0138] The scene generation time Tc(n) represents the time required to generate 3D scene data in the control server with sign n, that is, the time required to construct a scene in a three-dimensional virtual space. Therefore, for example, the scene generation time in control server 51 is Tc(51). For the sake of simplicity, the information processing system 11 is shown as having only one control server, 51, but there may be multiple control servers.

[0139] The transmission time Tc(n,m) represents the transmission time from the control server with sign n to the rendering server 61 with sign m, until the 3D scene data is delivered. Therefore, for example, the transmission time of the 3D scene data from the control server 51 to the rendering server 61-1 is Tc(51,61-1).

[0140] The 3D rendering processing time Tr(n) represents the processing time for the 3D rendering process performed in rendering server 61, where the sign is n. Therefore, for example, the 3D rendering processing time for rendering server 61-1 is Tr(61-1).

[0141] The transmission time Te(n,m) represents the data transmission time between the server with sign n and the server with sign m. Here, a server refers to any of the control server 51, rendering server 61, or edge server 41.

[0142] Therefore, for example, the transmission time between rendering server 61-2 and rendering server 61-1 is Te(61-2,61-1), and the transmission time between rendering server 61-1 and edge server 41-1 is Te(61-1,41-1). Also, for example, the data transmission time between control server 51 and edge server 41-4 is Te(51,41-4).

[0143] The transmission time Tp(m) represents the data transmission time between the edge server 41 and the client 31, where the sign is m. Therefore, for example, the data transmission time between the edge server 41-1 and the client 31 is Tp(41-1).

[0144] In the information processing system 11, there is a delay between the time the ROI image generated by the 3D rendering process on the rendering server 61 reaches the client 31.

[0145] For example, the time from when the control server 51 starts constructing the scene, that is, when it starts generating 3D scene data, until the ROI image (display image) reaches the client 31 is defined as the one-way transmission delay time. In this case, the one-way transmission delay time shown by arrow A11 in Figure 7 is the sum of the transmission time and processing time along the data transmission path, and can be expressed as Tc(n)+Tc(n,m)+Tr(n)+Te(n,m)+Tp(m).

[0146] Regarding the transmission time Te(n,m), for the sake of model simplification, the delay time can be the same for both the outbound and return paths, or if the delays differ between the outbound and return paths, they can be considered separately as Te(n,m) and Te(m,n). Furthermore, when transmitting data to the edge server 41 to which the client 31 is connected via multiple rendering servers 61, the number of Te(n,m) terms in the transmission delay time will increase according to the number of rendering servers 61 the data passes through.

[0147] Furthermore, the network transmission delay values ​​in the information processing system 11, i.e., the transmission time values ​​such as Te(n,m), are considered deterministic in the case of, for example, IOWN's (registered trademark) APN. In contrast, in the general internet, the transmission delay values ​​change moment by moment, so it is sufficient to determine the transmission time values ​​such as Te(n,m) by speed measurement or other means immediately before determining the data route.

[0148] A specific example of allocating rendering servers 61 based on one-way transmission delay time will be described. Here, we will describe an example of determining which rendering server 61 will generate the ROI video to be distributed to client 31-3 in Figure 8, that is, which rendering server 61 to assign.

[0149] For example, when all four rendering servers 61-1 to 61-4 are available, the transmission delay time until the ROI image generated by each rendering server 61 is delivered to the client 31-3 is as follows. That is, if the transmission delay time when the ROI image is generated in each of the rendering servers 61-1 to 61-4 is denoted as transmission delay time (61-1) to transmission delay time (61-4), then each transmission delay time can be expressed as follows.

[0150] Transmission delay time (61-1) = Tc(51) + Tc(51,61-1) + Tr(61-1) + Te(61-1,61-4) + Te(61-4,41-3) + Tp(41-3) Transmission delay time (61-2) = Tc(51) + Tc(51,61-2) + Tr(61-2) + Te(61-2,61-4) + Te(61-4,41-3) + Tp(41-3) Transmission delay time (61-3) = Tc(51) + Tc(51,61-3) + Tr(61-3) + Te(61-3,41-3) + Tp(41-3) Transmission delay time (61-4) =Tc(51)+Tc(51,61-4)+Tr(61-4)+Te(61-4,41-3)+Tp(41-3)

[0151] For example, the transmission delay time (61-1) is the delay time when data is transmitted from the control server 51 to the client 31-3 via the rendering server 61-1, rendering server 61-4, and edge server 41-3.

[0152] Rendering servers 61-1 and 61-2 are geographically separated from client 31-3. Therefore, in the example shown in Figure 8, when data is transmitted from rendering server 61-1 or rendering server 61-2 to edge server 41-3, the network configuration involves passing through rendering server 61-4. Consequently, the transmission delay time (61-1) and transmission delay time (61-2) include terms for the transmission time between rendering server 61-4 and the server, Te(61-1,61-4) and Te(61-2,61-4).

[0153] In reality, the transmission delay time (61-1) and transmission delay time (61-2) include the time it takes for the rendering server 61-4 to receive data from rendering servers 61-1 and 61-2 and transfer it to the edge server 41-3, i.e., the processing time within the rendering server 61-4. However, since the processing time within the rendering server 61-4 is sufficiently small compared to the processing time for 3D rendering, which is the main role of the rendering server 61, we will assume that the processing time related to the transfer is 0 for the sake of simplifying the model. Similarly, the internal processing time from when the edge server 41-3 receives data from the rendering server 61 until it transfers it to the client 31-3 is also assumed to be 0.

[0154] The control server 51 (control unit 103) compares the transmission delay times (61-1) to (61-4) calculated for each rendering server 61, and selects the rendering server 61-4 with the shortest transmission delay time as the rendering server 61 that generates the ROI video to be distributed to the client 31-3. In other words, rendering server 61-4 is assigned to client 31-3.

[0155] This section describes a specific example of assigning a rendering server 61 based on M2P latency. Here, we will describe an example of determining which rendering server 61 to assign to client 31-3 in Figure 8.

[0156] When assigning a rendering server 61 to a client 31-3 based on M2P latency, the latency of each rendering server 61 should be compared via the edge server 41-3 to which the client 31-3 is connected.

[0157] The delay time from when client 31-3 sends information indicating the user's head movement, i.e., ROI information, to when rendering server 61 receives that information, generates 6DoF video corresponding to the head movement, and delivers it to client 31-3 is M2P latency.

[0158] If we denote the M2P latency when 6DoF video (ROI video) is generated on rendering servers 61-1 to 61-4 as M2P latency(61-1) to M2P latency(61-4), then each M2P latency can be expressed as follows.

[0159] M2P latency (61-1) = Tp(41-3)+Te(41-3,61-4)+Te(61-4,61-1)+Tr(61-1)+Te(61-1,61-4)+Te(61-4,41-3)+Tp(41-3) M2P latency(61-2) =Tp(41-3)+Te(41-3,61-4)+Te(61-4,61-2)+Tr(61-2)+Te(61-2,61-4)+Te(61-4,41-3)+Tp(41-3) M2P latency(61-3) =Tp(41-3)+Te(41-3,61-3)+Tr(61-3)+Te(61-3,41-3)+Tp(41-3) M2P latency(61-4) =Tp(41-3)+Te(41-3,61-4)+Tr(61-4)+Te(61-4,41-3)+Tp(41-3)

[0160] The control server 51 (control unit 103) can assign the rendering server 61 with the shortest M2P latency among the rendering servers 61 that satisfy the delay time constraint for one-way transmission delay time to client 31-3. This makes it possible to provide the user with 6DoF video that tracks the user's movements with low latency.

[0161] This section describes a specific example of assigning rendering servers 61 based on C2P latency. Here, we will describe an example of determining which rendering server 61 to assign to client 31-3, as shown in Figure 9.

[0162] When assigning rendering servers 61 to clients 31-3 based on C2P latency, the delay time from when client 31-3 sends action information until the ROI image generated by each rendering server 61 reaches client 31-3 should be compared. In this case, the action information sent from client 31-3 is received by the control server 51, and 3D scene data reflecting the action information is transmitted to the rendering server 61. The rendering server 61 then uses the 3D scene data to perform 3D rendering processing, and the resulting ROI image is transmitted to client 31-3.

[0163] For example, when allocating a rendering server 61 based on C2P latency, the delay time can be considered separately for the outbound and return paths as data transmission routes.

[0164] There are multiple paths for transmitting action information from client 31-3 to control server 51, which constitutes the outbound path of the data transmission route. When determining the outbound path, even if a candidate path is the shortest distance in terms of network topology, that path is not necessarily the path that minimizes C2P latency in terms of round-trip time. Furthermore, when the rendering server 61 is shared with other clients 31, or when there are significant differences in processing time within each rendering server 61, the outbound and return paths are not necessarily the same.

[0165] Since the control server 51 is connected to the rendering servers 61, if we focus on each rendering server 61 and calculate the forward path delay time passing through those rendering servers 61, we get the following. That is, if we denote the forward path delay time when passing through rendering servers 61-1 to 61-4 as forward path delay time (61-1) to forward path delay time (61-4), then each forward path delay time can be expressed as follows. Note that forward path delay time (61-m) is the delay time when passing through rendering server 61-m (m = 1, 2, 3, 4) which is directly connected to the control server 51.

[0166] Outbound journey delay time (61-1) = Tp(41-3) + Te(41-3,61-4) + Te(61-4,61-1) + Tc(61-1,51) Outbound journey delay time (61-2) = Tp(41-3) + Te(41-3,61-4) + Te(61-4,61-2) + Tc(61-2,51) Outbound journey delay time (61-3) = Tp(41-3) + Te(41-3,61-3) + Tc(61-3,51) Outbound journey delay time (61-4) = Tp(41-3) + Te(41-3,61-4) + Tc(61-4,51)

[0167] Similarly, for the return journey, if we denote the return journey delay time when ROI images are generated at rendering servers 61-1 to 61-4 as return journey delay time (61-1) to return journey delay time (61-4), then each return journey delay time can be expressed as follows.

[0168] Return trip delay time (61-1) = Tc(51) + Tc(51,61-1) + Tr(61-1) + Te(61-1,61-4) + Te(61-4,41-3) + Tp(41-3) Return trip delay time (61-2) = Tc(51) + Tc(51,61-2) + Tr(61-2) + Te(61-2,61-4) + Te(61-4,41-3) + Tp(41-3) Return trip delay time (61-3) = Tc(51) + Tc(51,61-3) + Tr(61-3) + Te(61-3,41-3) + Tp(41-3) Return trip delay time (61-4) =Tc(51)+Tc(51,61-4)+Tr(61-4)+Te(61-4,41-3)+Tp(41-3)

[0169] Focusing solely on the return delay time, the return delay time (61-1) to the return delay time (61-4) are the same as the transmission delay time (61-1) to the transmission delay time (61-4) mentioned above.

[0170] Since the data transmission path is a combination of the forward and return paths, in the example shown in Figure 9, there are a total of 16 (4 x 4) possible transmission paths, and the C2P latency for each transmission path is the sum of the forward delay time and the return delay time. Therefore, by selecting and combining the forward path with the smallest forward delay time and the return path with the smallest return delay time, it is possible to identify the transmission path with the minimum (shortest) C2P latency.

[0171] For example, suppose the total delay time, i.e., the sum of the forward and return delay times, is minimized when combining the forward delay time (61-3) via rendering server 61-3 and the return delay time (61-4) via rendering server 61-4. In such a case, the combination of forward and return paths becomes the transmission path that minimizes C2P latency. In this example, rendering server 61-4 on the transmission path that constitutes the return path is assigned to client 31-3, and 3D rendering processing is performed at rendering server 61-4.

[0172] By determining the transmission path in this manner, ROI video reflecting user actions can be provided to the user with low latency.

[0173] The allocation of rendering servers 61, i.e., the transmission path, is determined based on various delay times such as transmission delay time and M2P latency. In addition, the determination of the transmission path (rendering server 61) may also take into account the frame rate and resolution of the ROI video, whether or not it is shared with other clients 31, changes in ROI (viewing mode), and the cost of using the rendering server 61.

[0174] Refer to Figures 10 to 14 to see specific examples of the allocation of rendering servers 61, i.e., the determination of transmission paths. Note that in Figures 10 to 13, some devices and connection relationships constituting the information processing system 11 have been omitted for clarity.

[0175] The example shown in Figure 10 illustrates a data transmission path when a client 31-3 requests a connection with high frame rate video from the control server 51, and the optimal rendering server 61 is assigned to that connection request.

[0176] In this example, the rendering server 61-4 has high processing power and is capable of generating high-frame-rate ROI video as requested by the connection request. The control server 51 determined that when the rendering server 61-4 generates high-frame-rate ROI video, the delay time (transmission delay time) when transmitting that ROI video to the client 31-3 satisfies the delay time constraint, and therefore the rendering server 61-4 is assigned to the client 31-3.

[0177] Furthermore, the determination of whether the delay time constraint is met may be made by targeting a rendering server 61 that has the processing power to generate ROI video at the requested frame rate and resolution, or by considering the frame rate and resolution when determining whether the delay time constraint is met. For example, the 3D rendering processing time Tr(n) may be determined by considering the frame rate and resolution of the ROI video. Alternatively, a threshold value used to determine whether the delay time constraint is met, i.e., the delay time that serves as the threshold, may be determined (changed) depending on the frame rate and resolution of the ROI video.

[0178] As described above, the transmission delay time includes the 3D rendering processing time Tr(n). In other words, the transmission delay time reflects the processing capacity of the rendering server 61. Therefore, if the transmission delay time when transmitting ROI video to clients 31-3 satisfies the delay time constraint, then the requirements should also be met from the perspective of the processing capacity of the rendering server 61. In other words, the rendering server 61 should have the processing capacity to provide ROI video to clients 31-3 with a sufficiently short delay time. For this reason, the control server 51 does not need to compare the processing capacity of each rendering server 61 individually.

[0179] The example shown in Figure 11 illustrates a data transmission path when a client 31-3 requests a connection with low frame rate video from the control server 51, and the optimal rendering server 61 is assigned to that connection request. In this example, since the connection request is for low frame rate video, the delay time constraint is set more loosely than in the example in Figure 10.

[0180] In this example, rendering server 61-3 has lower processing power than rendering server 61-4, but because the connection request is for low frame rate video, processing can be done by rendering server 61-3 as well as rendering server 61-4.

[0181] The control server 51 determines whether the delay time (transmission delay time) when transmitting ROI video to client 31-3 when ROI video is generated at a low frame rate satisfies the delay time constraint for each assignable rendering server 61. For example, suppose rendering server 61-3 and rendering server 61-4 are determined to satisfy the delay time constraint.

[0182] In this case, the control server 51 decides to allocate the less expensive rendering server 61-3 to client 31-3 because it can process the request even with rendering server 61-3, which has lower processing power, and rendering server 61-3 has lower usage costs than rendering server 61-4.

[0183] Figure 12 shows an example in which a rendering server 61 is assigned to each of the two clients 31.

[0184] In this example, client 31-3 first starts viewing ROI "A," that is, viewing ROI video A, and rendering server 61-4 is assigned to client 31-3. In other words, client 31-3 connects to rendering server 61-4 and receives ROI video A from rendering server 61-4.

[0185] Subsequently, when client 31-1 attempted to start viewing ROI "B," which is different from ROI "A," that is, to start viewing ROI video B, and requested a connection to control server 51, rendering server 61-1 was assigned to client 31-1.

[0186] Since clients 31-1 and 31-3 view different ROI videos, each client 31 is assigned a different rendering server 61. In particular, each client 31 is assigned a rendering server 61 that satisfies latency constraints and is geographically close to the client 31.

[0187] Figure 13 shows an example where, for example, from the state shown in Figure 12, client 31-1 begins to view ROI video A in addition to ROI video B simultaneously. In this case, the control server 51 dynamically optimizes the allocation of rendering servers 61.

[0188] Specifically, since client 31-1 will also view ROI video A in addition to client 31-3, the control server 51 determines which rendering server 61 to assign to client 31-1 and client 31-3 for ROI video A.

[0189] In this example, it was determined that if rendering server 61-2 is responsible for generating ROI video A, the latency constraints will be satisfied for both client 31-1 and client 31-3. Therefore, the same (common) rendering server 61-2 is assigned to both client 31-1 and client 31-3. Consequently, client 31-1 and client 31-3 will share ROI video A generated by a single rendering server 61-2.

[0190] In this case, the assignment for client 31-3 will be changed from rendering server 61-4 to rendering server 61-2, which is geographically further away. However, since rendering server 61-2 satisfies the latency constraint for both clients 31-1 and 31-3, generating ROI image A on a single rendering server 61-2 is more efficient in utilizing rendering server 61 than generating the same ROI image A on two separate rendering servers 61. In other words, the limited resources (rendering server 61) can be used more efficiently, and operating costs can be reduced.

[0191] After the allocation is determined, the ROI video A generated by rendering server 61-2 is transmitted to client 31-3 via rendering server 61-4 and edge server 41-3. In addition, the ROI video A generated by rendering server 61-2 is also transmitted to rendering server 61-1.

[0192] Rendering server 61-1 generates ROI video B and transmits the generated ROI video B and ROI video A transmitted from rendering server 61-2 to client 31-1. At this time, ROI video A and ROI video B are transmitted from rendering server 61-1 to edge server 41-1. Then, for example, edge server 41-1 generates a display video that simultaneously displays ROI video A and ROI video B, and transmits this display video to client 31-1.

[0193] The relationship between the ROI and the rendering server 61 is not static; it is possible to dynamically change this relationship according to the user's viewing situation (viewing style). For example, when a user selects a low-resolution ROI image displayed in picture-in-picture and enlarges it to display a high-resolution image, the rendering server 61 that generates (renders) that ROI image may be dynamically switched to a rendering server 61 with higher rendering processing capabilities. By doing so, the ROI image can be provided without reducing the resolution or frame rate.

[0194] Figure 14 shows an example of a change in the viewing mode of client 31. Note that parts in Figure 14 that correspond to those in Figure 6 are denoted by the same reference numerals, and their explanations are omitted as appropriate.

[0195] For example, in client 31, ROI video P11 is displayed as the main video, as shown on the left side of the diagram, and multi-view viewing is performed with ROI videos P12 to P15 displayed in picture-in-picture (PiP) relative to ROI video P11.

[0196] At this point, the main ROI video P11 is a high-resolution, high-frame-rate video, and therefore has high processing power. It is assumed that it is generated by rendering server 61, which is located close to client 31 in the network topology. ROI videos P12 to P15 are low-resolution, medium-frame-rate videos. Therefore, these ROI videos have low processing power and are generated by rendering server 61, which is located far from client 31 in the network topology.

[0197] From this state, suppose the user selects ROI video P14, which is being displayed in PiP, and then enlarges that ROI video P14. In other words, suppose the user performs an operation so that only ROI video P14 is displayed.

[0198] Then, the control server 51 changes the rendering server 61 assigned to the client 31. For example, if the ROI video P14 is to be displayed at high resolution and a medium frame rate, the control server 51 assigns the rendering server 61 that has high processing power and is located far from the client 31 in the network topology to the client 31. In other words, the control server 51 changes the rendering server 61 responsible for generating the ROI video P14 to the rendering server 61 that has high processing power and is located far from the client 31 in the network topology.

[0199] <Explanation of Rendering Server Determination Process> The operation of the information processing system 11 will be explained.

[0200] For example, when a user operates client 31 and instructs it to start viewing video content, the generation unit 213 generates a connection request (initial connection request) in response to the signal supplied from input unit 201 via user input acquisition unit 212 and supplies it to transmission unit 214. The transmission unit 214 transmits the connection request supplied from generation unit 213 to control server 51 via edge server 41 or the like. This connection request is a request signal that requests a connection to rendering server 61. In other words, the connection request is a request signal that requests the rendering of ROI video (ROI) to be viewed as video content on client 31 when viewing video content. The connection request may include ID information indicating client 31, etc.

[0201] The receiving unit 171 of the edge server 41 receives the connection request sent from the client 31 and supplies it to the control unit 172. The control unit 172 supplies the connection request supplied from the receiving unit 171 to the transmitting unit 173, which then transmits it to the control server 51 via the rendering server 61.

[0202] When a connection request is sent to the control server 51 in this manner, the control server 51 performs the rendering server determination process shown in Figure 15. The rendering server determination process by the control server 51 will now be explained with reference to the flowchart in Figure 15.

[0203] In step S11, the receiving unit 101 receives a connection request (initial connection request) transmitted from the client 31 via the edge server 41, etc., and supplies it to the control unit 103.

[0204] In step S12, the control unit 103 acquires viewing mode information from the client 31 in response to the connection request. Specifically, the control unit 103 generates a transmission request to request viewing mode information from the client 31 and supplies it to the transmission unit 102. The transmission unit 102 transmits the transmission request supplied by the control unit 103 to the client 31 via the rendering server 61 or edge server 41 as appropriate.

[0205] For example, when a transmission request is transmitted via the edge server 41, the receiving unit 174 of the edge server 41 receives the transmission request sent from the control server 51 via the rendering server 61 as appropriate and supplies it to the control unit 172. The control unit 172 supplies the transmission request supplied by the receiving unit 174 to the transmitting unit 176, which then transmits it to the client 31.

[0206] In client 31, when the receiving unit 215 receives a transmission request, the generation unit 213 generates viewing mode information according to the control unit 211 and supplies it to the transmission unit 214. The transmission unit 214 transmits the viewing mode information supplied from the generation unit 213 to the control server 51 via the edge server 41 or the like. For example, the viewing mode information is transmitted to the control server 51 via the same path as the connection request. The receiving unit 101 of the control server 51 receives the viewing mode information transmitted from client 31 and supplies it to the control unit 103.

[0207] In step S13, the control unit 103 performs an assignment process based on viewing format information, i.e., the ROI specified by the client 31, and determines the rendering server 61 to be assigned to the client 31 that sent the connection request. The details of the assignment process will be described later, but in the assignment process, the rendering server 61 to be assigned to the client 31 is determined for each ROI (ROI video) that the client 31 is viewing. Once the assignment process is completed, the rendering server determination process ends.

[0208] As described above, the control server 51 acquires viewing format information in response to a connection request from the client 31 and determines which rendering server 61 to assign to the client 31.

[0209] <Explanation of Assignment Process> Referring to the flowchart in Figure 16, the assignment process corresponding to step S13 in Figure 15 will be explained. Note that, in the absence of dynamic reassignment by the rendering server 61, the assignment process to clients 31 is basically performed in the order in which connection requests were sent.

[0210] In step S41, the control unit 103 performs a 6DoF assignment process based on the viewing mode information and determines the rendering server 61 responsible for generating 6DoF video as ROI video, or more specifically, 6DoF video or 3DoF video. Details of the 6DoF assignment process will be described later.

[0211] In step S42, the control unit 103 performs a 0DoF assignment process based on the viewing mode information and determines the rendering server 61 responsible for generating the 0DoF video as ROI video. Details of the 0DoF assignment process will be described later. Once the 0DoF assignment process is completed, the assignment process ends.

[0212] As described above, the control server 51 determines the allocation of rendering servers 61 according to the viewing style of the video content on the client 31, based on the viewing style information. In this way, it is possible to efficiently utilize limited resources (rendering servers 61) and suppress operating costs while maintaining user satisfaction with the viewing experience.

[0213] <Explanation of 6DoF allocation process> Referring to the flowchart in Figure 17, the 6DoF allocation process corresponding to step S41 in Figure 16 will be explained.

[0214] In step S71, the control unit 103 determines, based on the viewing mode information acquired in step S12 of Figure 15, whether or not the ROI selected by the client 31 that sent the connection request, i.e., the ROI video that the client 31 is trying to view, includes 6DoF video. More specifically, in the explanation of the flowchart in Figure 17, when referring to 6DoF video, it is assumed that 3DoF video is also included.

[0215] For example, the viewing mode information includes viewing ROI information that indicates the ROI to be viewed, and display configuration information that indicates the display configuration of the ROI video, such as single-viewpoint viewing or multi-viewpoint viewing. Therefore, the control unit 103 can identify what degree of freedom of ROI video the client 31 is trying to view and in what display configuration.

[0216] If it is determined in step S71 that no 6DoF video is included, the 6DoF allocation process ends, and the process then proceeds to step S42 in Figure 16.

[0217] In response to this, if it is determined in step S71 that 6DoF video is included, in step S72 the control unit 103 calculates the M2P latency between the rendering server 61 and the client 31 for each rendering server 61 that is currently available.

[0218] For example, as explained with reference to Figure 8, the control unit 103 calculates M2P latency based on the data transmission time between each device, such as between the client 31 and the edge server 41, and the processing time of the 3D rendering process in the rendering server 61. In other words, M2P latency is calculated based on the processing speed of the rendering server 61 and the distance (data transmission delay) between each device in the network topology. Note that if the value of the transmission time between each device is a definitive delay value or a fixed delay value, such as the APN of IOWN (registered trademark), the transmission time may be obtained from a pre-prepared database, rather than measuring the transmission time each time the M2P latency is calculated.

[0219] In step S73, the control unit 103 selects a rendering server 61 from among the available rendering servers 61 in which the M2P latency between it and the client 31 satisfies a predetermined delay time constraint. In step S73, if there are multiple rendering servers 61 that satisfy the delay time constraint, all of those rendering servers 61 are selected.

[0220] The latency constraint referred to here is a constraint defined by the upper limit of the latency during which a user can comfortably perform operations such as changing the ROI when viewing 6DoF video on client 31, i.e., the acceptable latency (tolerant latency).

[0221] For example, while a delay of a few milliseconds to tens of milliseconds is considered the upper limit for comfortable viewing of 6DoF video by a user, the acceptable delay time varies greatly depending on the application and the viewing method of the client 31. Here, we assume that the acceptable delay time (upper limit of delay time) is predetermined as a design value for the application.

[0222] In step S73, if the M2P latency is less than or equal to the allowable delay time, it is determined that the delay time constraint is satisfied. That is, a low-latency rendering server 61 in which the M2P latency is less than or equal to the allowable delay time is selected. Alternatively, as described above, a rendering server 61 may be selected in which the delay time constraint for unidirectional transmission is satisfied, and the delay time constraint for M2P latency is also satisfied.

[0223] In step S74, the control unit 103 selects the rendering server 61 with the lowest usage cost from among the rendering servers 61 selected in step S73. That is, the control unit 103 assigns the client 31 the rendering server 61 whose M2P latency satisfies the latency constraint and has the lowest usage cost, and assigns that rendering server 61 to generate 6DoF video.

[0224] When the control unit 103 selects (determines) a rendering server 61 to be assigned to the client 31, it starts up that rendering server 61. Specifically, the control unit 103 generates a connection instruction to connect to the client 31 and supplies this connection instruction to the transmission unit 102. The transmission unit 102 transmits the connection instruction supplied by the control unit 103 to the rendering server 61. The control unit 103 also generates a connection instruction to connect to the assigned rendering server 61 and supplies it to the transmission unit 102. The transmission unit 102 transmits the connection instruction supplied by the control unit 103 to the client 31.

[0225] As a result, the client 31 and the rendering server 61 are connected, and ROI information is supplied from the client 31 to the rendering server 61, and 6DoF video (ROI video) is supplied from the rendering server 61 to the client 31.

[0226] Since 6DoF video is different for each client 31, unlike 0DoF video, it is not possible to assign a rendering server 61 that is generating 6DoF video for other clients 31 to a single client 31. In other words, multiple clients 31 cannot share the same rendering server 61. Therefore, the control server 51 needs to start the rendering server 61 each time it decides to assign a rendering server 61 to generate 6DoF video.

[0227] Furthermore, the control unit 103 supplies 3D scene data and rendering parameters to the activated rendering server 61 at an appropriate timing.

[0228] Specifically, when the control unit 103 receives scene control information from the performer terminal device 32 via the receiving unit 101, it supplies the scene control information received from the receiving unit 101 to the scene data generation unit 105 and instructs it to generate 3D scene data. The scene data generation unit 105 generates new 3D scene data for a new time by performing scene construction processing based on the scene control information received from the control unit 103 and the past time 3D scene data held in the scene data storage unit 104, and supplies it to the transmission unit 106. The transmission unit 106 transmits the 3D scene data supplied from the scene data generation unit 105 to the rendering server 61.

[0229] Furthermore, the control unit 103 supplies the scene data generation unit 105 with information on the viewing mode of the client 31, more specifically, information that can identify the ROI to be generated, and information indicating the frame rate and resolution of the 6DoF video (ROI video) to be generated, and instructs it to generate rendering parameters. The scene data generation unit 105 generates rendering parameters based on the information supplied from the control unit 103 and supplies them to the transmission unit 106. The transmission unit 106 transmits the rendering parameters supplied from the scene data generation unit 105 to the rendering server 61.

[0230] Furthermore, the 3D scene data and rendering parameters may be transmitted simultaneously, or they may be transmitted at different times. Also, when generating 6DoF video (ROI video), rendering parameters may not be transmitted from the control server 51 to the rendering server 61. Instead, the rendering server 61 may obtain ROI information, frame rate, and resolution information from the client 31 and generate the rendering parameters.

[0231] Once step S74 is performed and the rendering server 61 to be assigned to client 31 is determined, and that rendering server 61 is started, the 6DoF assignment process ends. After the 6DoF assignment process ends, the process proceeds to step S42 in Figure 16.

[0232] As described above, the control server 51 determines the rendering server 61 responsible for generating 6DoF video based on M2P latency and starts up that rendering server 61. In this way, by using M2P latency when determining the rendering server 61 to generate 6DoF video, it is possible to assign an appropriate rendering server 61 to the client 31 and provide 6DoF video with sufficiently low latency. In other words, it is possible to provide users with a highly satisfying viewing experience. Moreover, by assigning the rendering server 61 with the lowest usage cost among those that satisfy the latency constraint to the client 31, operational costs can be suppressed.

[0233] <Explanation of 0DoF allocation process> Referring to the flowchart in Figure 18, the 0DoF allocation process corresponding to the process in step S42 of Figure 16 will be explained.

[0234] In step S101, the control unit 103 determines whether there are any unassigned ROIs for the rendering server 61. For example, if there are ROIs (ROI images) to be viewed by the client 31, identified by the viewing format information acquired in step S12 of Figure 15, for which the rendering server 61 has not yet been assigned to the client 31, then it is determined that there are unassigned ROIs.

[0235] If it is determined in step S101 that there are no unassigned ROIs, the rendering server 61 to be assigned to client 31 has been determined for all ROIs (ROI images), and the 0DoF assignment process ends. Once the 0DoF assignment process is complete, the process in step S42 of Figure 16 has been performed.

[0236] In contrast, if it is determined in step S101 that there are unassigned ROIs, one unassigned ROI is selected as the ROI to be processed (ROI video), and the process then proceeds to step S102. At this time, the ROI to be processed is the ROI of 0DoF video.

[0237] In step S102, the control unit 103 determines whether the ROI (ROI video) to be processed is an ROI that has already been viewed by another client 31.

[0238] In step S102, if it is determined that the ROI is being viewed by another client 31, the process in step S103 is performed. In step S103, the control unit 103 determines whether the rendering server 61 assigned to the other client 31 satisfies the delay time constraint with respect to the ROI (ROI video) to be processed.

[0239] Specifically, the control unit 103 calculates the one-way transmission delay time between the rendering server 61 assigned to another client 31 and the client 31 to be processed for the ROI (ROI video) to be processed. For example, as explained with reference to Figure 8, the control unit 103 calculates the transmission delay time based on the data transmission time between each device, such as between the client 31 and the edge server 41, and the processing time in the control server 51 and the rendering server 61. The control unit 103 then determines that the delay time constraint is satisfied if the calculated transmission delay time is less than or equal to a predetermined allowable delay time.

[0240] If it is determined in step S103 that the delay time constraint is not met, the process then proceeds to step S104.

[0241] Furthermore, if it is determined in step S102 that the ROI is not being viewed by another client 31, the process then proceeds to step S104.

[0242] In step S104, the control unit 103 selects a rendering server 61 from among the available rendering servers 61 in which the one-way transmission delay time with respect to a predetermined delay time constraint satisfies the constraint. Specifically, a rendering server 61 is selected in which the one-way transmission delay time is less than or equal to the allowable delay time.

[0243] In this case, if there are multiple rendering servers 61 that satisfy the delay time constraint, the control unit 103 selects the rendering server 61 that satisfies the delay time constraint and has the lowest usage cost. Furthermore, if, for example, the ROI (ROI video) to be processed is viewed by multiple clients 31, including the client 31 to be processed, the control unit 103 may select a rendering server 61 that satisfies the delay time constraint for all or some of those clients 31.

[0244] The control unit 103 assigns the selected rendering server 61 to the client 31 to be processed for the ROI (ROI video) to be processed, and connects the client 31 to that rendering server 61.

[0245] For example, the control unit 103 supplies a connection instruction to the transmission unit 102 in the same manner as in step S74 of Figure 17, causing the transmission unit 102 to send the connection instruction to the rendering server 61 and the client 31. That is, the transmission unit 102 sends a connection instruction to the rendering server 61 indicating that it will connect to the client 31, and also sends a connection instruction to the client 31 indicating that it will connect to the rendering server 61.

[0246] Furthermore, the control unit 103 sends 3D scene data and rendering parameters to the rendering server 61 at an appropriate timing, similar to the case in step S74 of Figure 17. In step S104, a new rendering server 61 is started up, which is responsible for generating the ROI (ROI image) to be processed.

[0247] Once the processing in step S104 is performed, the process proceeds to step S106. Note that, for example, if it is determined in step S103 that the delay time constraint is not met and the processing in step S104 is performed, the allocation of the rendering server 61 to other clients 31 viewing the ROI (ROI video) to be processed remains unchanged in step S104. In this case, for example, in the dynamic reallocation process described later with reference to Figure 19, the allocation to other clients 31 will be changed to the rendering server 61, etc., that was assigned to the client 31 to be processed in step S104.

[0248] In 6DoF video, the viewpoint position and head direction, i.e., the shooting position and direction of the virtual camera 71, differ for each client 31, making it impossible to share a single rendering server 61 among multiple clients 31. In contrast, with 0DoF video, a single rendering server 61 can be shared among multiple clients 31 if the latency constraint conditions are met. That is, 0DoF video generated by a single rendering server 61 can be distributed and viewed by multiple clients 31.

[0249] If it is determined in step S103 that the delay time constraint is satisfied, the process in step S105 is then performed.

[0250] In step S105, the control unit 103 assigns the rendering server 61 responsible for the ROI (ROI video) to be processed to the client 31 as well. That is, the control unit 103 assigns the rendering server 61 that generates the ROI video to be processed to the client 31, as well as to the other clients 31, and connects the client 31 to that rendering server 61.

[0251] For example, the control unit 103 generates a connection instruction to connect to the rendering server 61, similar to the case in step S74 of Figure 17, and supplies it to the transmission unit 102, causing the transmission unit 102 to send the connection instruction to the client 31. At this time, the connection instruction may also be sent to the rendering server 61. When the process in step S105 is performed, the rendering server 61 assigned to the client 31 has already started up and is generating ROI images, so the client 31 will connect to the already started rendering server 61. As a result, one rendering server 61 is shared by multiple clients 31.

[0252] After the processing in step S104 or step S105 is performed, the processing in step S106 is then performed. In step S106, the control unit 103 completes the allocation of the rendering server 61 to the ROI (ROI video) that is to be processed among the ROIs viewed by the client 31. In other words, after step S104 or step S105 is performed, the allocation of the rendering server 61 to the ROIs that were not yet allocated to be processed is completed.

[0253] Once the process in step S106 is completed, the process returns to step S101, and the process described above is repeated. That is, any ROIs that are still unassigned are designated as new ROIs to be processed, and the allocation of rendering servers 61 to those ROIs is determined. Once the allocation of rendering servers 61 to all ROIs has been determined, it is determined in step S101 that there are no unassigned ROIs, and the 0DoF allocation process ends.

[0254] As described above, the control server 51 determines the rendering server 61 responsible for generating 0DoF video based on the one-way transmission delay time, and connects the client 31 to that rendering server 61. In this way, by using the one-way transmission delay time when determining the rendering server 61 responsible for generating 0DoF video, the appropriate rendering server 61 can be assigned to the client 31. This makes it possible to provide users with a highly satisfying viewing experience while suppressing operating costs, similar to the case of 6DoF assignment processing. Moreover, since the same rendering server 61 can be assigned to multiple clients 31, operating costs can be further reduced.

[0255] By the way, in the process described with reference to Figures 16 to 18, for the sake of simplicity, the interactivity of the video content, that is, the reflection of user actions in the three-dimensional virtual space (scene), is not considered. However, it is not limited to this, and the interactivity of the video content may be considered.

[0256] For example, when bidirectional behavior is considered, when client 31 views ROI video with bidirectional behavior, the control unit 103 of the control server 51 determines which rendering server 61 to assign to client 31 based on the C2P latency described with reference to Figure 9. That is, if the C2P latency is less than or equal to a predetermined allowable delay time, it is determined that the delay time constraint is satisfied. In this case as well, if the delay time constraint is satisfied between multiple clients 31 and the same rendering server 61 for the same ROI video, the same rendering server 61 may be assigned to those multiple clients 31.

[0257] Furthermore, when interactive viewing is performed, the control server 51 receives action information from the client 31 via the receiving unit 101. The control unit 103 supplies the scene control information received from the performer terminal device 32, as well as the action information supplied by the receiving unit 101, to the scene data generation unit 105 and instructs it to generate 3D scene data. In this case, the scene data generation unit 105 generates new 3D scene data based on the scene control information and action information supplied by the control unit 103, and 3D scene data from past times. As a result, 3D scene data is obtained in which the actions indicated by the action information are reflected in the three-dimensional virtual space (scene).

[0258] <Explanation of Dynamic Reassignment Process> In each client 31, while viewing a display video containing one or more ROI videos, the user can change the ROI video displayed as the display video, i.e., the ROI to be displayed, by performing operations on the display video. In other words, the user can change the way the display video is viewed. Also, when video content is provided (distributed), there are clients 31 that start viewing and clients 31 that stop viewing midway. That is, the number of clients 31 connected to the cloud computing environment 21 also increases or decreases (changes).

[0259] As described above, when the viewing pattern of client 31 changes, or when the number of connected clients 31 changes, the optimal allocation of rendering servers 61 to each client 31 for each ROI video also changes. Therefore, after the distribution of video content begins, the control server 51 performs a dynamic reassignment process to dynamically reassign rendering servers 61. The dynamic reassignment process performed by the control server 51 will be explained below with reference to the flowchart in Figure 19.

[0260] In step S131, the control unit 103 determines whether or not there is a client 31 among all clients 31 whose viewing mode has changed.

[0261] For example, when client 31 changes its viewing mode, client 31 sends information to the control server 51 that identifies the change in viewing mode, such as a new connection request or viewing mode information. When the control unit 103 receives information from the receiving unit 101 that identifies the change in viewing mode, it determines that there is a client 31 whose viewing mode has changed.

[0262] If it is determined in step S131 that there are no clients 31 whose viewing mode has changed, the process in step S131 is repeated until it is determined that there are clients 31 whose viewing mode has changed.

[0263] In response to this, if it is determined in step S131 that there is a client 31 whose viewing mode has changed, the process in step S132 is then performed. In step S132, the control unit 103 selects one client 31 to be processed from among all clients 31.

[0264] In step S133, the control unit 103 performs the assignment process shown in Figure 16 for the client 31 to be processed. That is, in step S133, the process of assigning the rendering server 61 to the client 31 to be processed is performed again. In other words, the rendering server 61 is reassigned to the client 31 to be processed.

[0265] In this case, for example, in the 0DoF allocation process in Figure 18, which corresponds to step S42 in Figure 16, the allocation is determined considering the sharing of the rendering server 61 with other clients 31. As an example, suppose the control server 51 receives connection requests for rendering the same ROI (ROI video) from multiple clients 31. In other words, suppose it receives connection requests for rendering the same ROI from a client 31 corresponding to a first display device and a client 31 corresponding to a second display device different from the first display device. In such a case, if there is a rendering server 61 that satisfies the delay time constraint for multiple clients 31 (the first display device and the second display device), the control unit 103 allocates that rendering server 61 to the multiple clients 31. By dynamically reallocating the rendering server 61 in response to changes in viewing patterns in this way, the utilization efficiency of the rendering server group 61 can be improved and operating costs can be reduced.

[0266] In step S134, the control unit 103 determines whether or not processing has been performed for all clients 31. That is, the control unit 103 determines whether or not the processing in step S133 has been performed with all clients 31 as processing targets.

[0267] If it is determined in step S134 that processing has not yet been performed for all clients 31, the process then returns to step S132, and the above-described process is repeated. In this case, the clients 31 that have not yet been targeted for processing are designated as new clients 31 to be processed, and the process in step S133 is performed.

[0268] In contrast, if it is determined in step S134 that processing has been performed for all clients 31, the dynamic allocation process ends.

[0269] As described above, the control server 51 reassigns rendering servers 61 to all clients 31 in response to changes in the viewing patterns of the clients 31. By doing so, it is possible to efficiently utilize the rendering server group 61 while maintaining user satisfaction with the viewing experience and suppressing operating costs.

[0270] The dynamic allocation process described with reference to Figure 19 is the simplest example. In this example, when there is a change in the viewing status (connection status) of one client 31, the optimal rendering server 61 is reassigned to all clients 31 in the information processing system 11. However, the example shown in Figure 19 is merely one example, and the dynamic allocation process can be performed in any way. For example, by determining the order in which to select the clients 31 to be processed according to the viewing status, the utilization efficiency of the rendering server group 61 can be further improved by allowing more clients 31 to share the rendering server 61.

[0271] Here, we will explain a specific example of how the rendering server 61 assignment is changed by the process in step S133 of Figure 19.

[0272] For example, suppose a predetermined client 31-1 corresponding to the first display device was viewing only 0DoF video, which has a lower processing load during generation than 6DoF video, as ROI video. Also, suppose that in the rendering server determination process shown in Figure 15, which was executed on client 31-1 before the processing in step S133, rendering server 61-2 was assigned to client 31-1 as the rendering server 61 that generates 0DoF video.

[0273] Suppose that, from this state, client 31-1's viewing was switched from 0DoF video to viewing only 6DoF video, which has a higher processing load, or more precisely, that the switch was instructed. In other words, the ROI (viewing mode) of client 31-1 has changed. In this case, the change in ROI corresponds to an increase in the degree of freedom of viewpoint regarding the video content (ROI video).

[0274] Suppose a command is issued to switch ROI images, and client 31-1 is selected as the processing target and the process in step S133 is performed. At this time, the control unit 103 determines that the transmission delay between the rendering server 61-2, which had been responsible for generating the 0DoF images to be viewed by client 31-1 (the first display device), and client 31-1, more specifically the M2P latency, is greater than the allowable delay time, and the delay time constraint is not met. In other words, the delay time constraint is no longer met between the rendering server 61-2 and client 31-1 because the ROI has changed.

[0275] Then, based on the determination result that the M2P latency (delay time) with rendering server 61-2 does not satisfy the delay time constraint, the control unit 103 decides which rendering server 61 to assign to client 31-1 (first display device) from among the multiple other rendering servers 61. For example, instead of rendering server 61-2, the control unit 103 assigns to client 31-1 a rendering server 61-1 that has less (shorter) M2P latency with client 31-1 after the ROI change than rendering server 61-2.

[0276] Rendering server 61-1 is located closer to client 31-1 than rendering server 61-2 in the network topology. Therefore, when rendering server 61-1 is assigned after the ROI change, the latency constraint between it and client 31-1 is met, and client 31-1 can view 6DoF video with sufficiently low latency.

[0277] Furthermore, when the ROI changes to accommodate an increase in the degree of freedom of the viewpoint, such as switching from viewing 0DoF video to viewing only 3DoF video, or from viewing 3DoF video to viewing only 6DoF video, the same changes as the above-mentioned changes to the rendering server 61 allocation can be made.

[0278] Suppose that after the rendering server 61 assigned to client 31-1 is changed to rendering server 61-1, the viewing mode (ROI) on client 31-1 changes further. Here, for example, suppose that the viewing on client 31-1 is switched from 6DoF video to viewing only 0DoF video, which has a lower degree of freedom of viewpoint than 6DoF video. In this case, the change in ROI corresponds to a decrease in the degree of freedom of viewpoint related to the video content (ROI video). Suppose that in response to this change in ROI (viewing mode), the processing in step S133 is performed again for client 31-1.

[0279] For example, when the process in step S133 is performed, the control unit 103 determines whether the change in ROI (viewing mode) corresponds to a decrease in the degrees of freedom of the viewpoint related to the video content, and if it is determined that it is a change corresponding to a decrease in the degrees of freedom, then the degrees of freedom of the viewpoint have decreased because the ROI video being viewed has been changed from a 6DoF video to a 0DoF video.

[0280] Then, based on the determination result that the change in ROI (viewing mode) corresponds to a decrease in degrees of freedom, the control unit 103 decides which rendering server 61 to assign to client 31-1 from among the multiple other rendering servers 61. For example, instead of rendering server 61-1, the control unit 103 assigns to client 31-1 a rendering server 61 that has a larger (longer) transmission delay time with client 31-1 after the change in ROI than rendering server 61-1. In this case, the rendering server 61 assigned to client 31-1 may be rendering server 61-2, or a different rendering server 61 from rendering server 61-2, as long as the transmission delay time with client 31-1 satisfies the delay time constraint.

[0281] <Explanation of the process for determining the allowable delay time> As described above, in the allocation process explained with reference to Figure 16, the allowable delay time is used to determine whether the delay time constraint is met. That is, for example, in step S73 in Figure 17, step S103 in Figure 18, and step S104 in Figure 18, the allowable delay time is used as the threshold for determining whether the delay time constraint is met.

[0282] For example, the control server 51 performs the allowable delay time determination process shown in Figure 20 at an appropriate timing to determine the allowable delay time to be used for the client 31. The allowable delay time determination process by the control server 51 will be explained below with reference to the flowchart in Figure 20.

[0283] In step S201, the control unit 103 obtains viewing mode information from the client 31 that sent the connection request. For example, in step S201, the viewing mode information is obtained by the same process as in step S12 in Figure 15. Alternatively, in step S201, the viewing mode information obtained in step S12 in Figure 15 may be read out.

[0284] In step S202, the control unit 103 determines, based on the viewing format information, whether the client 31 that sent the connection request is requesting to view 6DoF video as ROI video, or more specifically, whether it is requesting to view 6DoF video or 3DoF video. In the explanation of Figure 20, 6DoF video is assumed to include 3DoF video. The number of ROI videos that the client 31 is requesting to view and the number of degrees of freedom of the viewpoint for each ROI video can be determined from the viewing ROI information.

[0285] If it is determined in step S202 that the user is requesting to view 6DoF video, the process then proceeds to step S205.

[0286] If, in step S202, it is determined that the client has not requested to view 6DoF video, then the process in step S203 is performed. In step S203, the control unit 103 determines, based on the viewing format information, whether the client 31 that sent the connection request has requested to view two or more 0DoF videos.

[0287] If it is determined in step S203 that the client has not requested to view two or more 0DoF videos, then the process in step S204 is performed. In step S204, the control unit 103 sets the allowable delay time used to determine whether the delay time constraint for the client 31 is met to a predetermined time A.

[0288] When step S204 is performed, viewing on client 31 is single viewing of 0DoF video, that is, viewing of only one 0DoF video. Therefore, unlike when multiple ROI videos are displayed simultaneously, synchronization accuracy between multiple ROI videos is not required, and a relatively lenient condition, namely a relatively long delay time A, is used to determine whether the delay time constraint is met. For example, the allowable delay time A is used in steps S103 and S104 in Figure 18. Once the allowable delay time is determined after the processing in step S204, the allowable delay time determination process ends.

[0289] Furthermore, if it is determined in step S203 that the user is requesting to view two or more 0DoF videos, the process in step S205 is then performed. Also, as described above, if it is determined in step S202 that the user is requesting to view a 6DoF video, the process in step S205 is also performed.

[0290] In step S205, the control unit 103 sets the allowable delay time used to determine whether the delay time constraint for client 31 is satisfied to B, which is shorter than the allowable delay time A.

[0291] When step S205 is performed, viewing on client 31 involves viewing one or more ROI videos, including at least 6DoF videos, or simultaneously viewing multiple 0DoF videos. When simultaneous viewing of multiple 0DoF videos, or when simultaneous viewing of 6DoF videos and other ROI videos (0DoF videos), it is necessary to minimize the synchronization delay between these videos. Therefore, in step S205, an allowable delay time B, which is a stricter condition than the allowable delay time A, is used to determine whether the delay time constraint is met. For example, the allowable delay time B is used in step S73 in Figure 17, and in steps S103 and S104 in Figure 18. Once the allowable delay time is determined after the processing in step S205, the allowable delay time determination process ends.

[0292] The specific values ​​for allowable delay time A and allowable delay time B will be determined according to the use case and application. Furthermore, the method for determining the allowable delay time explained with reference to Figure 20 is merely an example, and the allowable delay time may be determined in any other way.

[0293] As described above, the control server 51 determines an allowable delay time that defines the delay time constraint based on the viewing style of the client 31. By doing so, the rendering server 61 can be allocated more appropriately using an allowable delay time suitable for the viewing style, and users can comfortably view video content without experiencing delays. In other words, it is possible to suppress operating costs while providing a highly satisfying viewing experience.

[0294] <Explanation of video generation process> When a connection instruction to client 31 is sent from the control server 51 to the rendering server 61, the receiving unit 142 of the rendering server 61 receives the connection instruction. Then, the control unit 141 controls the transmitting unit 145 and the receiving unit 142 to communicate with client 31 and establish a connection with client 31.

[0295] Furthermore, when 3D scene data and rendering parameters are transmitted from the control server 51 to the rendering server 61, the receiving unit 142 receives the 3D scene data and rendering parameters. The receiving unit 142 supplies the received 3D scene data to the scene data storage unit 143 and supplies the received rendering parameters to the image generation unit 144.

[0296] When the rendering server 61 connects to the client 31, it starts the video generation process shown in Figure 21. The video generation process by the rendering server 61 will be explained below with reference to the flowchart in Figure 21.

[0297] In step S241, the video generation unit 144 generates an ROI video based on rendering parameters and the like supplied from the receiving unit 142 and the 3D scene data stored in the scene data storage unit 143, and supplies it to the transmission unit 145.

[0298] For example, the video generation unit 144 reconstructs a scene in a three-dimensional virtual space based on 3D scene data and generates video of a predetermined ROI of the reconstructed scene by performing 3D rendering processing. In this case, video of an ROI with a specified ROI, frame rate, and resolution, as specified by rendering parameters, is generated.

[0299] Specifically, for example, when a 0DoF image is generated as an ROI image, the image generation unit 144 generates the 0DoF image based on rendering parameters. Also, for example, when a 6DoF image or a 3DoF image is generated as an ROI image, the image generation unit 144 generates the 6DoF image or a 3DoF image based on rendering parameters and ROI information received from the client 31 by the receiving unit 142.

[0300] In step S242, the transmission unit 145 transmits the ROI video, or more specifically, the video data of the ROI video, supplied from the video generation unit 144 to the client 31. This video data is transmitted to the client 31 via, for example, an edge server 41. More specifically, if a single ROI video is viewed by the client 31, that ROI video is transmitted to the client 31 as a display image. In contrast, if multiple ROI videos are viewed simultaneously by the client 31 (multi-viewpoint viewing), the edge server 41 generates a display image from the multiple ROI videos, including the ROI video transmitted in step S242, and transmits that display image to the client 31.

[0301] In step S243, the control unit 141 determines whether or not to terminate the process of generating the ROI image.

[0302] If it is determined in step S243 that the process is not yet complete, the process then returns to step S241, and the process described above is repeated.

[0303] In contrast, if it is determined in step S243 that the process should be terminated, the control unit 141 stops the operation of each part of the rendering server 61, and the video generation process is terminated.

[0304] As described above, the rendering server 61 performs 3D rendering processing based on the 3D scene data and rendering parameters, and transmits the resulting ROI video data to the client 31. By performing 3D rendering processing according to the assignments of the control server 51, the rendering server 61 can provide ROI video to the client 31 with low latency. This enables the user to receive a highly satisfying viewing experience.

[0305] <Explanation of Content Playback Process> When the receiving unit 215 receives a connection instruction in response to a connection request in the client 31, the control unit 211 controls the transmitting unit 214 and the receiving unit 215 according to the connection instruction, and establishes a connection with the rendering server 61 by having them communicate with the rendering server 61. Then, the rendering server 61 transmits video data of the display image via the edge server 41.

[0306] After connecting with the rendering server 61, client 31 performs the content playback process shown in Figure 22. The content playback process by client 31 will be explained below with reference to the flowchart in Figure 22.

[0307] In step S281, the receiving unit 215 receives video data of the display image (ROI image) transmitted from the rendering server 61 via the edge server 41 and supplies it to the video display control unit 216.

[0308] In step S282, the video display control unit 216 supplies the video data received from the receiving unit 215 to the display 202, causing the display 202 to display a video containing one or more ROI videos. This plays the video content. More specifically, if there is audio accompanying the ROI video, the receiving unit 215 also receives the audio data. The control unit 211, etc., then supplies the audio data to a speaker (not shown), and the speaker plays the audio of the video content.

[0309] Furthermore, when client 31 displays 6DoF video, or more specifically, video containing 6DoF or 3DoF video, ROI information is sent to rendering server 61 as appropriate, in response to user input.

[0310] Specifically, for example, when a user operates the input unit 201 or moves while the sensor acting as the input unit 201 is attached, an input that changes the ROI in the three-dimensional virtual space is made. A signal indicating this input is acquired by the user input acquisition unit 212 and supplied to the generation unit 213. Based on the signal supplied from the user input acquisition unit 212, the generation unit 213 generates ROI information indicating the changed ROI and supplies it to the transmission unit 214, which then transmits it to the rendering server 61. As a result, the ROI image (6DoF image or 3DoF image) generated by the rendering server 61 reflects the user's movement (change in ROI).

[0311] In addition, for example, if a user switches the ROI video displayed on the displayed video while viewing the displayed video, the control unit 211 controls the transmission unit 214 to send new connection requests, viewing format information, etc., to the control server 51 as appropriate.

[0312] In step S283, the control unit 211 determines whether or not to terminate the process of playing the video content.

[0313] If it is determined in step S283 that the process is not yet complete, the process then returns to step S281, and the process described above is repeated.

[0314] In contrast, if it is determined in step S283 that the process should be terminated, the control unit 211 stops the operation of each part of the client 31, and the content playback process ends.

[0315] In this manner, client 31 receives the video data and displays (plays) the video. This allows the user to view the video content.

[0316] <Other Application Examples> This technology can be applied not only when providing events and other content in a three-dimensional virtual space, but also when providing live events and other content in the real world.

[0317] One example of a use case is using multiple remotely controlled cameras to film events such as live music concerts or sporting events held in the physical space.

[0318] In such cases, for example, a service could be provided that generates a 3D model that mimics the real space in a three-dimensional virtual space from multiple images with different ROIs obtained by photographing the real space, and then provides the client with real-time rendering of images with an arbitrary ROI. In such an example, after generating the 3D model, the same processing as in the first embodiment described above would be performed.

[0319] On the other hand, one possible use case is to provide a client with multiple videos, each with a different ROI, captured by multiple cameras, in their original 2D format.

[0320] In such a case, the configuration of the information processing system can be considered as shown in Figure 23, for example. In the information processing system 301 shown in Figure 23, a music live performance or the like taking place at the live venue LV31, which is a real space, is captured by multiple cameras, including camera 311-1 and camera 311-2. Hereafter, when there is no need to distinguish between cameras such as camera 311-1 and camera 311-2, they will simply be referred to as camera 311. For example, camera 311 may be a mobile PTZ camera that can be operated via a network.

[0321] The live venue LV31 is equipped with multiple cameras 311, as well as speakers such as speaker 312-1 and speaker 312-2. Each camera 311 and speaker is connected to an edge server on the venue side. Edge servers 313-1 and 313-2 are provided as edge servers on the venue side.

[0322] Hereinafter, when there is no need to distinguish between edge server 313-1 and edge server 313-2, they will simply be referred to as edge server 313, and when there is no need to distinguish between speaker 312-1 and speaker 312-2, they will simply be referred to as speaker 312.

[0323] Each edge server 313 is connected to the cloud computing environment 314. The cloud computing environment 314 is connected to clients 316-1 to 316-4 via client-side edge servers such as edge servers 315-1, 315-2, and 315-3. Furthermore, the cloud computing environment 314 has servers 321-1 to 321-3.

[0324] Hereafter, when there is no need to distinguish between edge servers 315-1 to 315-3, they will simply be referred to as edge server 315, and when there is no need to distinguish between clients 316-1 to 316-4, they will simply be referred to as client 316. Also, hereafter, when there is no need to distinguish between servers 321-1 to 321-3, they will simply be referred to as server 321.

[0325] For example, the edge server 313 can operate the camera 311 to change the ROI, supply ROI images captured by the camera 311 to the server 321, and output sound from the speaker 312. For example, the camera 311 can be operated to focus, control the pan / tilt head, and seamlessly switch between tracking targets. The edge server 313 can also perform image processing for tracking a specific person using the camera 311, such as person estimation and person movement estimation.

[0326] Server 321 performs tasks such as issuing control instructions to the edge server 313 for the camera 311, processing video such as effect compositing and picture-in-picture (PiP) compositing for ROI video, and mixing audio related to ROI video as needed. However, server 321 does not perform 3D model generation or 3D rendering.

[0327] Server 321 processes the ROI video supplied from edge server 313 as appropriate, and supplies the resulting ROI video to client 316 via edge server 315. More specifically, server 321 also generates audio accompanying the ROI video. In this process, for example, mixing may be performed, and the audio of the audience (users) uploaded from client 316 may be combined with the audio from the live venue LV31.

[0328] The edge server 315 distributes ROI video from server 321 to client 316, uploads audio from client 316 to server 321, and controls the viewpoint switching of the ROI video viewed by client 316.

[0329] In an information processing system 301 with this configuration, the processing performed in the cloud computing environment 314 (server 321) does not include 3D rendering. Instead, relatively low-processing-load video processing such as effect compositing and picture-in-picture (PiP) compositing is performed on server 321. Therefore, the selection (assignment) of server 321 to client 316 is primarily done using the network topology-based transmission delay time, which is determined by the geographical location of server 321, as a parameter.

[0330] For example, suppose camera 311 is a network-enabled mobile PTZ camera whose shooting position and direction can be changed by remote operation from client 316, and the ROI video captured by such camera 311 is viewed on client 316. In such a case, by having the video processing and other operations on the ROI video obtained from camera 311 performed on server 321, which is located close to client 316, the video response time can be shortened and the user's viewing experience can be improved.

[0331] On the other hand, instead of the client 316 operating the camera 311, it may be desirable to implement a camera that automatically and continuously follows a specific group member performing at a music live concert (a so-called "fan camera"). In such cases, the edge server 313 located at the live venue LV31 may perform person recognition within the ROI video, generate control signals to control the movement of the camera 311 to follow the person identified by the person recognition, and transmit the control signals to the camera 311. Performing this processing at the edge server 313 close to the live venue LV31 allows for low-latency control of the camera 311 and enables the capture of highly tracking video of a specific person.

[0332] <Description of a computer to which this technology is applied> The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up the software are installed on the computer. Here, the term "computer" includes computers built into dedicated hardware, as well as general-purpose personal computers, for example, that can perform various functions by installing various programs.

[0333] Figure 24 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above using a program.

[0334] In a computer, the processing circuit 901, ROM (Read Only Memory) 902, and RAM (Random Access Memory) 903 are interconnected by a bus 904.

[0335] An input / output interface 905 is further connected to the bus 904. An input / output interface 905 is connected to an input unit 906, an output unit 907, a recording unit 908, a communication unit 909, and a drive 910.

[0336] The input unit 906 may include physical or virtual means of operation that the user operates to input information, such as a keyboard, mouse, or touch panel, as well as means of inputting information by the user through voice, eye gaze, etc. Furthermore, the input unit 906 may include sensors for inputting various physical quantities to the computer.

[0337] For example, the input unit 906 may include sensors that acquire physical quantities such as light (including infrared light other than visible light) and sound, such as cameras and microphones. Alternatively, the input unit 906 may include sensors that acquire other physical quantities such as temperature, moisture content, acceleration, and distance.

[0338] The output unit 907 may include means for presenting information to the user by stimulating the user's senses, such as a display, speaker, or haptic device. The recording unit 908 consists of a hard disk, non-volatile or volatile memory, etc., and records various types of information (including programs).

[0339] The communication unit 909 is a network interface, etc., and performs wired or wireless communication with the outside. The drive 910 drives removable media 911 such as magnetic disks, optical disks, magneto-optical disks, or semiconductor memory.

[0340] The processing circuit 901 includes a processor that executes programs such as a CPU (Central Processing Unit) and a DSP (Digital Signal Processor). The processing circuit 901 (its processor) loads the program recorded in the recording unit 908 into the RAM 903 via the input / output interface 905 and the bus 904, and executes it, thereby performing the series of processes described above.

[0341] The processing circuit 901 can output the processing results of a series of processes from the output unit 907, for example, via the bus 904 and the input / output interface 905, as needed. The processing circuit 901 can also record the processing results in the recording unit 908 or transmit them from the communication unit 909.

[0342] The program executed by the computer (processing circuit 901) can be provided by recording it on a removable medium 911, such as a package medium. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.

[0343] In a computer, a program can be installed in the recording unit 908 via the input / output interface 905 by inserting the removable media 911 into the drive 910. Alternatively, a program can be received by the communication unit 909 from another device, such as a server, via a wired or wireless transmission medium, and installed in the recording unit 908. Furthermore, programs can be pre-installed in the ROM 902 or the recording unit 908.

[0344] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.

[0345] The processes that a computer performs according to a program do not necessarily have to follow the order described in the flowchart. In other words, the processes that a computer performs according to a program include processes that are executed in parallel or individually (e.g., parallel processing and object-based processing).

[0346] The program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Furthermore, the program may be transferred to a remote computer and executed there.

[0347] When the computer executes a program to perform the series of processes described above, for example, the communication unit 909 functions as the receiving unit 101, transmitting unit 102, and transmitting unit 106 in Figure 2, and the recording unit 908 functions as the scene data storage unit 14. Also, for example, the processing circuit 901 (its processor) functions as the control unit 103 and the scene data generation unit 105 by executing a program.

[0348] In this specification, a system means one component or a collection of multiple components (devices, modules (parts), etc.). Therefore, one or more components of a computer, for example, only the processor, or a combination of the processor and memory (for example, only the processing circuit 901, or a combination of the processing circuit 901 to the bus 904, etc.), constitute a system. Regarding a collection of multiple components, it is not necessary whether all components reside in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, or a single device containing multiple modules within a single enclosure, are all systems. Furthermore, for example, the entire computer, or a combination of a computer and other devices such as a server (not shown), also constitute a system.

[0349] The components (blocks) of the apparatus illustrated in this specification are functional conceptual blocks, and the actual apparatus does not need to have the illustrated configuration. That is, the apparatus can have any configuration in which the functions of the illustrated components are divided into any units and / or integrated, for example, a configuration having one block in which the functions of all components are integrated.

[0350] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.

[0351] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0352] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.

[0353] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.

[0354] Furthermore, this technology can also be configured as follows:

[0355] (1) A cloud rendering system comprising: a plurality of rendering devices each configured to render a predetermined video content; a control server configured to receive a request signal from an edge server for rendering a region of interest of a first display device for the predetermined video content; assign a first rendering device from the plurality of rendering devices to the first display device based on the region of interest; determine whether a first transmission delay of the first rendering device in response to a change in the region of interest satisfies a delay time constraint for the first display device; and, based on the determination that the first transmission delay does not satisfy the delay time constraint, assign a second rendering device having a second transmission delay smaller than the first transmission delay in response to a change in the region of interest to the first display device from the plurality of rendering devices, instead of the first rendering device. (2) The cloud rendering system according to (1), wherein the control server determines whether the delay time constraint for the first display device is satisfied based on at least one of the transmission delay time from the first rendering device to the first display device, or the responsiveness of the first display device to user input to the first display device. (3) The cloud rendering system according to (2), wherein the user input includes at least one of a gesture operation, a button input operation, or a touch input operation on the first display device. (4) The cloud rendering system according to (2) or (3), wherein the control server determines whether the delay time set for the viewing mode of the predetermined video content on the first display device satisfies the delay time constraint, and determines the rendering device to be assigned to the first display device based on the determination result. (5) The cloud rendering system according to (4), wherein when 0DoF video is viewed on the first display device, the control server determines whether the data transmission delay time to the first display device satisfies the delay time constraint.(6) The cloud rendering system according to (4), wherein the control server determines whether the M2P latency satisfies the delay time constraint when 6DoF video or 3DoF video is viewed on the first display device. (7) The cloud rendering system according to (4), wherein the control server determines whether the C2P latency satisfies the delay time constraint when video having bidirectional interaction with the control server is viewed on the first display device. (8) The cloud rendering system according to any one of (4) to (7), wherein the control server determines an allowable delay time used to determine whether the delay time constraint is satisfied based on the viewing mode of the predetermined video content on the first display device. (9) The cloud rendering system according to (8), wherein the control server sets the allowable delay time as a first allowable delay time when one 0DoF video is viewed on the first display device, and sets the allowable delay time as a second allowable delay time shorter than the first allowable delay time when 6DoF video or 3DoF video is viewed on the first display device, or when multiple videos are viewed. (10) A cloud rendering system according to any one of (1) to (9), wherein the change in the region of interest corresponds to an increase in the degree of freedom of the viewpoint with respect to the predetermined video content. (11) A cloud rendering system according to (10), wherein the increase in the degree of freedom of the viewpoint with respect to the predetermined video content corresponds to at least one of a change from 0DoF video to 3DoF video or 6DoF video, or a change from 3DoF video to 6DoF video. (12) A cloud rendering system according to (10) or (11), wherein the second rendering device is closer to the first display device than the first rendering device in the network topology.(13) The control server further determines whether the change in the region of interest corresponds to a decrease in the degree of freedom of the viewpoint with respect to the predetermined video content, and based on the determination result that the change in the region of interest corresponds to a decrease in the degree of freedom, assigns to the first display device a third rendering device from among the plurality of rendering devices, which has a third transmission delay that is greater than the second transmission delay with respect to the change in the region of interest, instead of the second rendering device. The cloud rendering system according to any one of (1) to (12). (14) The predetermined video content is provided as video captured by a virtual camera in a three-dimensional virtual space as the region of interest. The cloud rendering system according to any one of (1) to (13). (15) When the control server receives a request signal for rendering the same region of interest as the first display device as the request signal for a second display device different from the first display device, it assigns a fourth rendering device from a plurality of rendering devices to the first display device and the second display device, the fourth rendering device satisfying the delay time constraint for the first display device and the second display device. The cloud rendering system according to any one of (1) to (14).(16) A cloud rendering method for a cloud rendering system having a plurality of rendering devices, each configured to render a predetermined video content, and a control server, the method comprising: the control server receiving a request signal from an edge server for rendering a region of interest of a display device for the predetermined video content; assigning a first rendering device from the plurality of rendering devices to the display device based on the region of interest; determining whether a first transmission delay of the first rendering device in response to a change in the region of interest satisfies a delay time constraint for the display device; and, based on the determination result that the first transmission delay does not satisfy the delay time constraint, assigning a second rendering device from the plurality of rendering devices to the display device, in place of the first rendering device, the second rendering device having a second transmission delay smaller than the first transmission delay in response to a change in the region of interest.

[0356] 11 Information processing system, 21 Cloud computing environment, 31-1 to 31-3, 31 Client, 32 Performer terminal device, 41-1 to 41-4, 41 Edge server, 51 Control server, 61-1 to 61-4, 61 Rendering server, 101 Receiving unit, 102 Transmitting unit, 103 Control unit, 104 Scene data storage unit, 105 Scene data generation unit, 106 Transmitting unit

Claims

1. A cloud rendering system comprising: a plurality of rendering devices, each configured to render a predetermined video content; a control server configured to receive a request signal from an edge server regarding the rendering of a region of interest of a first display device for the predetermined video content; assign a first rendering device from the plurality of rendering devices to the first display device based on the region of interest; determine whether a first transmission delay of the first rendering device in response to changes in the region of interest satisfies a delay time constraint for the first display device; and, based on the determination that the first transmission delay does not satisfy the delay time constraint, assign a second rendering device having a second transmission delay smaller than the first transmission delay in response to changes in the region of interest to the first display device from the plurality of rendering devices, instead of the first rendering device.

2. The cloud rendering system according to claim 1, wherein the control server determines whether the delay time constraint for the first display device is satisfied based on at least one of the transmission delay time from the first rendering device to the first display device, or the responsiveness of the first display device to user input to the first display device.

3. The cloud rendering system according to claim 2, wherein the user input includes at least one of a gesture operation, a button input operation, or a touch input operation to the first display device.

4. The cloud rendering system according to claim 2, wherein the control server determines whether the delay time defined for the viewing mode of the predetermined video content on the first display device satisfies the delay time constraint, and determines the rendering device to be assigned to the first display device based on the determination result.

5. The cloud rendering system according to claim 4, wherein the control server determines whether the data transmission delay time to the first display device satisfies the delay time constraint when 0DoF video is viewed on the first display device.

6. The cloud rendering system according to claim 4, wherein the control server determines whether the M2P latency satisfies the delay time constraint when 6DoF video or 3DoF video is viewed on the first display device.

7. The cloud rendering system according to claim 4, wherein the control server determines whether the C2P latency satisfies the delay time constraint when bidirectional video is viewed on the first display device with respect to the control server.

8. The cloud rendering system according to claim 4, wherein the control server determines an allowable delay time used to determine whether the delay time constraint is satisfied, based on the viewing mode of the predetermined video content on the first display device.

9. The cloud rendering system according to claim 8, wherein the control server sets the allowable delay time to a first allowable delay time when one 0DoF video is viewed on the first display device, and sets the allowable delay time to a second allowable delay time that is shorter than the first allowable delay time when 6DoF video or 3DoF video is viewed on the first display device, or when multiple videos are viewed.

10. The cloud rendering system according to claim 1, wherein the change in the region of interest corresponds to an increase in the degree of freedom of the viewpoint with respect to the predetermined video content.

11. The cloud rendering system according to claim 10, wherein the increase in the degree of freedom of viewpoint with respect to the predetermined video content corresponds to at least one of a change from 0DoF video to 3DoF video or 6DoF video, or a change from 3DoF video to 6DoF video.

12. The cloud rendering system according to claim 10, wherein the second rendering device is closer to the first display device than the first rendering device in the network topology.

13. The cloud rendering system according to claim 1, wherein the control server further determines whether the change in the region of interest corresponds to a decrease in the degree of freedom of the viewpoint with respect to the predetermined video content, and based on the determination result that the change in the region of interest corresponds to a decrease in the degree of freedom, assigns to the first display device a third rendering device from among the plurality of rendering devices, which has a third transmission delay that is greater than the second transmission delay with respect to the change in the region of interest, instead of the second rendering device.

14. The cloud rendering system according to claim 1, wherein the predetermined video content is video captured by a virtual camera in a three-dimensional virtual space, which is provided as the region of interest.

15. The cloud rendering system according to claim 1, in which, when the control server receives a request signal for rendering the same region of interest as the first display device as the request signal for a second display device different from the first display device, it assigns a fourth rendering device from a plurality of rendering devices to the first display device and the second display device, the rendering device that satisfies the delay time constraint for the first display device and the second display device.

16. A cloud rendering method for a cloud rendering system having a plurality of rendering devices, each configured to render a predetermined video content, and a control server, the method comprising: the control server receiving a request signal from an edge server regarding the rendering of a region of interest of a display device for the predetermined video content; assigning a first rendering device from the plurality of rendering devices to the display device based on the region of interest; determining whether a first transmission delay of the first rendering device in response to a change in the region of interest satisfies a delay time constraint for the display device; and, based on the determination result that the first transmission delay does not satisfy the delay time constraint, assigning a second rendering device having a second transmission delay smaller than the first transmission delay in response to a change in the region of interest to the display device from the plurality of rendering devices, instead of the first rendering device.