Title - METHOD FOR TRANSITIONING VIDEO TRACKS IN MULTI-TRACK IMMERSIVE VIDEOS, RELATED CLIENT DEVICE AND SERVER DEVICE

AR127086B1Active Publication Date: 2026-08-26KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ARP20210103377
Authority / Receiving Office
AR · AR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-11
Filing Date
2021-12-06
Publication Date
2026-08-26
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

Immersive videos with large fields of view face challenges in transitioning between video tracks, leading to sudden changes that can cause visible objects and disorientation due to resource constraints and viewport changes.

Method used

A method for transitioning between video tracks by applying a decreasing or increasing rendering priority function over time, using a decaying function to gradually reduce or increase the priority of video tracks, along with strategies like reducing resolution, frame rate, and bit rate to manage resource demands.

Benefits of technology

Reduces sudden changes in immersive videos, enhancing the viewing experience by minimizing visible objects and disorientation during transitions, while optimizing resource usage.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method for transitioning from a first set of video tracks, VT1, to a second set of video tracks, VT2, when rendering a multi-track video, wherein each video track has a corresponding rendering priority. The method comprises receiving an instruction to transition from a first set of first video tracks, VT1, to a second set of second video tracks, VT2, obtaining the video tracks VT2, and, if the video tracks VT2 are different from the video tracks VT1, applying a decreasing function to the rendering priority of one or more of the video tracks in the first set of video tracks, VT1, and / or an increasing function to the rendering priority of one or more video tracks in the second set of video tracks, VT2. The decreasing and increasing functions decrease and increase the rendering priority over time, respectively.The rendering priority is used in determining the weighting of a video track and / or elements of a video track used to render a multi-track video.
Need to check novelty before this filing date? Find Prior Art

Description

1 / 26 CHANGING VIDEO TRACKS IN IMMERSIVE VIDEOS FIELD OF INVENTION The invention relates to the field of multi-track immersive videos. In particular, the invention relates to the transition between sets of video tracks in multi-track immersive videos. BACKGROUND OF THE INVENTION Immersive video, also known as 3DoF+ or 6DoF video (where DoF stands for "degrees of freedom"), allows a viewer within a viewing space to have a large field of view (FoV). In real-world scenes, the viewing space and FoV are determined by the camera system used to capture the video data. Computer graphics offer more flexibility, allowing for the direct representation of, for example, a 360-degree equirectangular projection. Encoding and streaming immersive video is challenging due to the resources required. There are also practical limitations on the total bitrate, the number of decoders, and the total frame size that a client can decode and render. Similarly, there are practical limits to the amount of video that can be uploaded from a live event, via satellite or 5G, and between cloud services, to maintain cost-effective streaming. It is common for modern video streaming services to adapt the transmission to the client's capabilities. Standards such as Dynamic Adaptive Streaming over HTTP MPEG (MPEG-DASH) 1603437 of 27 2 / 26 ISO / IEC 23009-1, the Common Media Application Format (CMAF) ISO / IEC 23000-19, and Apple's HTTP Live Streaming (HLS) specify how to divide video into smaller segments. Each segment is available in multiple versions, for example, at different bitrates. The client can measure available bandwidth and CPU usage and decide to request lower or higher quality segments from the server. The MPEG ISO / IEC 23090-2 Omnidirectional Media Format (OMAF) standard specifies how to transmit VR360 (3DoF) video, with profiles for both DASH and CMAF. Note that VR360 is also referred to as "3DoF," and the name "3DoF+" is derived from this nomenclature by adding some parallax to VR360. Both viewport-independent and viewport-dependent modes are supported. In the latter, the video frame is divided into tiles, and each tile is a separate video track. This allows the total video frame to exceed the capabilities of common hardware video decoders. For example, HEVC's main 5.2 level 10 is limited to 4K x 2K frames, which is below the resolution of a virtual retina display when used for a 360-degree field of view. For 360 VR, a resolution of 8K x 4K or even 16K x 8K is preferred. A new standard (ISO / IEC 23090-12) for MPEG Immersive Video (MIV) describes how to crop and package multiview+background video into multiple texture and background atlases for encoding using 2D video codecs. While this standard itself does not include adaptive streaming, it does include support for accessing sub-bitstreams, and it is expected that system aspects such as adaptive streaming will be addressed in future standards. 1603437 of 27 3 / 26 MPEG (or non-MPEG) related standards such as ISO / IEC 23090-10, Visual Volumetric Video-based Encoding Data Carrier. The idea behind bitstream access is that some units in a bitstream can be removed, resulting in a smaller, yet still valid, bitstream. A common form of bitstream access is temporal access (i.e., reducing the frame rate), but the MIV also supports spatial access, which is implemented by allowing the removal of a subset of the atlases. Each atlas provides a video track, and combinations of video tracks are used to represent the MIV. However, changing video tracks can produce visible objects due to a jump that would be especially noticeable in disocclusion zones and any non-Lambertian scene elements. For example, a change in viewports can cause a change in video tracks, which can produce objects due to the sudden change. BRIEF DESCRIPTION OF THE INVENTION The invention is defined by the claims. According to the examples according to one aspect of the invention, a method is provided for transitioning from a first set of video tracks, VT1, to a second set of video tracks, VT2, when rendering a multi-track video, wherein each video track has a corresponding rendering priority; the method comprises: receive an instruction to transition from VT1 video tracks to VT2 video tracks; obtain the VT2 video tracks; and 1603437 of 27 4 / 26 If the VT2 video tracks are different from the VT1 video tracks, apply: a decreasing function for the rendering priority of one or more of the video tracks in the first set of VT1 video tracks, wherein the decreasing function decreases the rendering priority over time, and / or an increasing function for the rendering priority of one or more of the VT2 video tracks, wherein the increasing function increases the rendering priority over time, wherein the rendering priority of one or more of the video tracks is to be used in selecting a weighting of the respective video track and / or video track elements when rendering multitrack video. The method may also include the stage of selecting the weight and the stage of representation using the weight or weights. Immersive videos with large fields of view cannot typically be adjusted to fit on a single screen for a user to view; therefore, a viewport of the video (i.e., a section of the entire video) is displayed to the user. This viewport can then change based on user activity (e.g., sensors detecting user movement or the user moving a joystick / mouse) or other factors (e.g., the server). As described earlier, immersive videos can be based on multiple video tracks used to create the immersive video itself. However, when you change the viewport, you can trigger a change in the video tracks used to render the viewport. This change can produce objects visible in the video. A similar change in video tracks can also occur for other reasons, such as a change in the width of 1603437 of 27 5 / 26 bandwidth when transmitting multi-track video; therefore, the video tracks used to represent the multi-track video are reduced and possibly changed (e.g., to represent the video at a lower resolution; therefore, the bandwidth required to transmit the multi-track video is reduced). This invention is based on improving the transition between video tracks. When an instruction is received to transition from one set of video tracks to another (e.g., a user moves a joystick), the new set of video tracks is rendered. If the new set of video tracks differs from the previously used video tracks, the rendering priority is gradually lowered for some (or all) of the video tracks in the new set. Reduction of rendering priority is achieved by applying a decreasing function to the rendering priority of the video tracks, which gradually decreases a numerical value representative of the rendering priority over time. By gradually reducing the rendering priority during the transition between video tracks, sudden changes in immersive video (i.e., MIV) are reduced, and thus the visible objects produced by sudden changes are also reduced, thereby increasing the viewing experience for a user. Similarly (or alternatively), the rendering priority of new VT2 video tracks can be gradually increased for the same effect. New video tracks can also be gradually introduced into rendering by gradually increasing their rendering priority using the increasing function. This gradual "introduction" into rendering also 1603437 of 27 6 / 26 reduces sudden changes (e.g., sudden increases in quality) that can distract a user and make them feel disoriented. Rendering priority determines the weighting used by the CPU / GPU for the different video tracks used when rendering immersive video. For example, the color of a pixel in the rendered immersive video might be based on the colors of the corresponding pixels in three provided video tracks. The final color of the rendered pixel would be a combination of the colors of the three pixels, each of which is weighted. The weighting for each of the three colors typically depends, for example, on the importance of the video track from which the pixel is taken. However, by gradually changing the rendering priority of the video track, the weighting of the pixel can also be reduced. Generally, the rendering priority is based on the characteristics of the viewport and video tracks (position, field of view, etc.). When a viewer moves around a virtual scene, the viewport gradually changes, and therefore so does the rendering priority. This invention further modifies that rendering priority during the process of changing video tracks. The falling and / or rising function can be configured to change the rendering priority over a period of between 0.2 seconds and two seconds or between three frames and 60 frames, for example, five frames. Alternatively, the decreasing and / or increasing function can be an exponential decay function. In this case, a video track could be removed when the corresponding rendering priority reaches a predefined value. The predefined value would depend on the possible range of values ​​for the rendering priority. For example, for a range of 0-100, the value 1603437 of 27 The default value 7 / 26 can be 0.01, 0.1, 1, or 5, depending on the half-life of the decay function. Similarly, the video track could be deleted after a certain number of half-lives have elapsed. Video tracks do not strictly need to be "deleted" (e.g., not requested from the server) when the rendering priority reaches a low value (e.g., 0, 0, 1, etc.) or after some time has passed. For example, a video track could be retained while it has a rendering priority of 0 and is therefore not being used effectively to render immersive video. However, if low-priority video tracks are not deleted and new video tracks are added, the computing resources required to render immersive video will increase significantly. Some video tracks can be retained for a relatively short time after, for example, the rendering priority reaches a predefined value. Obtaining VT2 video tracks can be based on requesting VT2 video tracks from a server or receiving VT2 video tracks from a server. Video tracks can be requested based on the user's desired viewport. For example, the user may require a particular viewport from the immersive video and therefore request the video tracks necessary to represent that viewport. Alternatively, the server can send the user a specific set of video tracks based on the user's desired output. For example, the server might want the user to render a particular viewport and therefore send the user the video tracks necessary to render that viewport. The method may also include retaining one or more of the latest available frames from the multitrack video and / or the VT1 video track for retouching. 1603437 of 27 8 / 26 the missing data from one or more of the subsequent frames when transitioning from video tracks VT1 to video tracks VT2. Retouching missing data in new frames based on older frames can also be used to further reduce the number of objects created due to sudden changes in the video tracks used to represent the frames. The frames used for retouching could come from the multitrack video rendered from the previous video tracks, or they could be frames from the video tracks themselves, depending on the change and what information is required to retouch the new frames. The method may also include: reduce the resolution of one or more VT1 video tracks and / or one or more VT2 video tracks; reduce the frame rate of one or more VT1 video tracks and / or one or more VT2 video tracks; and / or reduce the bit rate of one or more VT1 video tracks and / or one or more VT2 video tracks. During the transition between sets of video tracks, the resources used to render multi-track video can become overloaded with data because they have to process data related to the previous video tracks and data related to the new video tracks simultaneously. Therefore, reducing the resolution, frame rate, and / or bit rate of the video tracks can reduce the amount of data that needs to be processed during the transition, ensuring that the transition can be fully rendered. In particular, this can ensure that the overall data transmitted by a user fits within any bit rate and / or pixel rate limitations. 1603437 of 27 9 / 26 The invention also provides a computer program product comprising computer program code means that, when executed on a computer device having a processing system, cause the processing system to perform all the steps of the method defined above to make the transition from a first set of video tracks, VT1, to a second set of video tracks, VT2. The invention also provides a client system for transitioning from a first set of video tracks, VT1, to a second set of video tracks, VT2, when rendering a multi-track video, wherein each video track has a corresponding rendering priority; the client system comprises: a client communications module configured to receive video tracks and the corresponding rendering priority for each video track from a server system; and a client processing system configured to: receive an instruction to transition from video tracks VT1 to video tracks VT2; receive the VT2 video tracks from the client communications module; and if the VT2 video tracks are different from the VT1 video tracks, apply: a decreasing function to the rendering priority of one or more of the video tracks in the first set of VT1 video tracks, wherein the decreasing function reduces the rendering priority over time, and / or 1603437 of 27 10 / 26 an increasing function for the rendering priority of one or more of the VT2 video tracks, wherein the increasing function increases the rendering priority over time, wherein the rendering priority of one or more of the video tracks is for use in selecting a weighting of the respective video track and / or video track elements when rendering multitrack video. The client processing system can also be configured to select the weighting and perform the rendering using the weighting or weightings. The client processing system can also be configured to retain one or more of the latest available frames from the multitrack video to retouch missing data from one or more of the subsequent frames when switching from VT1 video tracks to VT2 video tracks. The client processing system can also be configured to: request a lower resolution version of one or more VT1 video tracks and / or one or more VT2 video tracks from the server system via the client communications module; Request a lower frame rate version of one or more VT1 video tracks and / or one or more VT2 video tracks from the server system via the client communications module; and / or request a lower bit rate version of one or more VT1 video tracks and / or one or more VT2 video tracks from the server system via the client communications module. The client processing system can be configured to request a lower resolution version, a lower frame rate version, and / or a 1603437 of 27 11 / 26 lower bit rate version based on the processing capabilities of the client processing system. The invention also provides a server system for transitioning from a first set of video tracks, VT1, to a second set of video tracks, VT2, when rendering a multi-track video, wherein each video track has a corresponding rendering priority; the server system comprises: a server communications module configured to send the video tracks and the corresponding rendering priority for each video track from a client system; and a server processor system configured to: receive an instruction to transition from video tracks VT1 to video tracks VT2; send the VT2 video tracks through the server communications module; and if the VT2 video tracks are different from the video tracks corresponding to the first graphics window, VT1, an instruction is sent to the client system, through the server communications module, to apply: a decreasing function for the rendering priority of one or more of the video tracks in the first set of video tracks VT1, wherein the decreasing function decreases the rendering priority over time, and / or 1603437 of 27 12 / 26 apply an increasing function to the rendering priority of one or more of the VT2 video tracks, wherein the increasing function increases the rendering priority over time, wherein the rendering priority of one or more of the video tracks is for use in selecting a weighting of the respective video track and / or video track elements when rendering multitrack video. The server processing system can also be configured to: send a lower resolution version of one or more VT1 video tracks and / or one or more VT2 video tracks to the client system via the server communications module; Send a lower frame rate version of one or more VT1 video tracks and / or one or more VT2 video tracks to the server system via the server communications module; and / or send a lower bit rate version of one or more VT1 video tracks and / or one or more VT2 video tracks from the server to the client system via the server communications module. The server processing system can be configured to send a lower resolution version, a lower frame rate version, and / or a lower bit rate version based on the processing capabilities of a client processing system within the client system. These and other aspects of the invention will become evident from the embodiment(s) described below and will be clarified with reference to them. 1603437 of 27 13 / 26 BRIEF DESCRIPTION OF THE FIGURES For a better understanding of the invention, and to show more clearly how it can be carried out, reference will now be made, by way of example only, to the accompanying figures, in which: Figure 1 shows a 2D projection of a 3D scene in 360 degrees; Figure 2 shows an illustration of three atlases; Figure 3 shows a server system sending video cues to a client system; Figure 4 shows a first example of a video track change; and Figure 5 shows a second example of a video track change. DETAILED DESCRIPTION OF THE MODALITIES The invention will be described with reference to the figures. It should be understood that the detailed description and specific examples, while indicating illustrative embodiments of the apparatus, systems, and methods, are provided for illustrative purposes only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatus, systems, and methods of the present invention will be better understood from the following description, appended claims, and accompanying figures. It should be understood that the figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numbers are used in all figures to indicate identical or similar parts. 1603437 of 27 14 / 26 The invention provides a method for transitioning from a first set of video tracks, VT1, to a second set of video tracks, VT2, when rendering a multi-track video, wherein each video track has a corresponding rendering priority. The method comprises receiving an instruction to transition from a first set of first video tracks VT1 (hereinafter referred to simply as “VT1 video tracks”) to a second set of second video tracks VT2 (hereinafter referred to simply as “VT2 video tracks”), obtaining the VT2 video tracks, and, if the VT2 video tracks are different from the VT1 video tracks, applying a decreasing function to the rendering priority of one or more video tracks in the first set of VT1 video tracks and / or an increasing function to the rendering priority of one or more video tracks in the second set of VT2 video tracks.The decreasing and increasing functions decrease and increase rendering priority over time, respectively. Rendering priority is used to determine the weighting of a video track and / or elements within a video track used to render a multi-track video. Figure 1 shows a 2D projection of a 3D scene in 360 degrees. The 3D scene in Figure 1(a) is of an entire classroom. The 2D projection could be a 360-degree screenshot of an immersive video of a class. In an immersive video, the user may be able to rotate around the classroom (e.g., VR360) and / or move around the classroom (e.g., 6DoF video) within the video. Figure 1(b) shows the 2D projection 102 of the 3D scene with an illustrative viewport 104. The viewport 104 has a field of view of approximately 90 degrees. The viewport 104 is the final section of the 3D scene that the user will see while watching an immersive video. As the user 1603437 of 27 15 / 26 rotates and / or moves in the immersive video, viewport 104 will change to show the section of the 3D scene that the user is facing. For a user to move around in an immersive video, the 3D scene must be visible from different positions within it. This can be achieved by using various sensors (e.g., cameras) to create the 3D scene, with each sensor positioned at different angles. These sensors are then grouped together in what are called atlases. Figure 2 shows an illustration of three atlas 202 units. Each atlas 202 contains one or more sensors (e.g., cameras, depth sensors, etc.). In this illustration, only three atlas 202 units are shown, and each atlas 202 has three sensors: a main camera 204 and two additional depth sensors 206. However, the number of atlas 202 units, the number of sensors, and the types of sensors per atlas 202 may vary depending on the use case. In an alternative example, cameras 204 could provide the complete (basic) views, and sensors 206 could provide additional views. Furthermore, the views from all sensors are grouped into atlases 202, which may or may not intersect. MPEG Immersive Video (MIV) specifies a bitstream format with multiple atlases. Preferably, each atlas has at least one geometry and one texture attribute video. In the case of adaptive streaming, each atlas is expected to generate multiple types of video data that form separate video tracks (e.g., one video track per atlas). Each video track is divided into short segments (e.g., on the order of a second) to allow a client to respond to a change in the viewport position by changing the subset of video tracks requested from a server. 1603437 of 27 16 / 26 The atlas system contains multiple camera views and, through range sensors, depth estimation, or other means, can also create depth maps. The combination of all these forms a video track per atlas 202 (in this case). Cameras 204 can also be registered, placing them in a common "scene" coordinate system. Therefore, the output of each atlas 202 has a multiview+depth representation (or geometry). Figure 3 shows a server system 304 sending video tracks 302 to a client system 306. The server system 304 receives output from each of the atlases 202 and encodes each track. In this example, the encoded output of a single atlas 202 forms one video track 302. The server system 304 then determines which video tracks 302 to send to the client system 306 (e.g., based on the viewport 104 requested by the client system 306). In Figure 3, only two video tracks 302 are sent to the client system 306 out of the three potential video tracks 302 received by the server system 304 from the atlases 202. For transmission, the server system 304 can encode all the outputs from each of the atlases 202. However, since it is impractical to transmit all the video tracks 302, only a subset of video tracks 302 are fully transmitted to the client system 306. Other video tracks 302 may also be partially transmitted if certain areas of a viewport 104 cannot be predicted by the video tracks 302 that were fully transmitted. Each atlas 202 contains a group of views, and the output of each atlas 202 is encoded separately, forming a video track 302. In Figures 2 and 3, it is assumed that there is only one atlas 202 per video track 302. The atlas 202 will contain the complete views and some patches (i.e., trunk pieces) of the additional views. In other cases, an atlas 202 may contain only views 1603437 of 27 17 / 26 additional or just contain the full view and a video track 302 can be formed from the output of multiple atlas 202. When the target viewport 104 is close to one of the views present on a video track 302, the client system 306 may not need any additional video tracks 302. However, when the target viewport 104 is somewhat far from the source views, the client system 306 would benefit from having multiple video tracks 302. Multiple 302 tracks can also be used to extend the field of view for headset-dependent streaming of VR360 content. For example, each camera might have a field of view of only 50 degrees, while the entire scene could have a field of view of more than 180 degrees. In that case, the 306 client system can request 302 video tracks that match the headset's orientation instead of having the full 180-degree scene available at full resolution. There is a delay between the request for a video track and the delivery of video track 302, and the position of viewport 104 may have changed intermittently. Therefore, client system 306 may also perform sub-bit stream access to reduce resources or improve rendering quality. This does not reduce the transmitted bit rate but lowers the CPU, GPU, and memory requirements of client system 306. Furthermore, in some cases, rendering quality can be improved by removing video tracks 302 that extend beyond a target viewport 104. The requested subset of video tracks 302 is then rendered. It's important to consider that not all rendering errors are visible to a human. For example, when a large patch shifts by a few pixels (i.e., due to the encoding of depth values), this 1603437 of 27 18 / 26 would not be visible to a human observer, while the same error would cause a significant drop in an objective quality estimate such as the peak signal-to-noise ratio (PSNR). During the video track 302 switch, problematic scene elements could suddenly be rendered differently, causing abrupt visual changes. Additionally, the human visual system is particularly sensitive to motion. Therefore, the inventors proposed gradually changing the rendering priority of visual elements in a 302 video track to reduce the visibility of these objects. Generally, the number of rendering objects will increase when the viewport 104 position is farther from the basic view positions within a video track 302. However, it is possible to render a viewport 104 that is outside the display space of a video track 302, particularly when it is needed for only a couple of video frames and therefore may not be fully visible to the viewer. A sudden change in pixel values ​​due to a change in the availability of video tracks 302 is highly noticeable to humans; therefore, phasing out the rendering priority allows the change to be made over a longer period, thus reducing the perceptibility of the change. The rendering priority for each video track can be described as a single variable for each video track or a set of variables (e.g., different rendering priority variables for foreground, background, etc.). Therefore, when the rendering priority is changed (increased / decreased), this can mean that the different variables of the rendering priority change at the same rate or at different rates. For example, a 1603437 of 27 19 / 26 Each set of patches generated from an atlas can have a different variable, and this set of variables would thus form the representation priority for the atlas. However, for simplicity, the representation priority will be described as a single element in the following examples. Figure 4 shows a first example of a 302 video track switch. Two video tracks, 302a and 302b, are initially used to render the immersive video. However, the former video track 302a is switched to a new video track 302c between segment 404 and segment 406. The rendering priority of the former video track 302a is gradually reduced before the end of segment 404, and the rendering priority of the new video track 302c is gradually increased from the beginning of segment 406. In this case, video track 302b is not removed and is therefore also used to render the video. When a 302 video track has a low rendering priority (402), it will only be used when no other 302 video track exists that contains the required information. When a 302 video track has a rendering priority of zero (i.e., 0%), it's as if the 302 video track isn't there. However, when information is missing, it can be patched using the 302 video track with a zero rendering priority (402). Gradually changing the rendering priority (402) of a 302 video track reduces the sudden appearance, disappearance, or replacement of rendered objects. A gradual decrease in rendering priority 402 is achieved by applying a decay function to the rendering priority 402 of a video track 302 that will no longer be used in future segments. The decay function gradually reduces the rendering priority 402 over a period of, for example, 0.2–2 seconds or 5–200 frames. The decay function can be linear, exponential, 1603437 of 27 20 / 26 quadratic, etc. The particular shape and timing taken from 100% to 0% are arbitrary and can be chosen based on the particular use case. The same reasoning applies to an increasing function used in the new 302c video track. Note that in Figure 4, it is assumed that client 306 is aware of the upcoming video track change 302. Typically, client 306 requests segments of video tracks 302. However, if server 304 is in control of the transmitted video tracks 302, server 304 can send a message as part of the video data (e.g., an SEI message) to inform client 306 of the change. Note that there is a disadvantage to encoding having very short segments, and therefore it is conceivable that a server 304 could communicate with a client 306 on a subsegment timescale. Client 306 could also (gradually) switch to a lower-resolution version of the older video track 302a before downsizing it completely and / or first request a low-resolution version of the new video track 302c. This would allow client 306 to maintain all three video tracks 302a, 302b, and 302c before switching to a higher resolution and downsizing the older video track 302a. A view blending engine can take resolution into account when prioritizing video tracks 302, and in this case, the rendering priority of a downscaled video track 402 can be further reduced. However, a sudden reduction in resolution may be observed. Therefore, it can be beneficial to make the change gradually. Client 306 could also (gradually) switch to a lower frame rate version of the older video track 302a before downgrading it completely, or first request a lower frame rate version of the new video track 302c. This could allow client 306 to maintain multiple tracks. 1603437 of 27 21 / 26 of video 302 (within the resource constraints for client 306), before switching to a higher frame rate and bringing down the previous video track 302a. Figure 5 shows a second example of a 302 video track switch. The switch to 302 video tracks can be made more gradual when video data is received for multiple 302 video tracks by changing the rendering priority 402 for one or more of the multiple 302 video tracks. However, when client 306 receives only one 302 video track, a different strategy may be needed. In Figure 5, the quality of the previous 302a video track is also reduced to allow simultaneous transmission of the previous 302a video track and the new 302c video track for a single 502 segment. Some examples of splitting the resources required to transmit / decode / render a 302 video track are: - reduce the bit rate (e.g., 10 Mbps to 5 Mbps), - reduce spatial resolution (e.g. 8K χ 4K ^ 6K χ 3K), - reduce the frame rate (e.g., 120 Hz ^ 60 Hz). Specifically, Figure 7 shows that the old video track 302a is reduced from a frame rate of 60 FPS to 30 FPS when the new video track 302c is introduced. The new video track 302a is also introduced at 30 FPS and is then increased to 60 FPS after the old video track 302a is removed. The rendering priority 402 of the old video track 302a is gradually reduced, and the rendering priority 402 of the new video track 302c is gradually increased during a single segment 502. An additional enhancement that is especially suitable for static regions (such as contexts) is to retain the last frame of video track 302 1603437 of 27 22 / 26 available and use it for a few more frames to allow for a more gradual reduction. The last available frame of video track 302 is then extrapolated to reflect new viewpoints. This improvement is especially useful for low-latency networks where the time to be bridged is short. When client 306 is able to convert the video frame rate with motion compensation, this could be used to extrapolate the motion within video track 302 based on the motion vectors in the last available frame. A graphics window typically depends on the viewpoint (position, extrinsic data) and the camera characteristics (intrinsic data) such as field of view and focal length. The client system can also retain the rendered viewport (frame slider) and render from that viewport to fill in the missing data in a new frame. This strategy might be allowed only when a change is made to the 302 video tracks or, by default, as a fallback strategy. The frame slider can be rendered with an offset to compensate for changes in viewer position. This is also a good fallback strategy for when (for a short time, for example, due to network delays) a new 302c video track is not available. The change / transition in video tracks 302 for immersive videos refers to 3DoF, 3DoF+, and 6DoF videos with multiple tracks. They can also be called spherical videos. A rendering priority is essentially an umbrella term for one or more parameters that can be modified to change how much of a video element contributes to the final rendered output. In other words, rendering priority refers to one or more parameters that affect the 1603437 of 27 23 / 26 Weighting of particular video elements (e.g., pixels) during the rendering of a new image. Rendering priority can also be referred to as a final mix weight. For example, view blending is based on a weighted sum of contributions, whereby each view has a weight (uniform or varying per pixel). Each view can be rendered on a separate slider, and the sliders can be blended based on the viewport's proximity to each view. In this example, the viewport's proximity to each view is used to determine rendering priority. Conventionally, rendering priority is based on proximity to the viewport. It is also proposed to adjust rendering priority (i.e., increase or decrease it) when video tracks change. For example, when three views [v1, v2, v3] are being rendered, but for the next segment, views [v2, v3, v4] would be available, then v1's contribution to the weighted sum could be gradually reduced to zero. When v4 comes in, its contribution (i.e., rendering priority) could be gradually increased, starting from zero. Alternatively, it's also possible to render a depth map in the viewport perspective and then obtain the texture from multiple views. In this case, a weighting function will also be used to decide how to blend the multiple texture contributions. The parameters that can be used to determine rendering priority are proximity to the output view and the difference in depth value between the rendered depth and the source depth. Other representation solutions may represent only part of the data, in which case there may be a threshold below which the data is discarded. 1603437 of 27 24 / 26 For example, when rendering patches or blocks of a video frame, the rendering priority can be based on proximity, patch size, patch resolution, etc. It is proposed to adapt the rendering priority based on the future availability / need of the patch. In summary, rendering priority can be based on the proximity of the source view to the rendered view, differences in depth, patch size, and patch resolution. Other methods for determining rendering priority may be known to an expert. The expert would easily be able to develop a processor to carry out any method described herein. Therefore, each stage of a flowchart can represent a different action performed by a processor, and can be implemented using a respective processing module. The processor can be implemented in many ways, using software and / or hardware, to perform the various required functions. Typically, the processor uses one or more microprocessors that can be programmed using software (e.g., microcode) to perform the required functions. Alternatively, the processor can be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuit systems to perform other functions. Examples of circuit systems that can be used in various modalities of the present description include, but are not limited to, conventional microprocessors, application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs). 1603437 of 27 25 / 26 In various implementations, the processor can be associated with one or more storage media, such as volatile and non-volatile computer memory, including RAM, PROM, EPROM, and EEPROM. The storage media can be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform the required functions. Various storage media can be fixed within a processor or controller or be portable, so that one or more programs stored on them can be loaded into a processor. Those skilled in the art can understand and implement variations of the described embodiments of the claimed invention by studying the figures, the description, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" does not exclude a plurality. A single processor or other unit can perform the functions of several elements mentioned in the claims. The mere fact that certain measures are mentioned in mutually dependent claims does not indicate that a combination of these measures cannot be used advantageously. A computer program can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied with, or as part of, other hardware, but it can also be distributed in other ways, such as through the Internet or other wired or wireless telecommunications systems. 1603437 of 27 26 / 26 If the term “adapted for” is used in the claims or description, it is noted that the term “adapted for” is intended to be equivalent to the term “configured for”. Any reference sign in the claims shall not be construed as limiting the scope. 1603437 of 27 Santiago Ferrer Reyes - 20142227887 Digitally signed by PORTALTRAMITES - INPI Date: 2021.12.06 09:11:53 -03:00 Reason: Digitally Signed by INPI Location: Buenos Aires, Argentina 1603437

Claims

1. A method for transitioning video tracks in multi-track immersive videos, characterized in that it comprises: receiving an instruction to transition from first video tracks to second video tracks; obtaining the second video tracks; and applying a function to a rendering priority of at least one of the first video tracks if the first video tracks are different from the second video tracks, wherein the function decreases or increases the rendering priority over time, wherein the rendering priority is used to select a weighting of at least one of the first video tracks. Fifteen claims follow.