Method and device for viewport-based atlas selection
Patent Information
- Application Number
- KR1020240039196
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2044-03-21
Smart Images

Figure 112024031907000-PAT00024_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to a segmented decoding and reproduction technology for providing seamless omnidirectional video capable of supporting motion parallax in response to the viewer's left-right / up-down rotation as well as left-right / up-down movement, in order to play natural omnidirectional video through a VR terminal. Background Technology
[0002] Virtual reality services are evolving in a direction that maximizes immersion and a sense of presence by generating omnidirectional video in the form of real-life images or computer graphics (CG) and playing it on HMDs, smartphones, etc. It is known that currently, to play natural and immersive omnidirectional video through an HMD, it must support 6 degrees of freedom (DoF). 6DoF video must provide free video in six directions through the HMD screen, including (1) left-right rotation, (2) up-down rotation, (3) left-right movement, and (4) up-down movement. However, most omnidirectional videos based on real-life images currently only support rotational movement. Accordingly, research is actively underway in fields such as the acquisition and reproduction technologies of 6DoF omnidirectional video. The problem to be solved
[0003] The present disclosure aims to provide a method for selecting an optimal group or atlas based on user viewpoint information in partially decoding and rendering immersive video divided into groups to reproduce a wide range of immersive video without interruption. means of solving the problem
[0004] The viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure may include the steps of buffering and preprocessing metadata and user viewpoint information, calculating a correlation value representing the correlation of user viewpoint positions for each atlas based on the metadata and user viewpoint information, and selecting a group or atlas to be used for MIV decoding and viewpoint image synthesis based on the correlation value.
[0005] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on a baseline between the viewport and each camera.
[0006] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, when camera grouping is regular based on a camera arrangement structure, the step of selecting said group or atlas may be performed by considering only the decision boundary without considering baselines with all cameras.
[0007] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the selected group or atlas is selected as the group or atlas with the highest correlation value with the user viewpoint location among a plurality of groups or atlases, and the plurality of groups or atlases may be divided by the decision boundary.
[0008] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, in the step of selecting the group or atlas, if the highest correlation value of the user viewpoint location is at a location that exceeds a threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint location, the group or atlas with the highest correlation value with the user viewpoint location is selected, and if the highest correlation value of the user viewpoint location is at a location that does not exceed the threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint location, the group with the highest correlation value with the previous user viewpoint location may be selected.
[0009] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on the difference in viewpoint direction between the viewport and each camera.
[0010] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on the visibility of each camera to the viewport.
[0011] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the visibility may be determined based on how much the sample points in the frustum space of each camera are projected into the image plane of each camera.
[0012] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on a baseline between the viewport and each camera, a difference in viewpoint orientation between the viewport and each camera, and visibility of each camera with respect to the viewport.
[0013] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art from the description below. Effects of the invention
[0014] According to the configuration of the present disclosure, an optimal group or atlas can be selected based on user viewpoint information and metadata to reproduce a wide range of immersive videos seamlessly by partially decoding and rendering immersive videos divided into groups.
[0015] Correlation between each group can be calculated using viewport information and metadata, and the selected group and atlas IDs can be output based on this. In calculating the correlation, a function based on the baseline between the viewport and each camera, the viewpoint orientation between the viewport and each camera, and the visibility between the viewport and each camera can be used.
[0016] According to the configuration of the present disclosure, by switching the groups or atlases required for rendering as the user's viewpoint position changes and performing partial decoding and partial spatial rendering, a wide range of immersive video can be reproduced seamlessly even under the limited performance and limited system resources of the terminal.
[0017] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below. Brief explanation of the drawing
[0018] FIG. 1 is a block diagram of an immersive image processing device according to one embodiment of the present disclosure. FIG. 2 is a block diagram of an immersive image output device according to one embodiment of the present disclosure. FIG. 3 illustrates the overall structure of an immersive video receiving system based on the configuration of the present disclosure. Figure 4 is a diagram for specifically explaining the operation process of the viewport-based atlas selection unit. FIG. 5 illustrates embodiments for cases where the unit of switching is a group and an atlas. Figure 6 illustrates an example of a camera baseline-based VAS algorithm. FIG. 7 illustrates an example of selecting a group or atlas using only a decision boundary. FIG. 8 illustrates an embodiment for selecting a group or atlas based on the difference in viewpoint direction between the viewport and each camera. FIG. 9 illustrates an embodiment for selecting a group or atlas based on the visibility of each camera to the viewport. Specific details for implementing the invention
[0019] The present disclosure is subject to various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals in the drawings refer to the same or similar functions across various aspects. The shapes and sizes of elements in the drawings may be exaggerated for clearer explanation. The detailed description of exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments as examples. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present disclosure in relation to one embodiment. It should also be understood that the location or arrangement of individual components within each disclosed embodiment may be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of exemplary embodiments is limited only by the appended claims, together with all equivalents to those claimed therein, provided they are properly described.
[0020] In this disclosure, terms such as first, second, etc. may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of this disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0021] Where it is stated that any component of the present disclosure is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, or that there may be other components in between. On the other hand, where it is stated that a component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0022] The components shown in the embodiments of the present disclosure are depicted independently to represent different characteristic functions and do not imply that each component consists of separate hardware or a single software unit. That is, each component is listed and included as a separate component for convenience of explanation; however, at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separated embodiments of each component are included within the scope of the rights of the present disclosure as long as they do not depart from the essence of the present disclosure.
[0023] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this disclosure, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, the description in this disclosure that a specific configuration "comprising" does not exclude configurations other than that configuration, but means that additional configurations may be included within the scope of the practice or technical concept of this disclosure.
[0024] Some components of the present disclosure may not be essential components performing an essential function in the present disclosure, but may be optional components merely for enhancing performance. The present disclosure may be implemented by including only the components essential to embody the essence of the present disclosure, excluding components used merely for enhancing performance, and a structure including only the essential components, excluding optional components used merely for enhancing performance, is also included within the scope of the rights of the present disclosure.
[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of this specification, such detailed description is omitted, and the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0027] Immersive video refers to video in which the viewport can dynamically change when the user's viewing position changes. To implement immersive video, multiple input videos are required. Each of these multiple input videos can be referred to as a source video or a viewpoint video. A different view index can be assigned to each viewpoint video. Immersive video can be composed of videos with different viewpoints, and accordingly, immersive video may also be referred to as multi-viewpoint video.
[0028] Immersive images can be classified into 3DoF (Degree of Freedom), 3DoF+, Windowed-6DoF, or 6DoF types. 3DoF-based immersive images can be implemented using only texture images. On the other hand, to render immersive images containing depth information, such as 3DoF+ or 6DoF, depth images (or depth maps) are required in addition to texture images.
[0029] The embodiments described below are assumed to be for immersive image processing including depth information such as 3DoF+ and / or 6DoF. Additionally, the viewpoint image is assumed to consist of a texture image and a depth image.
[0031] FIG. 1 is a block diagram of an immersive image processing device according to one embodiment of the present disclosure.
[0032] Referring to FIG. 1, an immersive image processing device according to the present disclosure may include a view optimizer (110), an atlas generation unit (120), a metadata generation unit (130), an image encoder unit (140), and a bitstream generation unit (150).
[0033] An immersive image processing device receives multiple pairs of images, camera intrinsic variables, and camera extrinsic variables as input values and encodes the immersive image. Here, the multiple pairs of images include texture images (Attribute component) and depth images (Geometry component). Each pair may have a different view. Accordingly, a pair of input images may be referred to as a viewpoint image. Each of the viewpoint images may be distinguished by an index. In this case, the index assigned to each viewpoint image may be referred to as a view or view index.
[0034] Intrinsic camera variables include focal length and principal point position, etc., and extrinsic camera variables include camera position and orientation, etc. Intrinsic and extrinsic camera variables can be treated as camera parameters or (user) viewpoint parameters.
[0035] The viewpoint optimization unit (110) divides the viewpoint images into multiple groups. By dividing the viewpoint images into multiple groups, independent encoding processing can be performed for each group. For example, viewpoint images captured by N spatially consecutive cameras can be classified into one group. In this way, viewpoint images with relatively coherent depth information can be grouped into one group, and accordingly, rendering quality can be improved.
[0036] Additionally, the viewpoint optimization unit (110) can classify viewpoint images into a basic image and an additional image. The basic image represents a viewpoint image with the highest pruning priority that is not pruned, and the additional image represents a viewpoint image with a lower pruning priority than the basic image.
[0037] Additionally, the viewpoint optimization unit (110) can determine at least one of the viewpoint images as the base image. Viewpoint images that are not selected as the base image may be classified as additional images.
[0038] The atlas generation unit (120) can generate a pruning mask by performing pruning. Then, it can extract patches using the pruning mask and generate an atlas by combining the base image and / or the extracted patches. If the viewpoint images are divided into multiple groups, the above process can be performed independently for each group.
[0039] The generated atlas may consist of a texture atlas and a depth atlas. The texture atlas represents an image combining base texture images and / or texture patches, and the depth atlas represents an image combining base depth images and / or depth patches.
[0040] The atlas generation unit (120) may include a pruning unit (122), an aggregation unit (124), and a patch packing unit (126).
[0041] The pruning unit (122) performs pruning on additional images based on pruning priority. Specifically, pruning on additional images can be performed using a reference image that has a higher pruning priority than additional images.
[0042] As a result of performing pruning, a pruning mask may be generated that includes information on whether each pixel in the additional image is valid or invalid. The pruning mask may be a binary image indicating whether each pixel in the additional image is valid or invalid. For example, in the pruning mask, pixels determined to be duplicate data with the reference image may have a value of 0, and pixels determined not to be duplicate data with the reference image may have a value of 1.
[0043] The set unit (124) combines the pruning masks generated on a frame-by-frame basis on an intra-period basis.
[0044] Additionally, the collection unit (124) can extract patches from the combined pruning mask image through a clustering process. Specifically, a rectangular area containing valid data within the combined pruning mask image can be extracted as a patch. Since the patch is extracted in a rectangular shape regardless of the shape of the valid area, the patch extracted from the non-rectangular valid area may contain not only valid data but also invalid data.
[0045] At this time, the set unit (124) can re-divide an L-shaped or C-shaped patch that reduces encoding efficiency. Here, the L-shaped patch indicates that the distribution of the effective area is L-shaped, and the C-shaped patch indicates that the distribution of the effective area is C-shaped.
[0046] When the distribution of valid regions is L-shaped or C-shaped, the area occupied by non-valid regions within the patch is relatively large. Accordingly, L-shaped or C-shaped patches can be divided into multiple patches to improve coding efficiency.
[0047] For unpruned viewpoint images, the entire viewpoint image can be treated as a single patch. Specifically, the entire 2D image obtained by unfolding the unpruned viewpoint image into a predetermined projection format can be treated as a single patch. The projection format may include at least one of an Equirectangular Projection Format (ERP), a Cube-map, or a Perspective Projection Format.
[0048] Here, the unpruned viewpoint refers to the base image with the highest pruning priority. Alternatively, an additional image that does not contain duplicate data with the base image and the reference image may be defined as an unpruned viewpoint. Alternatively, an additional image that is arbitrarily excluded from the pruning target, regardless of whether duplicate data with the reference image exists, may also be defined as an unpruned viewpoint. In other words, even if an additional image contains data that duplicates the reference image, it can still be defined as an unpruned viewpoint.
[0049] The packing unit (126) can pack patches into a square-shaped image. When packing patches, modifications such as resizing, rotation, or flipping of the patches may be involved. An image in which patches are packed can be defined as an atlas.
[0050] Specifically, the packing unit (126) can generate a texture atlas by packing basic texture images and / or texture patches, and can generate a depth atlas by packing basic depth images and / or depth patches.
[0051] The base image can be treated as a single patch in its entirety. In other words, the base image can be packed into the atlas as is. When the entire image is treated as a single patch, that patch may also be referred to as a Complete View or a Complete Patch.
[0052] The number of atlases generated by the atlas generation unit (120) can be determined based on at least one of the arrangement structure of the camera rig, the accuracy of the depth map, or the number of viewpoint images.
[0053] The metadata generation unit (130) generates metadata for image synthesis. The metadata may include at least one of camera-related data, pruning-related data, atlas-related data, or patch-related data.
[0054] Pruning-related data may include information for determining pruning priorities among viewpoint images.
[0055] Atlas-related data may include at least one of atlas size information, atlas count information, priority information among atlases, or a flag indicating whether the atlas contains a complete image. The atlas size may include at least one of texture atlas size information and depth atlas size information. In this case, a flag indicating whether the depth atlas size is the same as the texture atlas size may be additionally encoded. If the depth atlas size is different from the texture atlas size, depth atlas reduction ratio information (e.g., scaling-related information) may be additionally encoded. Atlas-related information may be included in the "View parameters list" item within the bitstream.
[0056] Patch-related data includes information for specifying the location and / or size of the patch within the atlas image, the time-lapse image to which the patch belongs, and the location and / or size of the patch within the time-lapse image. For example, at least one of location information indicating the location of the patch within the atlas image or size information indicating the size of the patch within the atlas image may be encoded. Additionally, a source index for identifying the time-lapse image from which the patch originated may be encoded. The source index represents the index of the time-lapse image that is the original source of the patch. Additionally, location information indicating the location corresponding to the patch within the time-lapse image or location information indicating the size corresponding to the patch within the time-lapse image may be encoded. The patch-related information may be included in the "Atlas data" item within the bitstream.
[0057] The video encoder unit (140) encodes the atlas. If the viewpoint images are classified into multiple groups, an atlas can be generated for each group. Accordingly, video encoding can be performed independently for each group.
[0058] The image encoder unit (140) may include a texture image encoder unit (142) that encodes a texture atlas and a depth image encoder unit (144) that encodes a depth atlas.
[0059] The bitstream generation unit (150) generates a bitstream based on encoded video data and metadata. The generated bitstream can be transmitted to an immersive video output device.
[0061] FIG. 2 is a block diagram of an immersive image output device according to one embodiment of the present disclosure.
[0062] Referring to FIG. 2, an immersive video output device according to the present disclosure may include a bitstream parsing unit (210), a video decoder unit (220), a metadata processing unit (230), and a video synthesis unit (240).
[0063] The bitstream parsing unit (210) can parse video data and metadata from the bitstream. The video data may include data from an encoded atlas. If an arbitrary spatial access service is supported, only a partial bitstream including the user's viewing location may be received.
[0064] The image decoder unit (220) can decode parsed image data. The image decoder unit (220) may include a texture image decoder unit (222) for decoding a texture atlas and a depth image decoder unit (224) for decoding a depth atlas.
[0065] The metadata processing unit (230) can unformat the parsed metadata.
[0066] Unformatted metadata can be used to synthesize images at a specific point in time. For example, when user movement information is input to an immersive video output device, the metadata processing unit (230) can determine an atlas required for image synthesis, patches required for image synthesis, and / or the position / size of said patches within the atlas in order to reproduce a viewport image according to the user's movement.
[0067] The image synthesis unit (240) can dynamically synthesize a viewport image according to the user's movement. Specifically, the image synthesis unit (240) can extract patches necessary for synthesizing a viewport image from an atlas using information determined by the metadata processing unit (230) according to the user's movement. Specifically, it can extract an atlas containing information on a viewpoint image necessary for synthesizing a viewport image and patches extracted from the viewpoint image within the atlas, and synthesize the extracted patches to generate a viewport image.
[0069] To effectively deliver immersive video that supports 6DoF of rotational motion and positional movement to HMD users, a system based on MIV (MPEG Immersive Video), a standard for immersive video encoding and transmission by the MPEG standardization group, can be configured. The MIV encoder receives texture images, geometric images, and internal / external camera parameters for multiple acquired viewpoints, minimizes overlap between viewpoints, extracts only the pixel data necessary for rendering, and outputs texture atlas images, geometric atlas images packed with other information necessary for reproduction.
[0070] In addition, to reproduce a wide range of immersive video, the MIV encoder groups the images into multiple groups and encodes them in parallel, and the MIV decoder can partially decode and render only the groups or atlases required by the user. In this case, the multiple groups are data divided based on a predetermined number of viewpoint units or visual similarity between viewpoint images, and may be MIV data representing a certain portion of the total viewing space. By switching the groups or atlases required for rendering as the user's viewpoint position changes and performing the aforementioned partial decoding and partial space rendering, a wide range of immersive video can be reproduced seamlessly even under the limited performance and system resources of the terminal.
[0071] Here, in order for the configuration of the above invention to operate smoothly, a method for selecting an optimal group or atlas for segmented decoding and rendering using user viewpoint information (position, direction, movement speed, etc.) input from an external display device (commercial HMD, etc.) may be required.
[0073] FIG. 3 illustrates the overall structure of an immersive video receiving system based on the configuration of the present disclosure.
[0074] The overall structure of the immersive video receiving system of Fig. 3 can be based on the configuration of Fig. 2.
[0075] Immersive video can be encoded in MIV (MPEG Immersive Video) and transmitted to a receiving system. Among the MIV-encoded data, atlas video data can be encoded and transmitted.
[0076] The video decoding unit can receive an encoded atlas video bitstream and decode the atlas video. Here, the atlas video bitstream may be a bitstream selected based on a group or atlas ID.
[0077] The MIV decoding unit can perform MIV decoding using selected atlas videos and metadata among the decoded atlas videos. The metadata may include a group or atlas ID.
[0078] The viewpoint image synthesis and reproduction unit can synthesize and reproduce binocular images corresponding to the user's viewpoint using atlas video and viewpoint information decoded through the MIV decoding unit. In addition, the viewpoint image synthesis and reproduction unit can be linked with an external reproduction device, such as a commercial HMD, to receive user viewpoint information and synthesize and reproduce the corresponding viewpoint images in real time.
[0079] Figure 4 is a diagram for specifically explaining the operation process of the viewport-based atlas selection unit.
[0080] As mentioned above, to reproduce a wide range of immersive video, MIV-encoded data may be data encoded in parallel by grouping the entire space into multiple groups. In this case, one group of MIV data may be referred to as TinyMIV. In such a case, the immersive video receiving system may need to select the optimal atlas or group required by the user based on user viewpoint information. To this end, the viewport-based atlas selector (hereinafter VAS) of the present invention may operate as shown in FIG. 4.
[0081] The VAS may receive at least one of the metadata received by the immersive video receiving system or user viewpoint information (viewport information) received from the viewpoint image synthesis and reproduction unit. In this case, the user viewpoint information received from the viewpoint image synthesis and reproduction unit may be information that has undergone the synthesis and reproduction processes in the viewpoint image synthesis and reproduction unit. Alternatively, the user viewpoint information received from the viewpoint image synthesis and reproduction unit may be information that receives the viewport information received from an external reproduction device as is, without undergoing the synthesis and reproduction processes in the viewpoint image synthesis and reproduction unit. Or, the user viewpoint information received from the viewpoint image synthesis and reproduction unit may be information that receives the viewport information received from an external reproduction device as is, without undergoing the synthesis and reproduction processes in the viewpoint image synthesis and reproduction unit, and has only undergone coordinate system transformation.
[0082] Referring to FIG. 4, the VAS may include at least one of an input information buffering and preprocessing unit, a viewport-atlas correlation calculation unit, or a group or atlas selection unit.
[0083] The input information buffering and preprocessing unit can perform input information buffering and preprocessing steps. Specifically, it can select and buffer parameters (VAS parameters) required for VAS operation within the received metadata. Additionally, user viewpoint information (viewport information) can also be buffered. Furthermore, it can select parameters required for calculating viewport-atlas correlation and transmit them to the viewport-atlas correlation calculation unit.
[0084] In the viewport-atlas correlation calculation unit, a viewport-atlas correlation calculation step can be performed. Specifically, based on viewport information (user viewpoint information), a value corresponding to how much correlation there is with the viewpoint position for each atlas can be output as a correlation value C().
[0085] The group or atlas selection unit may perform a group or atlas selection step. Specifically, the viewport-atlas correlation calculation unit may receive an atlas correlation value from the viewport-atlas correlation calculation unit and output a group or atlas ID required for MIV decoding and viewpoint image synthesis.
[0086] The viewport-atlas correlation calculation unit can output a group ID if the unit for split-rendering multiple immersive videos in the receiving system is a group, and can output an atlas ID if the unit for split-rendering multiple immersive videos is an atlas unit.
[0087] FIG. 5 illustrates embodiments for cases where the unit of switching is a group and an atlas.
[0088] In the group or atlas selection section, the optimal group or atlas can be selected based on the correlation value C( ) for each output group or atlas. When the unit of switching is a group, C(V, calculated between the viewport and each group You can select the optimal group using the formula below.
[0089] [Formula 1]
[0090] The correlation calculated in the viewport-atlas correlation calculation unit can be defined differently depending on various factors such as camera placement, grouping structure, and content characteristics.
[0091] Based on this, the group or atlas selection section can select a group or atlas using any one of four methods.
[0092] [Method 1]
[0093] Figure 6 illustrates an example of a camera baseline-based VAS algorithm.
[0094] Referring to Fig. 6, when spatially adjacent cameras are grouped together, the VAS can select a group or atlas based on the baseline between the viewport and each camera.
[0095] [Equation 2]
[0096] As shown in the formula above, when the viewport (virtual camera) position is denoted as V, V and the k-th group All cameras constituting { |i∈ Baseline B between} By summing , V), it can be determined that the smaller this value, the higher the correlation between V and the k-th group. Here, M can be any positive number.
[0097] The above formula is an example and does not necessarily use an inverse relationship, and the larger the sum of the baselines All functions that produce a small ( ) can be used.
[0098] FIG. 7 illustrates an example of selecting a group or atlas using only a decision boundary.
[0099] In [Method 1], if the camera arrangement is structured and grouping is regular based on that structure, a group or atlas can be selected using only the decision boundary without calculating baselines with all cameras.
[0100] An example of such a case may be as shown in FIG. 7-(a). Referring to FIG. 7-(a), there are a total of four omnidirectional camera rigs in scene space, the centers of each camera rig are spaced apart by the same distance, and each camera rig forms a group, which may be an example of group-based MIV encoding G0 to G3.
[0101] In this case, the group with the highest correlation obtained using [Equation 2] above may be exactly the same as the decision boundary-based group selection calculated using only the distance from the center of the camera rig. Therefore, computational complexity can be reduced by selecting the optimal group based only on whether it has crossed the decision boundary.
[0102] In [Method 1], when the user's viewpoint moves irregularly near the location where the correlation values reverse between groups (e.g., the decision boundary in Fig. 7-(a)), switching between groups may occur too frequently. To compensate for this, switching between groups can be performed only when the correlation value of another group is higher than that of the current group by a certain threshold. An example of applying such a switching threshold is shown in Fig. 7-(b).
[0104] [Method 2]
[0105] FIG. 8 illustrates an embodiment for selecting a group or atlas based on the difference in viewpoint direction between the viewport and each camera.
[0106] VAS can select groups or atlases based on the difference in viewpoint orientation between the viewport and each camera.
[0107] [Equation 3]
[0108] In Equation 3, A(V, ) is virtual camera V and the i-th camera It can represent the difference (angle) in the viewpoint direction. As shown in Equation 3, V and the k-th group All cameras constituting { |i∈ Difference in viewpoint direction between} A( By summing , V), it can be determined that the smaller this value, the higher the correlation between V and the k-th group. As in [Method 1], M is an arbitrary positive number, and since the above formula is merely one example, the reciprocal relationship does not necessarily have to be used. In other words, the larger the sum of the differences in viewpoint directions, All functions that produce a small ( ) can be used.
[0110] [Method 3]
[0111] FIG. 9 illustrates an embodiment for selecting a group or atlas based on the visibility of each camera to the viewport.
[0112] VAS can select a group or atlas based on the visibility of each camera to the viewport.
[0113] In this embodiment, a frustum of the virtual viewpoint camera can be defined according to the viewport position.
[0114] If the near and far distances of the scene are known, the corresponding values are used; otherwise, they can be defined arbitrarily. After defining the frustum, N sample points can be defined within that space. Then, these N samples are assigned to each input camera It can be projected onto the image plane. Some points out of the total samples Based on whether it was projected into the image plane, V and You can obtain liver visibility.
[0115] [Equation 4]
[0116] As in Equation 4, The sum of visibility with V of all cameras within It can be used as a correlation between and V. The visibility value basically uses the ratio of how many of the defined N sample points are projected into the image, and each sample point can be assigned a different weight and reflected in Vis(). For example, a higher weight can be assigned as the sample point is located closer to the center of the Frustum, and a lower weight as it is located further out. Conversely, a higher weight can be assigned as the sample point is located further out of the Frustum, and a lower weight as it is located closer to the center.
[0118] [Method 4]
[0119] Based on a combination of all or part of [Method 1], [Method 2] and [Method 3], the VAS can select a group or an atlas.
[0120] [Formula 5]
[0121] Here, α, β, and γ can be 0 or any positive number. Here, 0 may indicate that the corresponding method is not used. For example, if only [Method 1] and [Method 2] are applied, γ may be 0. For example, if only [Method 1] and [Method 3] are applied, β may be 0. For example, if only [Method 2] and [Method 3] are applied, α may be 0.
[0122] Additionally, the values of α, β, and γ may be predefined values or values determined from VAS parameters or user viewpoint information. Additionally, the sum of α, β, and γ may be 1.
[0123] The group or atlas selection unit can output the ID of the group or atlas most suitable for the user based on the total correlation obtained, and based on this, the receiving system of FIG. 3 can select only the corresponding atlas.
[0125] The viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure may include the steps of buffering and preprocessing metadata and user viewpoint information, calculating a correlation value representing the correlation of user viewpoint positions for each atlas based on the metadata and user viewpoint information, and selecting a group or atlas to be used for MIV decoding and viewpoint image synthesis based on the correlation value.
[0126] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on a baseline between the viewport and each camera.
[0127] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, when camera grouping is regular based on a camera arrangement structure, the step of selecting said group or atlas may be performed by considering only the decision boundary without considering baselines with all cameras.
[0128] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the selected group or atlas is selected as the group or atlas with the highest correlation value with the user viewpoint location among a plurality of groups or atlases, and the plurality of groups or atlases may be divided by the decision boundary.
[0129] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, in the step of selecting the group or atlas, if the highest correlation value of the user viewpoint location is at a location that exceeds a threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint location, the group or atlas with the highest correlation value with the user viewpoint location is selected, and if the highest correlation value of the user viewpoint location is at a location that does not exceed the threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint location, the group with the highest correlation value with the previous user viewpoint location may be selected.
[0130] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on the difference in viewpoint direction between the viewport and each camera.
[0131] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on the visibility of each camera to the viewport.
[0132] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the visibility may be determined based on how much the sample points in the frustum space of each camera are projected into the image plane of each camera.
[0133] In the viewport-based atlas selection method, apparatus, and recording medium according to the present disclosure, the step of selecting the group or atlas may be performed based on a baseline between the viewport and each camera, a difference in viewpoint orientation between the viewport and each camera, and visibility of each camera with respect to the viewport.
[0135] The exemplary methods of the present disclosure are described as a series of operations for clarity of description, but this is not intended to limit the order in which the steps are performed, and if necessary, each step may be performed simultaneously or in a different order. To implement the method according to the present disclosure, additional steps may be included in addition to the steps exemplified, steps excluding some steps and including the remaining steps, or steps excluding some steps and including additional steps.
[0136] The various embodiments of the present disclosure are not intended to list all possible combinations but to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0137] In addition, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, it may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0138] The scope of the present disclosure includes software or machine-executable instructions (e.g., an operating system, an application, firmware, a program, etc.) that enable an operation according to a method of various embodiments to be executed on a device or computer, and a non-transitory computer-readable medium on which such software or instructions, etc. are stored and which is executable on a device or computer.
Claims
Claim 1 A step of buffering and preprocessing metadata and user viewpoint information in an input information buffering and preprocessing unit; a step of calculating a correlation value representing the correlation of user viewpoint positions for each atlas based on the metadata and user viewpoint information in a viewport-atlas correlation calculation unit; A viewport-based atlas selection method comprising the step of selecting a group or atlas used for MIV decoding and viewpoint image synthesis based on the correlation value in a group or atlas selection unit, wherein, when camera grouping is regular based on a camera arrangement structure, the step of selecting the group or atlas is performed by considering only the decision boundary without considering baselines with all cameras, and in the step of selecting the group or atlas, if the highest correlation value of the user viewpoint position is at a position exceeding a threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint position, the group or atlas with the highest correlation value with the user viewpoint position is selected, and in the step of selecting the group or atlas, if the highest correlation value of the user viewpoint position is at a position not exceeding the threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint position, the group with the highest correlation value with the previous user viewpoint position is selected. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 In claim 1, the step of selecting the group or atlas is performed based on the difference in viewpoint direction between the viewport and each camera, a viewport-based atlas selection method. Claim 7 A viewport-based atlas selection method according to claim 1, wherein the step of selecting the group or atlas is performed based on the visibility of each camera to the viewport. Claim 8 In claim 7, the visibility is determined based on how much the sample points in the frustum space of each camera are projected into the image plane of each camera, a viewport-based atlas selection method. Claim 9 A viewport-based atlas selection method according to claim 1, wherein the step of selecting the group or atlas is performed based on a baseline between the viewport and each camera, a difference in viewpoint direction between the viewport and each camera, and visibility for each camera relative to the viewport. Claim 10 An input information buffering and preprocessing unit that buffers and preprocesses metadata and user viewpoint information; a viewport-atlas correlation calculation unit that calculates a correlation value representing the correlation of user viewpoint positions for each atlas based on the metadata and user viewpoint information; and a group or atlas selection unit that selects a group or atlas used for MIV decoding and viewpoint image synthesis based on the correlation value, wherein, when camera grouping is regular based on a camera placement structure, the group or atlas selection unit selects a group or atlas by considering only the decision boundary without considering baselines with all cameras, and in the group or atlas selection unit, if the highest correlation value of the user viewpoint position is located at a position that exceeds a threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint position, the group or atlas with the highest correlation value with the user viewpoint position is selected, and in the group or atlas selection unit, the group or atlas with the highest correlation value of the user viewpoint position is determined to be the group or atlas with the highest correlation value with the previous user viewpoint position. A viewport-based atlas selection device in which, when a position is located at a position that does not exceed the threshold from the boundary, the group with the highest correlation value with the previous user viewpoint position is selected. Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 In claim 10, the group or atlas selection unit selects a group or atlas based on the difference in viewpoint direction between the viewport and each camera, a viewport-based atlas selection device. Claim 16 In claim 10, the group or atlas selection unit selects a group or atlas based on the visibility of each camera to the viewport, a viewport-based atlas selection device. Claim 17 In claim 16, the visibility is determined based on how much the sample points in the frustum space of each camera are projected into the image plane of each camera, a viewport-based atlas selection device. Claim 18 In claim 10, the group or atlas selection unit selects a group or atlas based on a baseline between the viewport and each camera, a difference in viewpoint direction between the viewport and each camera, and visibility for each camera relative to the viewport, a viewport-based atlas selection device. Claim 19 A computer-readable recording medium storing a bitstream generated by a viewport-based atlas selection method, wherein the viewport-based atlas selection method comprises: a step of buffering and preprocessing metadata and user viewpoint information in an input information buffering and preprocessing unit; and a step of calculating a correlation value representing the correlation of user viewpoint positions for each atlas based on the metadata and user viewpoint information in a viewport-atlas correlation calculation unit. A computer-readable recording medium comprising a step of selecting a group or atlas used for MIV decoding and viewpoint image synthesis based on the correlation value in a group or atlas selection unit, wherein, when camera grouping is regular based on a camera arrangement structure, the step of selecting the group or atlas is performed by considering only the decision boundary without considering baselines with all cameras, and in the step of selecting the group or atlas, if the highest correlation value of the user viewpoint position is at a position exceeding a threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint position, the group or atlas with the highest correlation value with the user viewpoint position is selected, and in the step of selecting the group or atlas, if the highest correlation value of the user viewpoint position is at a position not exceeding the threshold from the decision boundary of the group or atlas with the highest correlation value with the previous user viewpoint position, the group with the highest correlation value with the previous user viewpoint position is selected.
Citation Information
Patent Citations
A method and apparatus for delivering a volumetric video content
EP3793199A1
Multi-view video processing method and device
KR1020220101169A
Method for switching atlas according to user's watching point and device therefor
KR1020230140418A