Generate and process image attribute pixel structure

By converting the non-rectangular image attribute pixel structure into a rectangular structure, the low efficiency problem of the existing format is solved, more efficient virtual reality image data transmission and processing is achieved, and image quality and coding efficiency are improved.

CN113330479BActive Publication Date: 2025-09-12KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080010368.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-24
Filing Date
2020-01-16
Publication Date
2025-09-12
Estimated Expiration
2040-01-16

AI Technical Summary

Technical Problem

Existing image formats such as cubemap and ERP formats are inefficient in representing 360° views, resulting in increased data requirements and inability to efficiently transmit and process image data in virtual reality applications.

Method used

An image attribute pixel structure is adopted. By converting a non-rectangular first image attribute pixel structure into a rectangular second image attribute pixel structure, equal-area projection and sinusoidal projection are used to generate a central area and a boundary area, thereby reducing overlap and mirror symmetry processing and improving coding efficiency.

Benefits of technology

It achieves more efficient scene representation, reduces data rate requirements, improves image quality and coding efficiency, and is suitable for flexible and high-performance image processing in virtual reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113330479B_ABST
    Figure CN113330479B_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus for generating or processing an image signal. A first image attribute pixel structure is a two-dimensional non-rectangular pixel structure representing the surface of a viewing sphere for a viewpoint. A second image attribute pixel structure is a two-dimensional rectangular pixel structure and is generated by a processor (305) to have a central area derived from the central area of ​​the first image attribute pixel structure and at least a first corner area derived from a first boundary area of ​​the first image attribute pixel structure. The first boundary area is an area close to one of the upper boundary and the lower boundary of the first image attribute pixel structure. The image signal is generated to include the second image attribute pixel structure, and the image signal can be processed by a receiver to restore the first image attribute pixel structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and method for generating and / or processing an image attribute pixel structure, and in particular, but not exclusively, to generating and / or processing a rectangular pixel structure representing the depth or light intensity of a scene. Background Art

[0002] The variety and range of image and video applications have increased significantly in recent years as new services and ways to utilize and consume video have been developed and introduced.

[0003] For example, an increasingly popular service is one that provides image sequences in a way that allows the viewer to actively and dynamically interact with the system to change rendering parameters. In many applications, a very attractive feature is the ability to change the viewer's effective viewing position and viewing direction, allowing the viewer to, for example, move around and "look around" the rendered scene.

[0004] Such features can, in particular, allow users to be provided with virtual reality experiences. This can allow, for example, the user to move (relatively) freely within the virtual environment and dynamically change their position and where they are looking. Typically, such virtual reality applications are based on a three-dimensional model of the scene, where the model is dynamically evaluated to provide a specific requested view. This approach is well known, for example from gaming applications for computers and consoles (e.g., first-person shooter games).

[0005] It is also desirable, particularly for virtual reality applications, that the presented images be three-dimensional. Indeed, to optimize the viewer's immersion, users often prefer to experience the presented scene as a 3D scene. Indeed, a virtual reality experience should preferably allow the user to choose his / her own position, camera viewpoint, and time of day relative to the virtual world.

[0006] Typically, virtual reality applications are inherently limited because they are based on predetermined models of scenes, and often on artificial models of the virtual world. It is often desirable to provide a virtual reality experience based on real-world capture. However, in many cases, such approaches are limited to or often require constructing a virtual model of the real world based on the real-world capture. This model is then evaluated to generate the virtual reality experience.

[0007] However, current approaches tend to be suboptimal and often have high computational or communication resource requirements and / or provide a suboptimal user experience (eg, due to reduced quality or limited freedom).

[0008] In many applications such as virtual reality, a scene can be represented by an image representation, for example, by one or more images representing a particular view pose of the scene. In some cases, such images can provide a wide-angle view of the scene and can cover, for example, a full 360° view or a full viewing sphere.

[0009] In many applications, particularly virtual reality applications, an image data stream is generated based on data representing a scene, such that the image data stream reflects the user's (virtual) position within the scene. Such an image data stream is typically generated dynamically and in real time, such that the image data stream reflects the user's movements within the virtual scene. The image data stream can be provided to a renderer, which renders an image to the user based on the image data in the image data stream. In many applications, the image data stream is provided to the renderer via a bandwidth-limited communication link. For example, the image data stream can be generated by a remote server and transmitted to a rendering device, for example, via a communication network. However, for most such applications, maintaining a reasonable data rate to allow for efficient communication is crucial.

[0010] It has been proposed to provide virtual reality experiences based on 360° video streaming, where a full 360° view of the scene is provided by a server for a given viewer position, allowing the client to generate views in different directions. In particular, one of the promising virtual reality (VR) applications is omnidirectional video (e.g., VR360 or VR180). This approach tends to incur high data rates, so the number of viewpoints required to provide a full 360° viewing sphere is typically limited to a low number.

[0011] As a specific example, virtual reality glasses have already entered the market. These glasses allow viewers to experience captured 360-degree (panoramic) videos. These 360-degree videos are typically pre-captured using camera rigs, where individual images are stitched together into a single spherical image. In some such embodiments, an image representing the complete spherical view from a given viewpoint can be generated and transmitted to a driver, which is arranged to generate an image corresponding to the user's current view for the glasses.

[0012] In many applications, a scene can be represented by a single sphere image, possibly with associated depth. An appropriate image for the current viewer pose can then be generated by selecting appropriate portions of the full image. Additionally, for sufficiently small changes in the viewer's position, the depth information can be used to generate a corresponding image using view-shifting algorithms and techniques.

[0013] Key considerations for such systems and applications are image formats and how to efficiently represent large views. For example, representing a full spherical view at a reasonably high resolution would result in high data requirements.

[0014] Since the number of viewpoints for all (or part) of the spherical information is preferably kept low (typically only a small number of or even just one viewpoint of data is provided), the degradation in quality of pose changes starting from the optimal pose, which were previously often relatively limited, can become noticeable. Particularly attractive applications of this method are applications such as immersive video, where small pose changes are supported but larger changes are not. For example, a video service can be provided that presents a scene with correct stereo cues (e.g., parallax) to the user, which gives the user the effect of rotating his head or making small head movements rather than providing the user with a roughly moving position. Such an application can provide a very advantageous user experience in many cases, but is based on the relatively small amount of data provided (compared to the case where free movement in the scene must be supported). To give a specific example, it can provide a very immersive motion experience, and the viewer may even have an experience similar to that of an audience member sitting in a seat at an event.

[0015] For systems based on such image or video data, a very important issue is how to provide an efficient representation of view data from a given viewpoint, especially how to represent the view sphere.

[0016] A common format for representing such a view sphere is called the cubemap format (see, for example, https: / / en.wikipedia.org / wiki / Cube_mapping). In this format, six images form a cube around a view position. The view sphere is then projected onto the sides of the cube, with each side providing a flat and square (partial) image. Another common format is called the ERP format, in which the view sphere surface is projected onto a rectangular image using equirectangular projection (see, for example, https: / / en.wikipedia.org / wiki / Equirectangular_projection).

[0017] However, a disadvantage of these formats is that they are often relatively inefficient and require relatively large amounts of data to represent. For example, if the view sphere is divided into pixels of uniform resolution and the same resolution is considered the minimum resolution for representation in the ERP / cubemap format, these formats would require 50% more pixels than the view sphere. This would result in a significant increase in the number of pixels required. Currently used formats are often suboptimal in terms of required data rate / capacity, complexity, etc., and this often results in suboptimal systems using these formats.

[0018] Therefore, improved methods would be advantageous. In particular, systems and / or image attribute formats that allow for improved operation, increased flexibility, improved virtual reality experience, reduced data rates, increased efficiency, facilitated distribution, reduced complexity, facilitated implementation, reduced storage requirements, improved image quality, and / or improved performance and / or operation would be advantageous. Summary of the Invention

[0019] Accordingly, the Invention seeks to preferably mitigate, alleviate or eliminate one or more of the above mentioned disadvantages singly or in any combination.

[0020] According to one aspect of the present invention, there is provided an apparatus for generating an image attribute pixel structure representing the attributes of a scene seen from a viewpoint, the apparatus comprising: a first processor providing a first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of the surface of a viewing sphere for the viewpoint; and a second processor for generating a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a central area derived from the central area of ​​the first image attribute pixel structure and at least a first corner area derived from a first boundary area of ​​the first image attribute pixel structure, the first boundary area being an area close to one of the upper and lower boundaries of the first image attribute pixel structure, the at least one corner portion not overlapping with the central area of ​​the second image attribute pixel structure; wherein the central area of ​​the first image attribute pixel structure is defined by an upper horizontal line corresponding to the upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to the lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure is closer to the periphery than at least one of the upper and lower horizontal lines; and the first corner area is close to the relative boundary of the first boundary area.

[0021] The present invention can provide improved scene representation. In many embodiments, a more efficient scene representation can be provided, for example, allowing a given quality to be achieved at a reduced data rate. The method can provide an improved rectangular image attribute pixel structure that is suitable for processing by many conventional processes, operations, and algorithms designed for rectangular images. In particular, the second image attribute pixel structure can be suitable for encoding using many known encoding algorithms, including many standardized video or image encoding algorithms.

[0022] In many embodiments, the method can provide a pixel representation of image attributes of a scene suitable for flexible, efficient, and high-performance virtual reality (VR) applications. In many embodiments, it can allow or enable VR applications with a significantly improved trade-off between image quality and data rate. In many embodiments, it can allow for improved perceived image quality and / or reduced data rate.

[0023] The method may be particularly applicable, for example, to broadcast video services that support adjustments for movement and head rotation at the receiving end.

[0024] A point (pixel) on the viewing sphere for a given viewpoint may have a value reflecting the value of an image property (typically light intensity, depth, transparency) of a scene object first encountered in a ray direction having its origin at the viewing sphere and intersecting the sphere at that point. It will be appreciated that this is in principle independent of the size of the viewing sphere, since the points are not extended. Furthermore, for a pixelated viewing sphere, the pixel value depends only on the size of the pixel, and thus for uniform resolution, only on the number of pixels into which the viewing sphere is divided, and not on the size of the viewing sphere itself.

[0025] In many embodiments, the image attribute pixel structure may be a regular pixel grid filled with a shape, wherein each pixel represents a value of an image attribute. The shape of the first image attribute pixel structure is non-rectangular, while the shape of the second image attribute pixel structure is rectangular.

[0026] The image attribute pixel structure may in particular be an image or a map, such as an intensity image, a depth map and / or a transparency map. The image attribute may be a depth attribute, a transparency attribute or an intensity attribute (eg a color channel value).

[0027] The first image attribute pixel structure may be an equal-area projection of at least part of the surface onto a plane. The equal-area projection may be a projection that maintains a ratio of areas (eg, pixel areas) between the surface of the viewing sphere and the plane onto which the surface is projected.

[0028] The first image attribute pixel structure may be a sinusoidal projection of at least part of the surface.

[0029] The second processor may be arranged to generate the second image attribute pixel structure by applying a pixel position mapping between the first image attribute pixel structure and the second image attribute pixel structure.The mapping may be different for the central region and (each) border region.

[0030] According to an optional feature of the invention, the first image attribute pixel structure has a uniform resolution for the at least part of the surface.

[0031] This may provide advantageous representation and operation in many embodiments.

[0032] According to an optional feature of the invention, the central region of the image attribute pixel structure does not overlap with the first boundary region.

[0033] This may provide advantageous representation and operation in many embodiments. In many embodiments, it may allow a particularly efficient representation and in particular reduce the required data rate for a given quality.

[0034] Each pixel of the first image attribute pixel structure may belong to only one of the central area and (one or more) boundary areas. Each pixel of the second image attribute pixel structure may belong to only one of the central area and (one or more) corner areas.

[0035] According to an optional feature of the present invention, the second processor is arranged to generate the second image attribute pixel structure to have a second corner area derived from a second boundary area of ​​the first image attribute pixel structure, the second corner area does not overlap with the center area of ​​the second image attribute pixel structure, the second boundary area is an area close to one of the upper boundary and the lower boundary, and the first boundary area and the second boundary area are on different sides of a virtual vertical line of the first image attribute pixel structure.

[0036] This can provide advantageous representation and operation in many embodiments. In many embodiments, it can allow for particularly efficient representation and, in particular, for a given quality reduction required data rate. The virtual vertical line can be any vertical line superimposed on the first image attribute pixel structure. The virtual vertical line can be any vertical line that divides the first image attribute pixel structure into a left region and a right region.

[0037] The imaginary vertical line may be a center line.

[0038] According to an optional feature of the invention, the virtual vertical line separates the first boundary region from the second boundary region, and the first boundary region and the second boundary region are mirror-symmetrical about the virtual vertical line.

[0039] This can be particularly advantageous in many embodiments.

[0040] According to an optional feature of the invention, a horizontal direction from the first border region to the second border region is opposite to a horizontal direction from the first corner region to the second corner region.

[0041] This can be particularly advantageous in many embodiments. In many embodiments, when the border region is positioned in a corner region, it can provide an improved and / or tighter fit between the center region and the border region. If the horizontal direction from the first border region to the second border region is from left to right, the horizontal direction from the first corner region to the second corner region can be from right to left, and vice versa.

[0042] The first corner region is adjacent to an opposite boundary of the first boundary region.

[0043] This can be particularly advantageous in many embodiments. In many embodiments, if the first border region is closer to the upper border of the first image attribute pixel structure (i.e., it is in the upper half), then the first corner region will be closer to the lower border of the second image attribute pixel structure (i.e., it is in the lower half), and vice versa.

[0044] In some embodiments, the horizontal pixel order for the first vertical boundary region is opposite to the horizontal pixel order for the first corner region.

[0045] According to an optional feature of the invention, the second processor is arranged to extrapolate pixels of at least one of the first corner region and the central region of the second image attribute pixel structure to an unfilled region of the second image attribute pixel structure adjacent to at least one of the first corner region and the central region.

[0046] This may provide a more efficient representation in many embodiments and may in particular improve coding efficiency when encoding the second image attribute pixel structure.

[0047] According to an optional feature of the invention, the second processor is arranged to determine the pixel values ​​of the first corner region by at least one of a shift, a translation, a mirroring and a rotation of the pixel values ​​of the first boundary region.

[0048] This can be particularly advantageous in many embodiments.

[0049] According to an optional feature of the invention, the first processor is arranged to generate the first image attribute pixel structure by distorting a rectangular image attribute pixel structure representing the at least part of the surface by means of an equirectangular projection.

[0050] This can be particularly advantageous in many embodiments.

[0051] According to one aspect of the present invention, a device for generating an output image attribute pixel structure is provided, the device comprising a receiver and a processor. The receiver is configured to receive an image signal including a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a central region derived from a central region of a first image attribute pixel structure and at least a first corner region derived from a first boundary region of the first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of a surface of a viewing sphere for a viewpoint, the first boundary region being a region proximate to one of an upper boundary and a lower boundary of the first image attribute pixel structure, the at least one corner portion not overlapping with the central region of the second image attribute pixel structure, the central region of the first image attribute pixel structure being defined by an upper horizontal line corresponding to an upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to a lower edge of the second image attribute pixel structure, the first boundary region of the first image attribute pixel structure being more peripheral than at least one of the upper horizontal line and the lower horizontal line, and the first corner region being proximate to an opposing boundary of the first boundary region. The processor is used to generate a non-rectangular output image attribute pixel structure representing at least a portion of the surface of the viewing sphere for the viewpoint, the non-rectangular output image attribute pixel structure having a central area derived from the central area of ​​the second image attribute pixel structure and a boundary area derived from the first corner area of ​​the second image attribute pixel structure, the boundary area being an area close to one of the upper and lower boundaries of the output image attribute pixel structure.

[0052] According to one aspect of the present invention, a method for generating an image attribute pixel structure representing the attributes of a scene seen from a viewpoint is provided, the method comprising: providing a first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of the surface of a viewing sphere for the viewpoint; and generating a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a central area derived from the central area of ​​the first image attribute pixel structure and at least a first corner area derived from a first boundary area of ​​the first image attribute pixel structure, the first boundary area being an area close to one of the upper and lower boundaries of the first image attribute pixel structure, the at least one corner portion not overlapping with the central area of ​​the second image attribute pixel structure; wherein the central area of ​​the first image attribute pixel structure is defined by an upper horizontal line corresponding to the upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to the lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure is closer to the periphery than at least one of the upper and lower horizontal lines; and the first corner area is close to the relative boundary of the first boundary area.

[0053] According to one aspect of the present invention, a method for generating an output image attribute pixel structure is provided, the method comprising: receiving an image signal including a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a central area derived from a central area of ​​a first image attribute pixel structure and at least a first corner area derived from a first boundary area of ​​the first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of a surface of a visual sphere for a viewpoint, and the first boundary area being an area close to one of an upper boundary and a lower boundary of the first image attribute pixel structure, the at least one corner portion not overlapping with the central area of ​​the second image attribute pixel structure, the central area of ​​the first image attribute pixel structure being formed by a pixel having the same central area as the first image attribute pixel structure. The second image attribute pixel structure is defined by an upper horizontal line corresponding to the upper edge and a lower horizontal line corresponding to the lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure is closer to the periphery than at least one of the upper horizontal line and the lower horizontal line; and the first corner area is close to the relative boundary of the first boundary area; a non-rectangular output image attribute pixel structure representing at least part of the surface of the viewing sphere for the viewpoint is generated, the non-rectangular output image attribute pixel structure having a central area derived from the central area of ​​the second image attribute pixel structure and a boundary area derived from the first corner area of ​​the second image attribute pixel structure, the boundary area being an area close to one of the upper boundary and the lower boundary of the output image attribute pixel structure.

[0054] According to one aspect of the present invention, there is provided an image signal including a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a central area derived from a central area of ​​the first image attribute pixel structure and at least a first corner area derived from a first boundary area of ​​the first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of a surface of a visual sphere for a viewpoint, and the first boundary area being an area close to one of an upper boundary and a lower boundary of the first image attribute pixel structure, the at least one corner portion not overlapping with the central area of ​​the second image attribute pixel structure, wherein the central area of ​​the first image attribute pixel structure is defined by an upper horizontal line corresponding to an upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to a lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure is closer to the periphery than at least one of the upper horizontal line and the lower horizontal line; and the first corner area is close to the relative boundary of the first boundary area.

[0055] These and other aspects, features and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0057] Figure 1 An example of an arrangement for providing a virtual reality experience is illustrated;

[0058] Figure 2 An example of an ERP projection for a spherical image of the visual sphere is illustrated;

[0059] Figure 3 illustrates examples of elements of an apparatus according to some embodiments of the present invention; and

[0060] Figure 4 An example of a sinusoidal projection of a spherical image onto a viewing sphere is illustrated;

[0061] Figure 5 illustrates an example of a mapping from a sinusoidal projection of a spherical image for a viewing sphere to a rectangular image representing the viewing sphere, according to some embodiments of the present invention;

[0062] Figure 6 illustrates an example of a mapping from a sinusoidal projection of a spherical image for a viewing sphere to a rectangular image representing the viewing sphere, according to some embodiments of the present invention;

[0063] Figure 7 illustrates an example of a rectangular image representing a viewing sphere according to some embodiments of the present invention;

[0064] Figure 8 illustrates an example of a rectangular image representing a viewing sphere according to some embodiments of the present invention;

[0065] Figure 9 illustrates examples of rectangular images representing two viewing spheres according to some embodiments of the present invention; and

[0066] Figure 10 Examples of elements of an apparatus according to some embodiments of the present invention are illustrated. DETAILED DESCRIPTION

[0067] Virtual experiences that allow users to walk around in a virtual world are becoming increasingly popular, and services are being developed to meet this demand. However, providing efficient virtual reality services is very challenging, especially if the experience is based on a capture of the real-world environment rather than a fully virtual generated artificial world.

[0068] In many virtual reality applications, viewer gesture input is determined to reflect the pose of a virtual viewer in the scene. The virtual reality device / system / application then generates one or more images corresponding to the view and viewport of the scene for the viewer corresponding to the viewer gesture.

[0069] Typically, virtual reality applications generate three-dimensional output in the form of separate view images for the left and right eyes. These outputs can then be presented to the user via suitable means (e.g., separate left-eye and right-eye displays, typically of a VR headset). In other embodiments, the images can be presented, for example, on an autostereoscopic display (in which case a large number of view images can be generated for the viewer's posture), or indeed in some embodiments, only a single two-dimensional image can be generated (e.g., using a conventional two-dimensional display).

[0070] Viewer gesture input can be determined in different ways in different applications. In many embodiments, the user's body movements can be tracked directly. For example, a camera surveying the user's area can detect and track the user's head (or even eyes). In many embodiments, the user can wear a VR headset that can be tracked externally and / or internally. For example, the headset can include accelerometers and gyroscopes that provide information about the movement and rotation of the headset and, therefore, the head. In some examples, the VR headset can transmit a signal or include a (e.g., visual) identifier that enables external sensors to determine the movement of the VR headset.

[0071] In some systems, the viewer gesture may be provided manually, such as by the user manually controlling a joystick or similar manual input. For example, the user may manually move the virtual viewer around the scene by controlling a first analog joystick with one hand, and manually control the direction the virtual viewer is looking by manually moving a second analog joystick with the other hand.

[0072] In some applications, a combination of manual and automatic methods can be used to generate the input viewer pose. For example, a head-mounted device can track the orientation of the head and the viewer's movement / position in the scene can be controlled by the user using a joystick.

[0073] The generation of images is based on a suitable representation of the virtual world / environment / scene. In some applications, a complete three-dimensional model can be provided for the scene, and the view of the scene seen from a specific viewer posture can be determined by evaluating the model. In other systems, the scene can be represented by image data corresponding to views captured from different capture postures. For example, for one or more capture postures, a complete spherical image can be stored together with three-dimensional (depth data). In such an approach, view images for postures other than the (one or more) capture postures can be generated by three-dimensional image processing (for example, in particular using a view shift algorithm). In a system where the scene is described / referenced by view data stored for discrete viewpoints / positions / postures, these discrete viewpoints / positions / postures can also be referred to as anchor viewpoints / positions / postures. Typically, when a real-world environment is captured by capturing images from different viewpoints / positions / postures, these capture viewpoints / positions / postures are also anchor viewpoints / positions / postures.

[0074] Thus, a typical VR application provides (at least) an image corresponding to a viewport of the scene for the current viewer pose, where the image is dynamically updated to reflect changes in the viewer pose and the image is generated based on data representing the virtual scene / environment / world.

[0075] In this field, the terms placement and pose are used as general terms for position and / or direction / orientation. For example, the combination of the position and direction / orientation of an object, camera, head or view can be called a pose or placement. Therefore, a placement or pose indication can include six values / components / degrees of freedom, wherein each value / component generally describes an independent attribute of the position / position or orientation / direction of the corresponding object. Of course, in many cases, placement or pose may be considered or represented with fewer components, such as when one or more components are considered to be fixed or unrelated (for example, when all objects are considered to be at the same height and have a horizontal orientation, four components can provide a complete representation of the object's pose). In the following, the term pose is used to refer to a position and / or orientation that can be represented by one to six values ​​(corresponding to the maximum possible degrees of freedom).

[0076] Many VR applications are based on gestures with the maximum degrees of freedom, i.e., three degrees of freedom each in position and orientation for a total of six degrees of freedom. A gesture can therefore be represented by a set or vector of six values ​​representing the six degrees of freedom, and the gesture vector can therefore provide a three-dimensional position and / or three-dimensional direction indication. However, it should be understood that in other embodiments, a gesture can be represented by fewer values.

[0077] The posture may be at least one of an orientation and a position. The posture value may indicate at least one of an orientation value and a position value.

[0078] Systems or entities based on providing the most freedom to the viewer are often referred to as having 6 degrees of freedom (6DoF).Many systems and entities only provide orientation or position, and these are often referred to as having 3 degrees of freedom (3DoF).

[0079] In some systems, VR applications can be provided locally to viewers, for example, by a standalone device that does not utilize (or even access) any remote VR data or processing. For example, a device such as a game console may include storage for storing scene data, an input for receiving / generating viewer gestures, and a processor for generating corresponding images based on the scene data.

[0080] In other systems, VR applications can be implemented and executed remotely from the viewer. For example, a device local to the user can detect / receive movement / posture data, which is transmitted to a remote device that processes the data to generate the viewer's posture. The remote device can then generate a view image appropriate for the viewer's posture based on scene data describing the scene data. The view image is then transmitted to a device local to the viewer, where it is presented. For example, the remote device can directly generate a video stream (typically a stereo / 3D video stream) that is directly presented by the local device. Therefore, in such an example, the local device may not perform any VR processing other than transmitting the movement data and presenting the received video data.

[0081] In many systems, functionality may be distributed across local and remote devices. For example, a local device may process received input and sensor data to generate a viewer pose that is continuously transmitted to a remote VR device. The remote VR device may then generate corresponding view images and transmit these view images to the local device for presentation. In other systems, the remote VR device may not generate view images directly, but may select relevant scene data and transmit it to the local device, which may then generate the presented view images. For example, the remote VR device may identify the closest capture point and extract corresponding scene data (e.g., a spherical image and depth data from the capture point) and transmit it to the local device. The local device may then process the received scene data to generate an image for a specific current view pose. The view pose will typically correspond to a head pose, and references to the view pose may typically be considered equivalent to references to the head pose.

[0082] In many applications, particularly for broadcast services, a source may transmit scene data in the form of an image (including video) representation of the scene that is independent of the viewer's pose. For example, an image representation of a single view sphere for a single capture position may be transmitted to multiple clients. Individual clients may then locally synthesize a view image corresponding to the current viewer's pose.

[0083] Specific applications of particular interest are those that support a limited amount of movement, such that the rendered view is updated to follow small movements and rotations corresponding to a largely static viewer who performs only small head movements and rotations. For example, a seated viewer can turn their head and move their head slightly, with the rendered view / image being adjusted to follow these posture changes. Such an approach can provide a highly immersive experience, such as with video. For example, a viewer watching a sporting event might feel as if they are present at a specific location within the arena.

[0084] Such limited degrees of freedom applications have the following advantages: they can provide an improved experience while not requiring accurate representations of the scene from many different positions, significantly reducing capture requirements. Similarly, the amount of data required to be provided to the renderer can be greatly reduced. In fact, in many scenarios, only the image for a single viewpoint and often also depth data are required to be provided to the local renderer, which can generate the desired view. To support head rotation, it is often desirable to have a large area of ​​view from the viewpoint represented by the provided data, and preferably the entire surface of the view sphere centered on the viewpoint is covered by the provided image and depth data.

[0085] This approach may be particularly well-suited for applications that need to transmit data from a source to a destination over a bandwidth-constrained communication channel, such as broadcast or client-server applications.

[0086] Figure 1 An example of a VR system is shown in which a remote VR client device 101 communicates with a VR server 103, for example via a network 105 (e.g., the Internet). The server 103 may be arranged to support a potentially large number of client devices 101 simultaneously.

[0087] The VR server 103 may support a broadcast experience, for example, by transmitting image data and depth for a particular viewpoint, with the client device then being arranged to process this information to locally synthesize a view image corresponding to the current pose.

[0088] Therefore, many applications are based on transmitting image information corresponding to view positions that are much larger than conventional small viewports, which provide relatively narrow left-eye and right-eye images. In particular, in many applications, it is desirable to transmit image attribute information (e.g., light intensity and depth) for the entire view sphere for one or more view / capture positions. For example, in VR360 video applications, light intensity and depth values ​​for the entire view sphere are transmitted. However, a key issue for such applications is how to represent this information so that particularly efficient communication can be achieved.

[0089] For example, it is desirable to be able to use existing algorithms and functions for encoding and formatting image attribute information. However, such functions are often designed almost exclusively for flat rectangular image formats, and the surface of a three-dimensional surface does not inherently correspond to a two-dimensional rectangle. To address this issue, many approaches use a cubemap format, in which a cube is positioned around a viewing sphere, and the surfaces of the viewing sphere are then projected onto the cube's square edges. Each of these edges is planar and can therefore be processed accordingly using conventional techniques. However, one disadvantage is that if the resolution of the cubemap is the same as the resolution of the viewing sphere at the point where the cubemap touches the viewing sphere (so as not to incur any resolution loss), then the cubemap requires a significantly larger number of pixels than the viewing sphere (the projection of the viewing sphere onto the outer regions of each edge causes each viewing sphere pixel to be projected onto an area larger than a single pixel (and in particular, an area that may correspond to a relatively large number of pixels)). It can be shown that a cubemap representation requires approximately 50% more pixels of a given uniform size than a spherical representation would require for a given uniform size.

[0090] Another commonly used format is to project the surface of the scene onto a two-dimensional rectangle using equirectangular projection (ERP). Figure 2 An example of such an image is shown. It can be seen that the deformation caused by the projection also significantly increases the resulting projected area of ​​some regions relative to others (in particular, the increase in projected area towards the vertical edges reflects that a single point above (or below) the view pose is stretched across the full width of the rectangle). As a result, the required number of (constant-sized) pixels required without reducing the center resolution will increase significantly. It can be shown that the ERP representation requires about 50% more pixels of a given uniform size than the spherical representation.

[0091] This increased pixel count leads to increased processing complexity and higher data requirements. In particular, higher data rates may be required to transmit image information.

[0092] Figure 3 Elements of an apparatus by which a representation of (at least part of) the surface of a viewing sphere may be generated are illustrated.

[0093] The apparatus comprises a first processor 301 arranged to provide a first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional (planar, Euclidean) non-rectangular pixel structure representing at least part of a surface of a viewing sphere for a viewpoint / view position. In the following example, the first processor 301 is arranged to process a (light intensity) image and a depth map, and the first image attribute pixel structure may be considered to correspond to either the image or the depth map, or indeed both simultaneously (wherein a pixel value is a combination of an image pixel value and a depth map pixel value).

[0094] In this example, a first processor 301 is coupled to a source 303 for a first image attribute pixel structure, and in particular, the source can provide an input image and a depth map. In particular, the source 303 can be a local memory storing image information, or the source 303 can be, for example, a suitable capture unit, such as a full spherical camera and / or a depth sensor.

[0095] The viewing sphere for a viewpoint is a (nominal) sphere surrounding the viewpoint, where each point of the surface represents an image property value for the scene in a direction from the viewpoint through the point on the surface. For intensity image property values, the value at a point on the surface corresponds to the intensity of light reaching the viewpoint from the direction of that point. Correspondingly, for depth or range image property values, the depth value at a given point on the surface of the viewing sphere corresponds to the distance from the viewpoint to the first object of the scene in a direction from the viewpoint to (through) the point on the surface.

[0096] The first image attribute pixel structure represents an image attribute that represents an attribute of a scene that can be used by an image rendering process to generate a particular view image. Thus, the image attribute can be an attribute that can support image generation / composition functionality for generating images of the scene (for one or more viewports, e.g., corresponding to different view poses). Specifically, the image attribute can be at least one of the following attributes: a light intensity attribute, a depth attribute, or a transparency attribute.

[0097] In some embodiments, the image attribute may be a composite attribute, for example, the image attribute may include multiple color channel light intensity values ​​(e.g., a red value, a green value, and a blue value), and may also include a depth value. For example, each pixel of the first image attribute pixel structure may include multiple values ​​for each value, or the pixel value may be a multi-component vector value. Equivalently, the first processor 301 may be considered to provide multiple single-value image attribute pixel structures, wherein each of these single-value image attribute pixel structures is processed as described below.

[0098] The following description will focus on the properties of light intensity and depth. Therefore, for the sake of brevity and clarity, the (one or more) image attribute pixel structures are also referred to by the more common terms of (light intensity) image and depth map respectively.

[0099] In particular, the image attribute pixel structure can be a planar area or region divided into a plurality of pixels. Each pixel includes one or more values ​​that indicate the value of the image attribute for the area covered by the pixel. Typically, the pixels are all of the same size, i.e., the resolution is uniform. Typically, the pixels are square, or at least rectangular, and are arranged in an equidistant grid. Therefore, a conventional image or depth map is an example of an image attribute pixel structure.

[0100] However, the first image attribute pixel structure is not a rectangular image attribute pixel structure, but a two-dimensional non-rectangular pixel structure. The first image attribute pixel structure also represents at least a portion of the surface of the viewing sphere, and typically the entire surface of the viewing sphere. Because the surface of the viewing sphere has a three-dimensional curved property, the corresponding planar representation is generally not rectangular.

[0101] In particular, the surface of a sphere can be considered to be divided into a given number of pixels of equal area, where the area covered by the pixels is usually (substantially) square. If these pixels are instead rearranged on a plane, the resulting coverage area will not be rectangular or square. In particular, Figure 4 The resulting square pixel area is shown.

[0102] In this case, the first image attribute pixel structure may in particular be a sinusoidal projection of the surface of the viewing sphere. Figure 4 The area covered by the first image attribute pixel structure for the complete surface is shown. As shown in the figure, the surface of the viewing sphere is represented as a region where the width / horizontal extension is a sinusoidal function of the vertical position, where the vertical position is represented by a value ranging from 0 to π (180°), where the central vertical position corresponds to π / 2 (90°). In this example, the vertical positions of 0 and π (180°) therefore correspond to the directions directly downward and upward from the viewpoint, while the vertical position of π / 2 (90°) corresponds to the horizontal direction as seen from the viewpoint.

[0103] It should be understood that in some embodiments, the first image attribute pixel structure may represent only a portion of the surface of the viewing sphere. For example, in some embodiments, the first image attribute pixel structure may represent only a hemisphere (e.g., the upper hemisphere (e.g., corresponding to a camera positioned at ground level and capturing only scenes above the ground level)) or only a hemisphere in a given direction (e.g., only a hemisphere viewed in one general direction for a user). In such examples, the first image attribute pixel structure will also be non-rectangular, but will not directly correspond to Figure 4In some embodiments, the first image attribute pixel structure may still be a sinusoidal projection, but may only be a projection of a portion of the surface. For example, the resulting first image attribute pixel structure for a hemisphere may only correspond to Figure 4 The upper half or left (or right) half.

[0104] The first image attribute pixel structure is therefore a non-rectangular structure and is therefore not suitable for processing in many existing processes, including, for example, processing in image or video encoders. The first image attribute pixel structure also tends to be inconsistent with many existing standards and formats based on rectangular image representations. Therefore, it is desirable to convert the non-rectangular first image attribute pixel structure into a rectangular second image attribute pixel structure. As described above, conventionally, this is often accomplished by projecting the surface of a sphere onto a rectangle or onto the side of a cube (cube diagram representation) using ERP.

[0105] However, in contrast, Figure 3 The apparatus comprises a second processor 305, which is arranged to generate a second image attribute pixel structure as a two-dimensional rectangular pixel structure using a method that results in a more efficient structure and in particular a method that maintains the resolution without significantly increasing the number of required pixels.

[0106] Accordingly, the output of the second processor 305 is a rectangular image structure, and in particular may be a rectangular image and / or depth map. This second image attribute pixel structure may be fed to an output generator 307, which is arranged to generate an image signal in the form of an output data stream that can be transmitted to a remote device. In particular, the output generator 307 may be arranged to encode the second image attribute pixel structure using a technique designed for rectangular images and include the encoded data in the output data stream. For example, image or video encoding may be applied to the rectangular image provided by the second processor 305 in order to generate a corresponding encoded video data stream that can be transmitted to a remote client.

[0107] In particular, the second processor 305 is arranged to determine different areas in the first image attribute pixel structure and position these areas differently and separately in a rectangular area of ​​the second image attribute pixel structure. In particular, the second processor 305 is arranged to derive the central area of ​​the second image attribute pixel structure based on the central area of ​​the first image attribute pixel structure. The second processor 305 may also derive one or more corner areas based on one or more boundary areas in the vertical direction and in particular one or more boundary areas close to the upper edge / border or lower edge / border of the first image attribute pixel structure. Therefore, in this method, the second processor 305 may occupy the position of the central area of ​​the rectangular output image based on the image data of the central area of ​​the non-rectangular image, and occupy the position of (one or more) corner areas based on the image data of (one or more) outer areas of the input image in the vertical direction (close to the top or bottom of the input image).

[0108] You can refer to Figure 5-7 to illustrate an example of a method performed by the second processor 305, wherein, Figure 6 and Figure 7 Shown is the corresponding Figure 3 A specific image (where the Figure 5 An example of the principle of

[0109] In this example, the first image attribute pixel structure in the form of an image represents the surface of the viewing sphere using a sinusoidal projection and therefore corresponds to a flat region 501 corresponding to a shape formed by half a sine wave cycle and its mirror image, as shown in the figure. In this example, four boundary regions are determined: an upper left region p1, an upper right region p2, a lower left region p3, and a lower right region p4.

[0110] The second processor 305 then generates a second image attribute pixel structure corresponding to a rectangular image. This image therefore corresponds to the rectangular area 503. This image is generated by maintaining the center portion of the input image and shifting the boundary areas p1-p4 diagonally to opposite corners. In this example, the input image has dimensions W×H0, while the height of the output image may be reduced by H.

[0111] Figure 6 Illustrate how a pixel with the same pixel size as the image can be first generated by copying the boundary region to the corner region of the intermediate image. Figure 4 An example of a rectangular intermediate image of the same height and width as the input image. This image may be significantly larger than the input image because it includes many redundant pixels. However, the output image can then be generated by vertically cropping it, thereby reducing the height and removing the redundant pixels. In particular, the height can be reduced to a level that minimizes the number of redundant pixels. In fact, in many embodiments, the height can be reduced by cropping to exclude redundant pixels.

[0112] This method can provide a very efficient view sphere representation of a scene by means of a rectangular image. The method is based on the inventors' recognition that the properties of the projection of the view sphere are suitable for being divided into different regions that can be closely fitted into rectangular areas. In fact, from Figure 5 、 Figure 6 and Figure 7 As can be seen from the example, it is possible to determine the boundary area that fits closely to the boundary area. Figure 7 As can be seen from the example of , it is possible to generate a Figure 7 In fact, it can be shown that in this example, a rectangular representation of the surface of the viewing sphere can be achieved with only about 5% more pixels, without any loss of resolution.

[0113] Therefore, a more efficient representation than ERP or cube graph representation can be achieved by the described method.

[0114] In the above examples, the image attribute pixel structure has been a (light intensity) image, but it will be appreciated that these methods can also be applied to other attributes, such as a depth map or a transparency map. For example, the image of the above example can be supplemented with a depth map and / or a transparency map, which can provide, for example, a depth value and a transparency value for each pixel of the image, respectively. These maps can then be processed in the same manner as described for the light intensity image to obtain a more suitable square image or encoded using, for example, conventional techniques.

[0115] In the above example, the second processor 305 is arranged to determine four boundary regions (p1, p2, p3, p4) in the first image attribute pixel structure, wherein the boundary structure is close to the upper boundary of the first image attribute pixel structure or the lower boundary of the image attribute pixel structure. In the example, the boundary regions are therefore the upper boundary region and the lower boundary region and are in particular (continuous) regions, for which part of its boundary is also the boundary of the first image attribute pixel structure itself.

[0116] In this example, the second processor 305 identifies the center area of ​​the second image attribute pixel structure and uses the center area of ​​the first image attribute pixel structure to occupy the center area's position. It also identifies the four corner areas of the second image attribute pixel structure and uses the four border areas of the first image attribute pixel structure to occupy the positions of the four corner areas. Thus, the four identified border areas of the first image attribute pixel structure can be efficiently moved to the corner areas of the second image attribute pixel structure.

[0117] In this example, the center region of the first image attribute pixel structure, which is used to occupy the position of the center region, is defined by upper and lower horizontal lines corresponding to the upper and lower edges of the second image attribute pixel structure. The boundary regions of the first image attribute pixel structure are more peripheral than these lines, that is, they are located above and below the dividing horizontal lines, respectively. Therefore, a position mapping is applied from the pixel positions of the center region of the first image attribute pixel structure to the pixel positions in the center region of the second image attribute pixel structure. If the same pixel position mapping is applied to the boundary regions, the positions will fall outside the second image attribute pixel structure.

[0118] In particular, the second image attribute pixel structure is a rectangular structure having an upper / top (horizontal) edge and a lower / bottom (horizontal) edge. The central area of ​​the second image attribute pixel structure can be defined by these edges, and the central area in the second image attribute pixel structure can be spliced ​​to the edge of the second image attribute pixel structure.

[0119] The upper edge of the second image attribute pixel structure may correspond to the upper / top horizontal line in the first image attribute pixel structure, and the lower edge of the second image attribute pixel structure may correspond to the upper / top horizontal line in the first image attribute pixel structure.

[0120] The central area of ​​the first image attribute pixel structure may be selected as a portion of the first image attribute pixel structure that falls between the two horizontal lines.

[0121] One, a plurality of and usually all of the boundary regions of the first image attribute pixel structure are regions that are more peripheral than at least one of the horizontal lines. Thus, the boundary region may be above the upper horizontal line or below the lower horizontal line.

[0122] In some embodiments, the described method may be applied only to the top or bottom of the first image attribute pixel structure. Thus, in some embodiments, only one of the upper and lower horizontal lines may be considered, or equivalently, one of the upper and lower horizontal lines may be considered to correspond to an edge of the first image attribute pixel structure. However, in most embodiments, the method will be applied symmetrically to both the top and bottom portions of the first image attribute pixel structure. Furthermore, in most embodiments, the method will be applied symmetrically to both the top and bottom portions of the first image attribute pixel structure.

[0123] In many embodiments, multiple boundary regions may be determined and assigned to corner regions in the second image attribute pixel structure. Each of these boundary regions may be more peripheral than the horizontal line, ie, may be above the upper horizontal line or below the lower horizontal line.

[0124] In many embodiments, the boundary region may include all regions of the first image attribute pixel structure that are more peripheral / remote / outer / outer than the horizontal line. Thus, in some embodiments, all pixels above the upper horizontal line and below the lower horizontal line may be included in the boundary region.

[0125] The method may in particular allow the generation of a second image attribute pixel structure that is smaller than the rectangular structure that would contain the first image attribute pixel structure. The number of rows in the second image attribute pixel structure may be lower than the number of rows in the first image attribute pixel structure, and may typically be at least 5%, 10% or 20% lower.

[0126] The height (vertical extension) of the second image attribute pixel structure may be significantly lower than the height (vertical extension) of the first image attribute pixel structure, and may typically be at least 5%, 10% or 20% lower.

[0127] Additionally, this can often be achieved while maintaining the number of columns / width / horizontal extent, thus significantly reducing the number of pixels required for a rectangular image.

[0128] In many embodiments, each boundary region may be a contiguous region comprising a relatively large number of pixels. In many embodiments, at least one boundary region may comprise no fewer than 1,000, or even 5,000, pixels. In many embodiments, a boundary region may comprise no fewer than 5% or 10% of the total number of pixels in the first image attribute pixel structure.

[0129] In some embodiments, the encoding of the second image attribute pixel structure may use an encoding algorithm based on image blocks (e.g., macroblocks known from, for example, MPEG encoding). In such embodiments, each boundary region may include an integer number of macroblocks, i.e., each boundary region may not include any portion of a coded block. In addition, in many embodiments, each boundary region may include multiple coded blocks.

[0130] Each boundary region can be reassigned to the boundary region as a block, ie, the relative positions between pixels are unchanged.

[0131] In many embodiments, each portion may be a portion that extends to a corner of the second image attribute pixel structure. A border region may be included in the second image attribute pixel structure such that it abuts a corner of the second image attribute pixel structure. A border region may be included in the second image attribute pixel structure such that it has an edge in common with an edge of the second image attribute pixel structure, and may have two edges in common with the second image attribute pixel structure.

[0132] In many embodiments, the boundary region may be included in the second image attribute pixel structure such that there is a certain distance between the boundary region and the center region of the second image attribute pixel structure, that is, there may be a guard band between them. Such a distance may be, for example, not less than 1, 2, 5, 10, or 50 pixels.

[0133] In many embodiments, a border region may be included in the second image attribute pixel structure such that the vertical positions of pixels in the border region in the second image attribute pixel structure are different from the vertical positions in the first image attribute pixel structure. Specifically, the vertical positions of pixels in the border region may be positioned more peripherally than the vertical positions of the upper vertical line and / or the lower vertical line, and more centrally / farther from the periphery than the vertical positions of the upper vertical line and / or the lower vertical line.

[0134] In a specific example, each boundary region is moved to a diagonal corner region, i.e., the upper left boundary region is moved to the lower right corner region; the upper right boundary region is moved to the lower left corner region; the lower left boundary region is moved to the upper right corner region; and the lower right boundary region is moved to the upper left corner region.

[0135] Therefore, in this example, the horizontal relationship between the two boundary regions in the first image attribute pixel structure and the two corner regions to which they are mapped in the second image attribute pixel structure is reversed. Thus, the first boundary region, located to the left of the second boundary region, is moved to the first corner region, located to the right of the second corner region to which the second boundary region is moved. In this method, the horizontal direction from the first boundary region to the second boundary region is opposite to the horizontal direction from the first corner region to the second corner region.

[0136] Similarly, in this embodiment, the vertical relationship between the two boundary regions in the first image attribute pixel structure and the two corner regions to which they are mapped in the second image attribute pixel structure is reversed. Therefore, the first boundary region located above the second boundary region (the first region is the upper boundary region, and the second region is the lower boundary region) will be moved to the first corner region located below the second corner region to which the second boundary region is moved (the first corner region is the lower corner region, and the second corner region is the upper corner region). In this method, the vertical direction from the first boundary region to the second boundary region is opposite to the vertical direction from the first corner region to the second corner region.

[0137] In many embodiments, such as those described above, the first corner area and the second corner area may include pixel data derived from a first boundary area and a second boundary area in the first image attribute pixel structure, respectively, wherein the two boundary areas are close to the same upper boundary or lower boundary (i.e., both boundary areas are at the upper boundary or both are at the lower boundary). For example, the two boundary areas may be areas p1 and p2 (or p3 and p4) in the accompanying drawings. The two boundary areas are horizontally shifted relative to each other, and in particular, the entirety of one of the two boundary areas may be completely to the right of the entirety of the other area. Therefore, the two boundary areas may be located on different sides of a virtual vertical line, in particular, the virtual vertical line may be the center line of the first image attribute pixel structure. In the example in the accompanying drawings, the virtual vertical line is the center line (i.e., p1 and p2 are located on different sides of the vertical center line, as are p3 and p4).

[0138] exist Figure 5-7 In the specific example of , the virtual vertical line is the center line of the first image attribute pixel structure. In addition, in this example, the virtual vertical line is the line separating the first boundary area from the second boundary area. In fact, the first boundary area and the second boundary area together form a continuous area, and the continuous area is subdivided into the first boundary area and the second boundary area by the virtual vertical line. In addition, in the specific example, the first boundary area and the second boundary area are mirror-symmetrical around the virtual vertical line. In this example, the upper boundary area (and / or the lower boundary area) is identified as the area above (or below) a given vertical coordinate (and therefore above (or below) a given horizontal line corresponding to the horizontal line). Therefore, the area is divided into two corresponding areas by the vertical center line. In the specific example, two identical but mirror-symmetrical boundary areas are obtained in this way.

[0139] In addition, in a specific example, both an upper boundary region and a lower boundary region are identified, and these regions are mirror-symmetrical with respect to the horizontal center line. In particular, in this example, four boundary regions are found that are mirror-symmetrical in pairs around the horizontal center line and the vertical center line.

[0140] In many embodiments, dividing the first image attribute pixel structure into such horizontally and / or vertically shifted and separated regions can provide an efficient and advantageous partitioning that allows for a relatively low complexity but efficient reorganization of the second image attribute pixel structure while reducing the amount of unused portions of the second image attribute pixel structure, thereby reducing overhead / waste. As described above, in many embodiments, the two boundary regions can be linked to the corner region such that the horizontal order is reversed and / or the vertical order is reversed, but it will be understood that this is not required and some embodiments may not employ (one or more) such reversals.

[0141] In particular, the approach allows for the generation of efficient rectangular and planar image structures, which allows for maintaining uniform resolution without requiring a large overhead for rectangular representation.

[0142] In many embodiments, the first image attribute pixels have a uniform resolution for the surface (or portion of the surface) of the viewing sphere. Thus, the resolution of the viewing sphere is the same in all directions, and all directions are represented with the same quality. Transforming the first image attribute pixel structure into a rectangular second image attribute pixel structure can be achieved by a direct rearrangement of the pixels, so that the pixel resolution remains unchanged. Compared to ERP or cube map formats, the described method generates a rectangular image in which the resolution of the viewing sphere is unchanged and thus also represents a uniform resolution of the viewing sphere. Furthermore, this is achieved with only a small overhead and an increased number of pixels.

[0143] A particular advantage of the described method is that it provides a method in which the boundary region is closely fitted within the selected boundary region. The shape and contour of the boundary region closely match the corner regions remaining after the central portion is copied into the second image attribute pixel structure. Furthermore, in a specific example, the boundary region is assigned to the corner regions so that the shapes match each other without introducing any additional operations (specifically, only requiring translation / shifting).

[0144] In many embodiments, pixels of the first image attribute pixel structure can be directly mapped to pixels of the second image attribute pixel structure, and in particular, each pixel in the second image attribute pixel structure can be a copy of a pixel in the first image attribute pixel structure. The processing of the second processor 305 can therefore be regarded as a mapping of pixels (pixel positions) in the first image attribute pixel structure to pixels (pixel positions) in the second image attribute pixel structure. However, it should be understood that in some embodiments, the second processor 305 may also include some processing of the pixel values, for example, the processing may include brightness adjustment, depth adjustment, filtering, etc.

[0145] In the illustrated example, the positions of the corner regions are occupied by direct shifting / offsetting / translating of the boundary regions.The internal spatial relationships between pixels in each of the boundary regions are maintained in the corner regions.

[0146] However, in other embodiments, the second processor 305 may alternatively or additionally be arranged to include, for example, a mirror image and / or rotation of the border region. In particular, this may ensure a closer fit between the shape of the border region and the shape of the corner region in which the border region is positioned.

[0147] This can be used, for example, to apply different mappings between border regions and corner regions. For example, instead of linking border regions to diagonally opposite corner regions (i.e., the upper left border region to the lower right corner region), a given border region can be mapped to a near-side corner region and a rotation (and / or mirroring) can be used to fit the shape of the border region to the shape of the corner region. For example, in the example in the accompanying drawings, the upper left border region p1 can be rotated 180° and shifted to the upper left corner region. Thus, a rotation can be performed so that the central portion of the border region becomes the lateral portion.

[0148] Such an approach of using translation and / or mirroring rather than just translation can be particularly advantageous in many embodiments in which only portions of the viewing sphere are represented by the first image attribute pixel structure and the second image attribute pixel structure. For example, in an example in which only the upper half of the viewing sphere is represented (corresponding only to the upper half of the image of the exemplary figure), the two boundary regions p1 and p2 can be fitted into two corner regions. For example, p1 and p2 can be fitted into the upper left corner region and the upper right corner region, respectively, after being rotated 180°, or into the upper right corner region and the upper left corner region, respectively, after being mirrored about the horizontal line.

[0149] In some embodiments, the first processor 201 is arranged to receive a representation of the surface of the viewing sphere as an image attribute pixel structure representing the surface by an equirectangular projection. For example, the first processor 201 may receive Figure 2 The representation shown.

[0150] In such an embodiment, the first processor 201 may be arranged to warp such a rectangular image attribute pixel structure into a non-rectangular image attribute pixel structure, which can then be processed into the first image attribute pixel structure as described above. In particular, in such an embodiment, the first processor 201 may be arranged to transform the received image attribute pixel structure from an equirectangular projection into an image attribute pixel structure corresponding to a sinusoidal projection.

[0151] The first processor 201 may for example be arranged to achieve this by performing a coordinate translation according to a cosine warp. An example implemented in matlab may be as follows:

[0152] W0=4000

[0153] H0=2000

[0154] W=W0*4

[0155] W2=W / 2

[0156] H=H0

[0157] H2=H / 2

[0158] AA = imread('in1.png');

[0159] BB=zeros(size(AA));

[0160] for x=1:W

[0161] for y=1:H

[0162] sc=abs(cos(3.14 / 2*(y-H2) / H2));

[0163] %sc=0;

[0164] x1=x-W2;

[0165] i=floor(x1*sc)+W2;

[0166] BB(y,x,1)=AA(y,i,1);

[0167] BB(y,x,2)=AA(y,i,2);

[0168] BB(y,x,3)=AA(y,i,3);

[0169] end

[0170] end

[0171] imwrite(uint8([BB]),['out.png']);

[0172] In the previous example, a second image attribute pixel structure was generated, which included a central area and one or more corner areas generated based on the central area and one or more boundary areas of the first image attribute pixel structure, respectively. In particular, the method can utilize the geometric properties of the first image attribute pixel structure to generate the central area and the boundary areas, so that the corner areas generated by filling the central area of ​​the second image attribute pixel structure from the central area of ​​the first image attribute pixel structure have geometric properties (particularly shape) that are relatively closely matched to the geometric properties (particularly shape) of the boundary areas. This allows an arrangement in which the entire first image attribute pixel structure is closely located within non-overlapping central and corner areas with only small gaps. Thus, efficient representation is achieved, with only a small number of pixels of the second image attribute pixel structure not representing pixels of the first image attribute pixel structure.

[0173] exist Figure 7This is illustrated in FIG by the black pixels between the corner regions and the center region. It can be seen that the method is able to exploit the geometry of the first image attribute pixel structure to ensure an efficient rectangular representation, while the non-rectangular representation of the first image attribute pixel structure has only a small number of extra pixels in the rectangular representation. Figure 4 In the example of a sinusoidal projection (shown in ), increasing the number of pixels by only 5% can produce a rectangular representation.

[0174] Thus, a second image attribute pixel structure is generated with one or more unfilled areas, however, these unfilled areas are kept within a relatively small area. The small overhead (e.g., compared to ERP or cubemap representations) allows the pixel count of the image attribute pixel structure to be reduced, which can significantly reduce the encoding data rate.

[0175] In some embodiments, the second processor 305 may be further configured to perform filling in of one or more of the unfilled regions. In particular, the filling in may be performed by generating pixel values ​​for pixels of the unfilled region(s) based on pixel values ​​of nearside pixels in the regions that have already been occupied, and in particular based on pixel values ​​of the central region and pixel values ​​of pixels of the nearest corner regions that have not yet been occupied according to the first image attribute pixel structure.

[0176] In many embodiments, one or more pixel values ​​generated according to the first image attribute pixel structure may be extrapolated into the unfilled region(s). It will be appreciated that a variety of filling techniques are known from the deocclusion process as part of the view synthesis technique, and any such suitable algorithm can be used.

[0177] In some embodiments, the filling can be performed by generating an intermediate image attribute pixel structure based on the first image attribute pixel structure, wherein the first image attribute pixel structure is extrapolated to the surrounding area. In this example, rather than simply moving the boundary area of ​​the first image attribute pixel structure to the corner area of ​​the second image attribute pixel structure (which would directly result in an unfilled area when the shape of the boundary area does not match the shape of the corner area), the area of ​​the intermediate image corresponding to the unfilled area is also moved, thereby filling the unfilled area.

[0178] Figure 8 An example of an intermediate image generated by extrapolating the first image attribute pixel structure into a rectangular image is shown. The second processor 305 can then generate a second image attribute pixel structure by copying the center area to the center area of ​​the second image attribute pixel structure, and copying the area for each boundary area to the corner area of ​​the second image attribute pixel structure, but selecting the shape of the copied area to exactly match the shape of the corner area.

[0179] The advantage of the filled region approach is that it provides a second image attribute pixel structure that tends to have more consistent pixel values ​​with less variation across the divisions between different regions. This results in more efficient encoding, resulting in a lower data rate for a given quality level.

[0180] The above examples focus on processing a single image. However, it will be appreciated that the method can also be applied to multiple images, such as individual frames of a video sequence.

[0181] Furthermore, in some embodiments, the method can be applied to parallel images, such as a left-eye image and a right-eye image of a stereoscopic image representation of a scene. In such a case, the second processor 305 can generate two rectangular image attribute pixel structures, which can then be encoded. In some embodiments, the rectangular image attribute pixel structures can be combined before encoding. For example, Figure 9 As shown, a single rectangular image attribute pixel structure may be generated by juxtaposing the two independent image attribute pixel structures generated by the second processor 305, and the resulting overall image may be encoded as a single image.

[0182] Therefore, the above apparatus can very efficiently generate an image signal including the described second image attribute pixel structure. In some embodiments, the image signal can be an uncoded image signal (e.g., corresponding to Figure 3 The output of the second processor 305 in the three examples is shown in FIG. 1 , but in many embodiments the image signal may be an encoded image signal (eg corresponding to Figure 3 The output of encoder 307 in the three examples).

[0183] It should be understood that the sink / client / decoder side can receive an image signal including a second image attribute pixel structure as described in the previous example and process it to recreate an image attribute pixel structure corresponding to the original first image attribute pixel structure (i.e. corresponding to a non-rectangular representation of the viewing sphere).

[0184] Figure 10 An example of such an apparatus is shown. The apparatus comprises a receiver 1001 arranged to receive an image signal comprising an image representation of a scene in the form of an image attribute pixel structure representing a viewing sphere as seen from a given viewpoint, as described for the second image attribute pixel structure.

[0185] The second image attribute pixel structure is fed to an inversion processor 1003 which is arranged to perform Figure 3The non-rectangular image attribute pixel structure is generated by performing the reverse operation of the operation performed by the second processor. In particular, it can perform pixel (position) inverse mapping so that the central area of ​​the received second image attribute pixel structure is mapped to the central part of the generated image attribute pixel structure, and the (one or more) corner areas of the second image attribute pixel structure are mapped to the boundary areas of the generated image attribute pixel structure.

[0186] This locally generated non-rectangular image attribute pixel structure can then be output to other functional parts for further processing. For example, in Figure 10 In the example, the generated image attribute pixel structure is fed to the local renderer, which can proceed to synthesize the view image corresponding to the current viewer pose, as known to those skilled in the art.

[0187] It should be understood that, for the sake of clarity, the above description has described embodiments of the present invention with reference to different functional circuits, units, and processors. However, it will be apparent that any suitable distribution of functionality between different functional circuits, units, or processors may be used without departing from the present invention. For example, functions illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Therefore, references to specific functional units or circuits are merely to references to appropriate modules for providing the described functionality, rather than to a strict logical or physical structure or organization.

[0188] The present invention can be implemented in any suitable form, including any combination of hardware, software, firmware or these items. The present invention can optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements and components of embodiments of the present invention can be implemented physically, functionally and logically in any suitable manner. In fact, this function can be implemented in a single unit, in multiple units or as the part of other functional units. Just for this reason, the present invention can be implemented in a single unit or can be distributed between different units, circuits and processors physically and functionally.

[0189] Although the present invention has been described in conjunction with certain embodiments, this is not intended to limit the invention to the specific forms set forth herein. Rather, the scope of the present invention is limited solely by the claims. In addition, although features may appear to have been described in conjunction with specific embodiments, those skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the present invention. In the claims, the term "comprising" does not exclude the presence of other elements or steps.

[0190] In addition, although listed separately, multiple modules, elements, circuits or method steps can be implemented by, for example, a single circuit, unit or processor. In addition, although individual features may be included in different claims, these features may be advantageously combined, and the inclusion of these features in different claims does not mean that the combination of these features is not feasible and / or not advantageous. Moreover, the inclusion of a feature in one claim category does not imply a limitation to that category, but rather indicates that the feature is equally applicable to other claim categories when appropriate. In addition, the order of features in the claims does not mean that the features must work in any particular order, and in particular, the order of the individual steps in a method claim does not mean that the steps must be performed in that order. Instead, the steps can be performed in any appropriate order. In addition, singular references do not exclude plural references. Therefore, references to "one", "first", "second", etc. do not exclude a plurality. The figure marks in the claims are provided merely to clearly illustrate examples and should not be interpreted as limiting the scope of the claims in any way.

Claims

1. A device for generating an image attribute pixel structure representing an attribute of a scene seen from a viewpoint, the device comprising: A first processor (301) provides a first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least part of a surface of a viewing sphere for the viewpoint; as well as A second processor (305) is used to generate a second image attribute pixel structure, which is a two-dimensional rectangular pixel structure and has a second central area derived from the first central area of ​​the first image attribute pixel structure and at least a first corner area derived from the first boundary area of ​​the first image attribute pixel structure, the first boundary area is an area close to one of the upper boundary and the lower boundary of the first image attribute pixel structure, and the first corner area does not overlap with the second central area of ​​the second image attribute pixel structure; wherein the central area of ​​the first image attribute pixel structure is defined by an upper horizontal line corresponding to the upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to the lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure is closer to the periphery than at least one of the upper horizontal line and the lower horizontal line; and the first corner area is close to the relative boundary of the first boundary area.

2. The device according to claim 1, wherein The first image attribute pixel structure has a uniform resolution for the at least portion of the surface.

3. The device according to claim 1 or 2, wherein: The first central area of ​​the first image attribute pixel structure does not overlap with the first boundary area.

4. The device according to claim 1 or 2, wherein: The second processor (305) is arranged to generate the second image attribute pixel structure to have a second corner area derived from a second boundary area of ​​the first image attribute pixel structure, the second corner area does not overlap with the second center area of ​​the second image attribute pixel structure, the second boundary area is an area close to the one of the upper boundary and the lower boundary, and the first boundary area and the second boundary area are on different sides of a virtual vertical line of the first image attribute pixel structure.

5. The device according to claim 4, wherein The virtual vertical line separates the first boundary area from the second boundary area, and the first boundary area and the second boundary area are mirror-symmetrical around the virtual vertical line.

6. The device according to claim 4, wherein A horizontal direction from the first boundary region to the second boundary region is opposite to a horizontal direction from the first corner region to the second corner region.

7. The device according to claim 1 or 2, wherein: The second processor (305) is arranged to extrapolate pixels of at least one of the first corner region and the second central region of the second image attribute pixel structure to an unfilled region of the second image attribute pixel structure adjacent to at least one of the first corner region and the second central region.

8. The device according to claim 1 or 2, wherein: The second processor (305) is arranged to determine the pixel value of the first corner area by at least one of shifting, translating, mirroring and rotating the pixel value of the first boundary area.

9. The device according to claim 1, wherein The first processor (301) is arranged to generate the first image attribute pixel structure by distorting a rectangular image attribute pixel structure representing the at least part of the surface by means of an equirectangular projection.

10. The device according to claim 1 or 2, wherein: The image attribute pixel structure is a depth map.

11. The device according to claim 1 or 2, wherein: The image attribute pixel structure is a light intensity image.

12. A device for generating an output image attribute pixel structure, the device comprising: A receiver (1001) for receiving an image signal including a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a second central area derived from a first central area of ​​the first image attribute pixel structure and at least a first corner area derived from a first boundary area of ​​the first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of a surface of a viewing sphere for a viewpoint, and the first boundary area being an area close to one of an upper boundary and a lower boundary of the first image attribute pixel structure, the first corner area not overlapping with the second central area of ​​the second image attribute pixel structure, the central area of ​​the first image attribute pixel structure being defined by an upper horizontal line corresponding to an upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to a lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure being more peripheral than at least one of the upper horizontal line and the lower horizontal line; and the first corner area being close to an opposite boundary of the first boundary area, and A processor (1003) for generating a non-rectangular output image attribute pixel structure representing at least part of the surface of the viewing sphere for the viewpoint, the non-rectangular output image attribute pixel structure having a central area derived from the second central area of ​​the second image attribute pixel structure and a boundary area derived from the first corner area of ​​the second image attribute pixel structure, the boundary area being an area close to one of an upper boundary and a lower boundary of the output image attribute pixel structure.

13. A method for generating an image attribute pixel structure representing attributes of a scene viewed from a viewpoint, the method comprising: providing a first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of a surface of a viewing sphere for the viewpoint; and Generate a second image attribute pixel structure, which is a two-dimensional rectangular pixel structure and has a second central area derived from the first central area of ​​the first image attribute pixel structure and at least a first corner area derived from the first boundary area of ​​the first image attribute pixel structure, the first boundary area is an area close to one of the upper boundary and the lower boundary of the first image attribute pixel structure, and the first corner area does not overlap with the second central area of ​​the second image attribute pixel structure; wherein the central area of ​​the first image attribute pixel structure is defined by an upper horizontal line corresponding to the upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to the lower edge of the second image attribute pixel structure; the first boundary area of ​​the first image attribute pixel structure is closer to the periphery than at least one of the upper horizontal line and the lower horizontal line; and the first corner area is close to the relative boundary of the first boundary area.

14. A method for generating an output image attribute pixel structure, the method comprising: receiving an image signal including a second image attribute pixel structure, the second image attribute pixel structure being a two-dimensional rectangular pixel structure and having a second central region derived from a first central region of the first image attribute pixel structure and at least a first corner region derived from a first boundary region of the first image attribute pixel structure, the first image attribute pixel structure being a two-dimensional non-rectangular pixel structure representing at least a portion of a surface of a viewing sphere for a viewpoint, the first boundary region being a region proximate to one of an upper boundary and a lower boundary of the first image attribute pixel structure, the first corner region not overlapping with the second central region of the second image attribute pixel structure, the central region of the first image attribute pixel structure being defined by an upper horizontal line corresponding to an upper edge of the second image attribute pixel structure and a lower horizontal line corresponding to a lower edge of the second image attribute pixel structure, the first boundary region of the first image attribute pixel structure being more peripheral than at least one of the upper horizontal line and the lower horizontal line, and the first corner region being proximate to an opposite boundary of the first boundary region; Generate a non-rectangular output image attribute pixel structure representing at least a portion of the surface of the viewing sphere for the view point, the non-rectangular output image attribute pixel structure having a central area derived from the second central area of ​​the second image attribute pixel structure and a boundary area derived from the first corner area of ​​the second image attribute pixel structure, the boundary area being an area close to one of an upper boundary and a lower boundary of the output image attribute pixel structure.

15. A computer program product comprising computer program code means adapted to perform all the steps of the method according to claim 13 or 14, when said program is run on a computer.

Citation Information

Patent Citations

  • Panoramic video polygon sampling method and panoramic video polygon sampling method

    CN106375760A

  • A method, an apparatus and a computer program product for coding a 360-degree panoramic images and video

    GB2548358A