Capture device information
The immersive video data signal encoding and decoding system addresses capture device imperfections by associating image regions with views and sensors, improving rendering quality and resource efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2024-07-05
- Publication Date
- 2026-07-29
AI Technical Summary
Existing immersive video capture technologies fail to account for the undesirable aspects of capture devices, such as parallax and distortion, leading to degraded rendering performance and resource-intensive processing requirements.
A data signal encoding and decoding system that includes parameters defining the differences between capture devices and a pinhole camera model, associating image regions with views and optical sensors, and compensating for undesirable aspects like parallax and distortion.
Enables high-quality rendering and efficient resource utilization by compensating for capture device imperfections, allowing for accurate scene representation and reduced processing demands.
Smart Images

Figure 2026525171000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the transmission of immersive video data including capture device information. The invention also relates to the conversion of such data.
Background Art
[0002] Typically, six degrees of freedom (6DoF) video, i.e., immersive video, is captured using one or more capture devices and processed to form an idealized representation that can be rendered relatively easily with depth image-based rendering (DIBR) technology. The six degrees of freedom are forward and backward, up and down, left and right, yaw, pitch, and roll. A renderer, for example, can convert samples within a view into a mesh, which is a collection of vertices, edges, and faces that defines the shape of an object. An edge is a connection between two vertices. A vertex is a data structure that describes the position of a point in two-dimensional or three-dimensional space or the position of a plurality of points on a surface, along with other information such as color, normal vector, texture coordinates, and other specific attributes.
[0003] A face is a closed set of edges. Vertex color is the value of the color applied to the vertices within a mesh. Vertex color is derived from one or more video components representing color, vertex position is derived from one or more video components representing depth, and the depth value is converted to the vertex position using the view parameters in the conversion step.
[0004] This conversion step is called unprojection, and the view parameters related to unprojection are the intrinsic elements of the camera (e.g., focal length, pixel size) and the extrinsic elements of the camera (e.g., position and orientation). The ability to unproject from an image sample (each sample is formed from one or more video components) to a point in the scene coordinate space is a feature of this 6DoF video.
[0005] A renderer can use a graphical processing unit (GPU) to render a mesh into a framebuffer. The renderer can select only the view closest to the target viewport, or it can select multiple nearby views and render each view into an intermediate framebuffer. (A viewport is the viewpoint from a (virtual) camera facing the target direction.) The intermediate framebuffers are blended to render the final framebuffer. Rendering from multiple views is beneficial when scene objects are partially occluded in some views, or when objects have view-dependent appearances. [Overview of the project] [Problems that the invention aims to solve]
[0006] Therefore, the present invention aims to mitigate, reduce, or eliminate one or more of the above-mentioned drawbacks, either individually or in combination. [Means for solving the problem]
[0007] The present invention is defined by the claims.
[0008] According to an example of an aspect of the present invention, a device for decoding a data signal is provided, the device having a receiver configured to receive a data stream, the data stream having one or more image frames having a plurality of video components, one or more view parameters, one or more capture device parameters, each capture device having one or more optical sensors, the one or more capture device parameters having one or more parameters relating to one or more undesirable aspects of each capture device, including parameters that define the difference between the capture device and a pinhole camera model, a first syntactic element relating each region of each image frame to a view, a second syntactic element relating each view to a capture device, a third syntactic element relating each capture device to one or more optical sensors, and a fourth syntactic element relating each optical sensor to a video component.
[0009] Multiple video components may have at least two of the following: depth components, color components, transparency components, reflection components, and occupancy components.
[0010] The parameters for one or more views include an index of the number of views, the position of each view within the scene (e.g., in scene units, x, y, z coordinates of the view's position), the orientation of the view within the scene, and / or other view parameters defined by the MIV standard.
[0011] One or more parameters relating to the undesirable aspects of each capture device depend on the structure of the capture device and define how this differs from the pinhole camera model. The pinhole camera model is a model for a capture device that assumes the aperture of the capture device occupies a single point in space. The pinhole camera model does not take into account parallax within the capture device (i.e., between different optical sensors of the capture device) or distortion / aberration due to the physical structure of the capture device (e.g., due to the lens of the capture device).
[0012] One or more parameters relating to one or more undesirable aspects of each capture device may include, for example, one or more parameters defining the difference in external parameters between different optical sensors of the capture device, one acquisition artifact index for each camera device, one acquisition artifact model for each camera device, and / or an index of the distance sensing principle for each camera device. One or more parameters relating to the undesirable aspects of each capture device may not include parameters that are conventionally used in 3D reconstruction, which assume each capture device operates as a pinhole camera (e.g., focal length, overall position and orientation of the capture device) (however, differences in the position and orientation of the optical sensors of the capture device, and the overall position and orientation of the capture device are undesirable aspects).
[0013] The first syntax element identifies the view from which each region of each image frame was acquired. In this way, the associated view parameters can be applied to the region during 3D reconstruction. For example, each region of each image frame associated with the view by the first syntax element can be an atlas patch (if the image frame is an atlas).
[0014] The second syntax element indicates the capture device used to acquire the video component in each view. In some examples, the second syntax element can identify the type of capture device (i.e., the capture device model) being used.
[0015] The third syntactic element specifies how many optical sensors each capture device contains, and / or which optical sensors it contains.
[0016] The fourth syntax element identifies which video component was acquired by each optical sensor.
[0017] In some examples, the device has a processor configured to compensate for at least one undesirable aspect.
[0018] In some examples, the device has a transmitter, and the processor and transmitter are coupled to each other, with the transmitter configured to output a signal that compensates for at least one of the undesirable aspects. In other words, the processor can be configured to control the transmitter to output a signal that compensates for at least one of the undesirable aspects.
[0019] According to an example of another aspect of the present invention, an apparatus for encoding a data signal is provided, the apparatus having a transmitter configured to transmit a data stream, the data stream having one or more image frames having a plurality of video components, one or more view parameters, one or more capture device parameters, each capture device having one or more optical sensors, the one or more capture device parameters having one or more parameters relating to one or more undesirable aspects of each capture device, including parameters that define the difference between the capture device and a pinhole camera model, a first syntactic element relating each region of each image frame to a view, a second syntactic element relating each view to a capture device, a third syntactic element relating each capture device to one or more optical sensors, and a fourth syntactic element relating each optical sensor to a video component.
[0020] The inventors have recognized that there are many situations in which a device encoding a data signal cannot compensate for undesirable aspects of a capture device that captures image data within the data signal (for example, encoders with limited computing resources, such as low-latency encoders for live transmission or encoders installed in mobile devices such as smartphones). By encoding information about undesirable aspects within the data signal, these undesirable aspects can be compensated for in such situations.
[0021] According to an example of yet another aspect of the present invention, a data signal is provided, the data signal comprising one or more image frames having a plurality of video components, parameters of one or more views, parameters of one or more capture devices, each capture device having one or more optical sensors, the parameters of the one or more capture devices comprising one or more parameters relating to one or more undesirable aspects of each capture device, including parameters that define the differences between the capture device and a pinhole camera model, a first syntactic element relating each region of each image frame to a view, a second syntactic element relating each view to a capture device, a third syntactic element relating each capture device to one or more optical sensors, and a fourth syntactic element relating each optical sensor to a video component.
[0022] In some examples, at least one capture device has a first optical sensor and a second optical sensor, and at least one of one or more parameters related to one or more undesirable aspects of each capture device has an indicator that the external camera parameters of the first optical sensor differ from the corresponding external parameters of the second optical sensor. For example, this indicator may include one or more external parameters of each optical sensor, and / or the difference between one or more external parameters of each optical sensor and the corresponding external parameters of the entire camera device.
[0023] In some examples, the parameters of at least one of the capture devices include the external parameters of the light source of each capture device.
[0024] In some examples, the fourth syntax element associates one of a plurality of video components with an infrared image sensor. In some examples, the third syntax element can include, for each capture device, an indicator indicating whether the capture device has an infrared image sensor, and the fourth syntax element can associate one of the plurality of video components with the infrared image sensor only if the third syntax element indicates that at least one capture device has an infrared image sensor.
[0025] In some examples, one of the plurality of video components indicates a depth confidence value of a depth estimation process. In some examples, the fourth syntax element has an indicator as to whether any video component represents a depth confidence map, and if so, has an indicator as to which optical sensor is associated with the video component that represents the depth confidence map.
[0026] In some examples, at least one parameter of one or more capture devices is encoded relative to a parameter of a view.
[0027] In some examples, the parameter of at least one of the one or more capture devices includes inertial motion unit data.
[0028] In some examples, the parameter of at least one of the one or more capture devices includes one or more acquisition artifact indicators, and the one or more acquisition artifact indicators include one or more of having a sensor with a rolling shutter, having transcoded video data, having standard (uncalibrated) camera parameters, lacking synchronization between sensors of the capture device, and / or lacking synchronization with other capture devices.
[0029] In some examples, the parameters of at least one of the capture devices include one or more acquisition artifact models, each including one or more of the following: lens aberration, sensor temperature, sensor age, production batch ID, and / or rolling shutter timing. These parameters are also called distortion parameters. In some examples, if the capture device includes multiple optical sensors, the parameters of the capture device may include one or more acquisition artifact models for each optical sensor.
[0030] In some examples, the parameters of at least one of the capture devices include an index of the distance sensing principle of each capture device, where the distance sensing principle is passive stereo, active stereo, structural optics, time-of-flight, light detection and ranging (LiDAR), or a combination thereof.
[0031] According to an example of yet another aspect of the present invention, a method for decoding a data signal is provided, the method comprising: receiving one or more image frames having a plurality of video components; receiving parameters of one or more views; receiving parameters of one or more capture devices, each capture device comprising one or more optical sensors, the parameters of the one or more capture devices being one or more parameters relating to one or more undesirable aspects of each capture device, the parameters relating to the undesirable aspects of the capture device defining the difference between the capture device and a pinhole camera model; receiving a first syntactic element relating each region of each image frame to a view; receiving a second syntactic element relating each view to a capture device; receiving a third syntactic element relating each capture device to one or more optical sensors; and receiving a fourth syntactic element relating each optical sensor to a video component.
[0032] In some examples, the method further includes a step to compensate for at least one of the undesirable aspects.
[0033] According to an example of yet another aspect of the present invention, a method for encoding a data signal is provided, the method comprising the steps of: transmitting one or more image frames having a plurality of video components; transmitting parameters of one or more views; transmitting parameters of one or more capture devices, each capture device comprising one or more optical sensors, the parameters of the one or more capture devices being one or more parameters relating to one or more undesirable aspects of each capture device, the parameters relating to the undesirable aspects of the capture device defining the difference between the capture device and a pinhole camera model; transmitting a first syntactic element relating each region of each image frame to a view; transmitting a second syntactic element relating each view to a capture device; transmitting a third syntactic element relating each capture device to one or more optical sensors; and transmitting a fourth syntactic element relating each optical sensor to a video component.
[0034] According to one aspect of the present invention, a device for decoding data signals is provided, the device having a receiver configured to receive a data stream, the data stream having one or more image frames including a plurality of video components, parameters of one or more views, and parameters of one or more capture devices, the region of the image frame being associated with a view, the view being associated with a capture device, the capture device having one or more optical sensors, the optical sensors being associated with video components.
[0035] There are several types of capture devices for immersive video, each with different physical principles. Generally, these capture devices acquire image data with undesirable aspects stemming from the device's structure and calibration. These undesirable aspects are often inherent in the physical principles and are expected to persist even as capture technology advances. Ignoring these undesirable aspects often degrades rendering performance, particularly when compositing new viewpoints from captured viewpoints. Extensive processing using computer vision techniques is typically required to convert the captured data into a pure multi-view format suitable for high-quality rendering.
[0036] For example, to calculate a depth map, a capture device may have two optical sensors spaced slightly apart. Another optical sensor may be used to capture a color image. In the context of immersive video, having a capture device that can output both color and depth images is beneficial, but it also has the disadvantage of parallax between the color and depth images. In a multiview representation consisting of two video components, one representing color and the other representing depth, it is useful and usually expected that the two video components share the same internal and external camera parameters as if they were recorded by a single optical sensor. One advantage of this is that there is a direct relationship between the image samples of the video component corresponding to color and the image samples of the video component corresponding to depth.
[0037] Despite some undesirable aspects, it is advantageous for the decoder to receive immersive video data in a multiview format. However, to enable high-quality rendering, it is also advantageous to receive information about the capture device used to capture the immersive video data. In particular, it is advantageous for the decoder to be able to receive information about the optical sensors within these capture devices. The decoder output containing this information is better suited for compositing new viewports.
[0038] In telepresence applications, mobile devices may lack sufficient processing capabilities to convert the data into a suitable format for transmission, such as encoding the data as multiview + depth data.
[0039] In archiving applications such as surveillance and cultural heritage, it is desirable to preserve the original data for future processing. In the future, better methods for processing the data may become available. Furthermore, it may be necessary to retrieve only a small portion of the archived data, and not all of it needs to be processed. Even if data is retrieved, processing it after it has been archived may be acceptable.
[0040] Beyond novel view synthesis, another application of decoded immersive video data is scene understanding, where captured data is used to characterize a scene. For example, data can be acquired by a robotic probe, and the decoded data can be used as input for scientific experiments. Information about undesirable aspects output by the decoder is advantageous because it improves the accuracy of the scientific experiment.
[0041] Having information about the capture device present in the data signal and received by the decoder is advantageous because it allows for a more accurate representation of the captured scene. In particular, the decoded data can be used as input to a processor to compensate for one or more undesirable aspects.
[0042] According to an optional feature of the invention, the capture device suffers from one or more undesirable aspects, and the device has a processor configured to compensate for at least one undesirable aspect.
[0043] Decoders benefit from receiving immersive video data that includes information about undesirable aspects arising from, for example, the mechanical structures, electronic circuits, and physical principles used to acquire the image data. This is because, without information about undesirable aspects, accurately reconstructing or rendering the immersive video data becomes difficult or impossible.
[0044] Preferably, the decoder has a processor to compensate for at least one undesirable aspect, because without that undesirable aspect, the decoder's output is more usable in subsequent actions such as 6DoF rendering. Generally, the more undesirable aspects are compensated for, the more general and versatile the decoder's output becomes. When the decoder's output has no undesirable aspects, the decoder's output has an ideal immersive video format that is generally applicable.
[0045] According to an optional feature of the present invention, the device has a transmitter, the processor and the transmitter are coupled (to each other), and the transmitter outputs a signal that compensates for at least one of the undesirable aspects.
[0046] In some applications, resource constraints exist on both sides of the transmission channel, such as computing power, memory, battery consumption, and / or network bandwidth. In this case, both the encoder (part of the server) and the decoder (part of the client) may not be able to process immersive video data in a way that reduces undesirable aspects. This can result in a degraded experience, for example, when rendering immersive video data to the viewport.
[0047] However, there may be other network nodes available, such as cloud servers, that have sufficient resources. Such nodes can function well as transcoders for immersive video data by decrypting the data, processing it to remove one or more undesirable aspects, and then transmitting the compensated data.
[0048] It is advantageous to couple a decoder with a processor to the transmitter. This is because a resource-constrained client can receive compensated immersive video data from the transcoder rather than from the server, and the decoded data signal from the transcoder provides a better representation of the scene.
[0049] Transcoders are particularly useful in telepresence applications where several mobile nodes act as both servers capturing image data and clients rendering the image data. This is because the mobile nodes can be connected via cloud services, which in turn provide transcoder nodes to offload processing to compensate for undesirable aspects.
[0050] Transcoders are also useful in (live) broadcast scenarios where one or more servers interact with numerous clients. By offloading the handling of undesirable aspects to one or more transcoders, infrastructure and operational costs can be shared across multiple productions. Furthermore, the number of devices that need to be installed on-site is reduced.
[0051] According to one aspect of the present invention, an apparatus for encoding a data signal is provided, the apparatus having a transmitter configured to transmit a data stream, the data stream comprising one or more image frames having a plurality of video components, parameters of one or more views, and parameters of one or more capture devices, wherein a region of the image frame is associated with a view, the view is associated with a capture device, the capture device has one or more optical sensors, and the optical sensors are associated with video components.
[0052] In multiview representations, the physical camera is replaced by a (logical) camera model that typically represents an idealized camera with a sensor plane and an ideal lens. This is called the pinhole model, and the equation that relates a point in scene space to its position on the sensor plane is called perspective projection. Examples of idealized camera models include equirectangular projection and orthographic projection.
[0053] Such multi-view representations involve the concept of image frames (with depth) that have multiple video components. For example, an image frame has color components and depth components. An image frame may also contain other components such as transparency and reflectivity information. Also, an image frame does not necessarily have color components and depth components. Some components, such as color itself, are composed of multiple channels. Color is usually composed of red (R), green (G), and blue (B) channels, or one luma (Y) and two chroma channels (C). B, C R It consists of ).
[0054] In some immersive video data formats, there is an indirect relationship between view parameters and video components, making it possible to associate image regions with views. In such formats, one or more image frames can be transmitted, each containing multiple image regions, and these image regions are associated with views. These views are then associated with parameters. Such generalized image frames are called image atlases, and their advantage is that only the essential parts of multiple views can be transmitted. Given a budget for luma sample rate, luma sample count, or bitrate, it is possible to transmit more views if similar parts of views can be omitted.
[0055] In other, more restrictive immersive video data formats, the image region is limited to the same size as the image frame, and each image frame represents only one view. In such formats, the image region implicitly exists, and it is sufficient to associate the view with the image frame.
[0056] Encoders are more efficient and therefore advantageous when transmitting immersive video data streams containing multiple views with multiple video components. Other immersive / 6DoF data formats, such as dynamic point clouds, dynamic meshes, or neural representations, may require more processing to convert the captured data into a transmittable format. With multiview formats, multiple capture devices can be connected to the video encoder, and the resulting encoded video data can be combined with metadata to form an immersive video data signal suitable for 6DoF rendering.
[0057] It is particularly advantageous for an encoder to transmit an immersive video data stream that includes information about one or more capture devices used to capture the image data used to form the immersive video data stream. This is because the client can compensate for undesirable aspects of the capture device, while the encoder can avoid processing those undesirable aspects. This allows the client to accurately represent the captured scene and reduces the amount of resources required by the encoder. For example, it may be possible to reduce the number of CPU cycles, reduce memory consumption, or reduce outbound network bandwidth.
[0058] A capture device may suffer from one or more undesirable aspects. A view can be associated with a capture device, and the parameters of the view include acquisition parameters, which are associated with aspects of acquisition.
[0059] Associating the view with the capture device is advantageous because the data signal represents typical immersive video data, and the acquisition parameters are additional information. The parsing and decoding processes are largely the same as for typical immersive video data.
[0060] Furthermore, there may be multiple capture devices of the same model. Associating a view with a capture device determines which capture parameters need to be notified for each view. Here, a view corresponds to a capture device, which is an instance of the capture device model. Some capture parameters are the same for all capture devices of the same capture device model and only need to be notified once, while other capture parameters may differ (slightly) for multiple capture devices of the same capture device model and these are notified for each view. If multiple capture device models exist, the capture parameters notified per model or per view may be specific to a particular capture device model.
[0061] In other words, by associating a view with a capture device and then associating the view with acquisition parameters dependent on that capture device, efficient transmission is achieved, and it is also practical from an implementation standpoint. General immersive video encoding optimizations can then be naturally applied to encoding immersive video data that includes undesirable (acquisition) aspects.
[0062] According to one aspect of the present invention, a data signal is provided, the data signal comprising one or more image frames having a plurality of video components, parameters of one or more views, and parameters of one or more capture device models, wherein a region of the image frame is associated with a view, the view is associated with a capture device, the capture device has one or more optical sensors, and the optical sensors are associated with video components.
[0063] Immersive video data signals can be represented by multiple video components and metadata. These multiple video components form an image frame (with depth). There may be multiple such image frames. The image region of an image frame is associated with a view. In some representations, the image region is the entire image frame, but in other cases, each view may be formed by multiple smaller image regions. This makes it possible to transmit partial views.
[0064] In practical acquisition, the view is captured by a capture device, which has multiple optical sensors. The characteristics of these optical sensors, the physical acquisition principle, and the processing of the output of these sensors within the capture device usually result in undesirable aspects of the acquisition. Examples of this are shown in this patent application.
[0065] Describing optical sensors and associated undesirable acquisition aspects (i.e., obtaining data signals containing information describing them) is advantageous because it leads to a more accurate description of the captured scene. Furthermore, compensating for these aspects immediately after capture is difficult. By saving data signals containing information about optical sensors and undesirable acquisition aspects, it is beneficial to postpone the processing to compensate for undesirable acquisition aspects until later in the transmission chain, or even later in time.
[0066] The capture device model may contain undesirable aspects. For this reason, it is advantageous for one or more optical sensors to include a first optical sensor and a second optical sensor. One undesirable parameter is an indicator that shows the external camera parameters of the first optical sensor differ from the parameters of the (corresponding) second optical sensor.
[0067] One physical principle that can be used to detect depth values is parallax (also known as triangulation), where the same scene is captured from two (slightly) different viewpoints. The shift of scene elements from one sensor to the other can be measured and converted into distance. Examples of types / classes of capture devices that utilize the parallax effect include passive stereo, active stereo, and structural optics. These essentially have a first optical sensor and a second optical sensor or light source.
[0068] Furthermore, in some devices, separate optical sensors are used for color information and depth information, and it is advantageous to optimize each sensor individually for its separate application. These optical sensors need to be placed next to each other, resulting in parallax.
[0069] Therefore, the data signal includes multiple video components associated with a view and captured by two optical sensors in a capture device associated with that view, with parallax existing between the two optical sensors. This causes a shift in the scene object, which can result in rendering artifacts if not corrected.
[0070] When parallax exists in the capture device (instance, model, class), it is advantageous for the data signal to include external camera parameters of the first and second optical sensors.
[0071] According to an optional feature of the present invention, the capture device is a depth camera including a light source, and the parameters of the capture device include external parameters of the light source.
[0072] The problem with camera devices that measure depth using the parallax principle is that they need to match scene elements (pixels) on two optical sensors, which usually results in ambiguity. Active sensors mitigate this problem by having a light source project a pattern onto the objects in the scene. In structural optics-class devices, the pattern is detected by optical sensors, and the depth value of each element of the pattern is estimated using the shift of the pattern relative to a reference. In active stereo-class devices, the pattern is detected by multiple optical sensors and used to reduce ambiguity by adding an additional texture to the scene objects.
[0073] In either case, knowing the relative position of the light source is advantageous because the quality of depth estimation depends on the projection of the pattern, whether current or future, in other classes of other capture devices. For example, one face of a scene object may be estimated more accurately than the other.
[0074] According to an optional feature of the present invention, the capture device has an infrared image sensor, and the infrared image sensor is associated with a video component.
[0075] The light sources described above typically emit invisible light so as not to disturb the scene for human observers. Infrared light, just outside the visible light range, for example, 800-900 nm, is the most common, but other wavelengths such as UV can also be used. A capture device that measures depth values using one or more light sources and one or more invisible light optical sensors usually has a processor implemented in an ASIC or FPGA that estimates parallax and derives depth values.
[0076] Including infrared video components in the data signal is advantageous because it allows the receiving node to have more resources and process non-visible light images more accurately.
[0077] Having multiple invisible light images is particularly advantageous for data signals and capture devices with multiple views. This is because the processor can take all images into account when estimating depth, which usually leads to more accurate estimations. Such a processor has more information than the processor in a single capture device.
[0078] According to an optional feature of the present invention, the video component is associated with one of one or more optical sensors. The video component displays the depth confidence value for the depth estimation process.
[0079] A capture device equipped with a depth value estimation processor can estimate depth values using multiple consecutive image exposures. Such a capture device aggregates and outputs multiple measurements, such as a mean or median depth estimate. Within such a capture device, depth confidence can be estimated by examining the inconsistencies between multiple depth value measurements. This inconsistency can be modeled, for example, by the standard deviation of the depth value estimates. Such depth confidence can be used within the capture device to filter out unreliable samples (e.g., by "inpainting") by deriving values from reliable neighboring samples.
[0080] It is advantageous for the data signal to have a video component that represents a depth confidence value for a depth value derived from one of one or more optical sensors, associated with that optical sensor, thereby allowing another processor to improve the quality of the immersive video data. Such a processor can perform better inpainting, for example, by capturing multiple views with multiple depth cameras. Alternatively, the processor can ignore samples if another view has more reliable samples of the same scene area.
[0081] According to an optional feature of the present invention, the parameters of the optical sensor are encoded with respect to the parameters of the view.
[0082] Typically, capture devices are spaced apart to cover the entire scene using a limited number of devices. Each capture device can be spaced, for example, 100-150 mm apart and 1 m apart. Within each capture device, multiple optical sensors are placed close together. Other parameters of the optical sensors, such as focal length, are typically adjusted to make the images from multiple optical sensors more similar. Therefore, the camera parameters of the optical sensors, including intrinsic camera parameters, extrinsic camera parameters, and distortion parameters, are likely to be similar. This view is an idealized (pinhole) model of the optical sensor, with extrinsic and intrinsic camera elements similar to those of the optical sensor.
[0083] Encoding optical sensor parameters relative to view parameters is advantageous because the difference (delta) between parameters is expected to be smaller than the parameters themselves, requiring fewer bits (and thus less entropy) to encode the difference. Furthermore, delta parameters, which represent relative camera external elements as translation and rotation, are the same across multiple camera devices of the same model, further reducing the bitrate by notifying such shared parameters only once.
[0084] According to an optional feature of the present invention, the capture device is a camera having an inertial motion unit, and the parameters of the capture device include data from the inertial motion unit.
[0085] In many cases, camera equipment is mounted on a sturdy stand (e.g., a tripod) or mount to avoid movement and vibration. Typically, camera equipment is a rigid box with multiple optical and electronic components mounted on one or more subframes. However, the actual materials, such as metal and plastic, have some flexibility; otherwise, the camera equipment and its mounting would be brittle and prone to breakage during normal use. Some camera equipment has an inertial motion unit (IMU), such as an accelerometer or gyroscope, to detect motion.
[0086] It is advantageous that the data signal includes data from the inertial motion unit, which the processor can use to perform better online calibration of camera parameters.
[0087] According to an optional feature of the present invention, the parameters of the capture device include one or more acquisition artifact indicators, the acquisition artifacts include one or more of the following: having a sensor with a rolling shutter, having transcoded video data, having standard (uncalibrated) camera parameters, lack of synchronization between sensors of the capture device, and lack of synchronization with other capture devices.
[0088] Ideally, image data from multiple views should be captured simultaneously, as this allows for easy rendering from multiple views. This is usually considered a given. If a video component derived from one optical sensor is not synchronized with another video component derived from a different optical sensor, moving scene objects will be rendered at different positions depending on how the views are blended. This can result in unsightly visual artifacts. Synchronizing optical sensors is particularly beneficial in sports applications with a lot of movement.
[0089] However, this is not always possible or advantageous. For example, multiple unsynchronized capture devices may be cheaper or require less setup, and such capture devices may only be synchronized to an accuracy of one or a few frames of delay.
[0090] In such situations, it is advantageous for the data signal to include an indicator of synchronization deficiency, as this can be compensated for after the data signal has been transmitted. Conversely, if it is indicated that the synchronization is accurate, such compensation becomes unnecessary, saving these resources.
[0091] If this cannot be compensated for, subsequent processing that would result in errors due to inaccurate synchronization can be avoided, and appropriate error messages can be provided as needed.
[0092] According to an optional feature of the present invention, the parameters of the optical sensor include one or more acquisition artifact models, the acquisition artifact models including one or more of lens aberrations, sensor temperature, sensor years of use, production batch ID, and roll shutter timing.
[0093] Undesirable aspects of acquisition can be compensated for by providing sufficient information on these aspects. Therefore, it is advantageous to provide information on these aspects, such as (but not limited to) lens aberrations, sensor temperature, sensor age, production batch ID, and rolling shutter timing, as data signals containing these parameters provide a more accurate scene representation.
[0094] According to an optional feature of the present invention, the capture device model includes an indicator of the distance detection principle, which includes passive stereo, active stereo, structural optics, time of flight, light detection and ranging (LiDAR), or a combination thereof.
[0095] Undesirable acquisition aspects often stem from the device's structure or the physical principles used to capture color and / or depth information. Compensating for undesirable aspects benefits from knowing which principles were used, because it allows for the application of specific strategies to compensate for those aspects.
[0096] According to one aspect of the present invention, a method for decoding a data signal is provided, the method comprising the steps of: receiving one or more image frames having a plurality of video components; receiving parameters of one or more views; receiving parameters of one or more capture devices; receiving parameters of one or more optical sensors; associating a region of the image frame with a view; associating a view with a capture device; associating a capture device with one or more optical sensors; and associating an optical sensor with a video component.
[0097] A capture device may suffer from one or more undesirable aspects. This method includes a step to compensate for at least one of these undesirable aspects.
[0098] According to one aspect of the present invention, a method for encoding a data signal is provided, the method comprising the steps of: transmitting one or more image frames having a plurality of video components; transmitting parameters of one or more views; transmitting parameters of one or more capture devices; transmitting parameters of one or more optical sensors; associating a region of an image frame with a view; associating a view with a capture device; associating a capture device with one or more optical sensors; and associating an optical sensor with a video component.
[0099] The present invention also provides a computer program carrier that, when executed on a computer, causes the computer to perform all steps according to any of the methods described above.
[0100] A computer program carrier can include computer memory (e.g., random access memory) and computer storage (e.g., hard drives, solid-state drives, etc.). A computer program carrier can also be a bitstream that transmits computer program code.
[0101] The present invention also provides a system comprising a processor configured to perform all steps by any of the methods provided by the present invention as described above.
[0102] The present invention also provides a computer implementation method comprising all steps of any of the methods provided by the present invention as described above.
[0103] These and other aspects of the present invention will become apparent from and be explained with reference to the embodiments described below. [Brief explanation of the drawing]
[0104] To better understand the present invention and to more clearly illustrate how it can be put into practice, refer to the accompanying drawings as merely examples. [Figure 1] A diagram illustrating how multiple depth-sensing devices are used to capture a single scene using prior art. [Figure 2] A diagram illustrating how multiple capture devices are used to form a single data signal transmitted to a client according to the present invention. [Modes for carrying out the invention]
[0105] The present invention will be described with reference to the drawings.
[0106] The detailed descriptions and specific examples illustrate exemplary embodiments of the apparatus, systems, and methods, but should be understood to be for illustrative purposes only and not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatus, systems, and methods of the invention will be better understood from the following description, the appended claims, and the appended drawings. The drawings are for illustrative purposes only and are not drawn to a specific scale. Also, the same reference numerals are used throughout the drawings to indicate the same or similar parts.
[0107] In the following, instead of terms such as "combined" or "joined," we will use terms such as "connected" or "connected" as examples. This is because terms such as "connected" are more commonly used in this technical field. However, in the context of this application, it should be emphasized that the use of terms such as "connected" includes not only directly connected things, but also things that are "connected through something else," such as a medium (like air), a buffer, or an interface such as an amplifier.
[0108] Figure 1 illustrates a method for capturing a single scene based on prior art using multiple depth-sensing devices. This represents the currently available Intel RealSense depth sensor, which is active stereo-based. A processor board connected to each sensor receives a so-called raw depth map, which is calculated using stereo matching between the first and second infrared sensors. In the case of the Intel RealSense module, this stereo estimation step is performed on the ASIC within the device (rather than on the processor board).
[0109] Figure 2 illustrates how multiple capture devices are used to form a single data signal transmitted to a client according to the present invention.
[0110] For example, in contrast to Intel, which provides software to improve depth maps using post-processing filtering on a connected host computer, the present invention delays this process until after transmission to a more powerful computer that can combine data from multiple sensors to improve the process. In the present invention, the processor board is suitable for transmitting depth data in real time and with low latency to a computer that improves the depth and generates a 6DoF data stream.
[0111] Each depth camera capture device transmits either a raw depth map (z_1, z_2) linearly encoded in, for example, raw millimeters (8 bits or more), or a depth map encoded as inverse depth. In addition to the depth map, the corresponding infrared map of the viewpoint (I_1, I_2) is also transmitted to the 6DoF format generator. These have already been used for depth estimation within each device (using the first and second infrared sensors), but are now useful for improving depth using depth enhancement between devices. Starting with the initial depth map estimated by each separate device, the overall quality of each depth map is further improved by combining infrared images from multiple cameras. This depth map enhancement step requires knowing whether the processor board connected to each device has already performed depth distortion correction and whether depth filtering has already been applied to the data. This information is transmitted via metadata (m_1, m_2). The metadata (m_1, m_2) can also indicate nominal pose information regarding how each device is positioned in common-world space. For example, this can help speed up the pose calibration procedure performed by the 6DoF format generator.
[0112] Following the depth sensor device, Figure 2 also shows a color sensor device 3. This device transmits color along with metadata, which can indicate whether the color sensor is a rolling shutter or a global shutter. Color sensor devices 4, 5, ... N may also exist (not shown in Figure 2), in which case the shutter type of the color sensor (rolling or global) is important. If multiple color sensors are of the global shutter type, they can be used for depth correction based on matching.
[0113] Various other methods can be considered for encoding and decoding data signals using capture device information.
[0114] Those skilled in the art can easily develop a processor to perform any of the methods described herein. Thus, each step in the flowchart represents a different action performed by the processor, which can be performed by each module of the processing processor.
[0115] One or more steps of any method described herein may be performed by one or more processors. A processor consists of electronic circuits suitable for processing data. Any method described herein may be computer-implemented, where computer implementation means that the steps of the method are performed by one or more computers, where a computer is defined as a device suitable for data processing. A computer is suitable for processing data according to given instructions.
[0116] As described above, the system utilizes a processor to perform data processing. The processor is implemented in various ways using software and / or hardware to perform the various functions required. The processor typically uses one or more microprocessors programmed to perform the required functions using software (e.g., microcode). The processor may also be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuits for performing other functions.
[0117] Examples of circuits used in various embodiments of this disclosure include, but are not limited to, conventional microprocessors, application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).
[0118] In various implementations, the processor may be associated with one or more storage media, which are volatile and non-volatile computer memories such as RAM, PROM, EPROM, and EEPROM. These storage media may be encoded with one or more programs that perform the required functions when executed on one or more processors and / or controllers. The various storage media may be mounted within the processor or controller, or they may be transportable so that one or more programs stored in the storage media can be loaded into the processor.
[0119] It should be noted that this invention is relevant and technically related to the current immersive video data transmission standard and future updates based on ISO / IEC 23090-12 MPEG Immersive Video (MIV). At the time of writing, the first edition of this standard is scheduled to be published in 2023, and the second edition is in the draft stage.
[0120] The data signal according to the standard consists of a sequence of so-called V3C units. A V3C unit consists of a header and a payload. The first V3C unit is called the V3C parameter set (VPS) and is used by the decoder to initialize or stop. All subsequent V3C units form multiple sub-bitstreams, with the V3C unit header identifying the sub-bitstream and the payload containing a portion of the bits of the video sub-bitstream.
[0121] There are sub-bitstreams containing encoded video components, sub-bitstreams containing so-called atlas data, and sub-bitstreams containing so-called common atlas data. The common atlas data contains view parameters for multiple views. One of the view parameters is a numerical view identifier (View ID). The atlas data contains patch parameters. The patch described by these parameters associates a rectangular image region of the atlas frame with a rectangular image region on the view's projection plane. One of the patch parameters is the projection ID, which associates the patch with one of the decoded views through its view ID.
[0122] In an embodiment of the present invention shown as an example, this standard is extended to include an association between a view and a capture device model, the capture device model having one or more optical sensors, and the optical sensors are associated with a video component.
[0123] The VPS MIV extension is part of the VPS and is extended with “vme_capture_device_information_present_flag”, which optionally notifies capture device information. This flag can be used to create subprofiles. A decoder (client) that can decode (and render) only immersive video data without unwanted capture aspects can use this flag to determine whether or not to decode the bitstream. [Table 1]
[0124] The capture device information syntax structure provides a subset of relevant parameters that help initialize the decoder. This information is also used to initialize a processor suitable for compensating for undesirable aspects of acquisition. In this example, multiple views can be associated with a single capture device model by notifying the device model ID. The model class is notified by an enumeration of specified values, e.g., 0=unknown, 1=single color camera, 2=passive stereo camera, 3=active stereo camera, etc. The number of sensors is indicated, and a parallax ID is indicated for each sensor. In this example, if multiple optical sensors have the same parallax ID, they are modeled as having no parallax. Optical sensors may be coaxial. A single physical optical sensor may also be modeled as two optical sensors. [Table 2]
[0125] To avoid parsing dependencies between the VPS and the common atlas subbitstream, some of the syntactic elements described above can be repeated within the Common Atlas Sequence Parameter Set (CASPS). In short, in this example, the syntactic elements described above are used directly within the CASPS without repetition.
[0126] The camera external elements of view v can be extended to indicate additional parameters in the MIV view parameter list. Similarly, capture device parameters can be updated in the same way that view parameter updates can be provided in the MIV. [Table 3]
[0127] It should be understood that standards are extended not only by syntax but also by appropriate semantics and decoding processes. In particular, there may be decoder processes (ISO / IEC 23090-12 Article 9) that output variables or arrays based on existing or inferred syntactic elements.
[0128] Modifications of the disclosed embodiments can be understood and implemented by those skilled in the art in carrying out the claimed invention, based on a review of the drawings, disclosures, and appended claims. In the claims, the words “comprising” do not exclude other components or steps, and the indefinite articles “a” or “an” do not exclude plurality.
[0129] A single processor or other unit can perform the functions of several of the items listed in the claims.
[0130] The mere fact that certain means are described in mutually different dependent claims does not indicate that combinations of these means cannot be used advantageously.
[0131] Computer programs can be stored / distributed on suitable media such as optical or solid-state media supplied together with or as part of other hardware, but they can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0132] When the term "adapt" is used in a claim or specification, it means that the term "adapt" is equivalent to the term "constituted".
[0133] No reference numeral in a claim should be construed as limiting the scope.
[0134] Any method described herein excludes methods of performing mental acts.
[0135] Modifications of the disclosed embodiments can be understood and implemented by those skilled in the art in carrying out the claimed invention, based on a review of the drawings, disclosures, and appended claims. In the claims, the words “comprising” do not exclude other components or steps, and the indefinite articles “a” or “an” do not exclude plurality.
[0136] A single processor or other unit can perform the functions of several of the items listed in the claims.
[0137] The mere fact that certain means are described in mutually different dependent claims does not indicate that combinations of these means cannot be used advantageously.
[0138] Computer programs can be stored / distributed on suitable media such as optical or solid-state media supplied together with or as part of other hardware, but they can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0139] When the term "adapt" is used in a claim or specification, it means that the term "adapt" is equivalent to the term "constituted".
[0140] No reference numeral in a claim should be construed as limiting the scope.
[0141] Any method described herein excludes methods of performing mental acts.
Claims
1. A device for decoding a data signal, the device having a receiver configured to receive a data stream, the data stream being One or more image frames with multiple video components, Parameters of one or more views, Parameters of one or more capture devices, each capture device having one or more optical sensors, and the parameters of the one or more capture devices include one or more parameters related to one or more undesirable aspects of each of the capture devices, and the parameters related to the undesirable aspects of the capture devices define the difference between the capture device and the pinhole camera model, A first syntax element that associates each region of each image frame with a view, A second syntactic element that associates each view with a capture device, A third syntactic element that associates each capture device with one or more optical sensors, A fourth syntax element that associates each optical sensor with a video component, A device including a device.
2. The apparatus according to claim 1, comprising a processor configured to compensate for at least one of the aforementioned undesirable aspects.
3. Having a transmitter, The processor and the transmitter are coupled to each other. The apparatus according to claim 2, wherein the transmitter is configured to output a signal that compensates for at least one of the undesirable aspects.
4. A device for encoding a data signal, configured to transmit a data stream, the device having a transmitter, the data stream is, One or more image frames with multiple video components, Parameters of one or more views, Parameters of one or more capture devices, each capture device having one or more optical sensors, and the parameters of the one or more capture devices include one or more parameters related to one or more undesirable aspects of each of the capture devices, and the parameters related to the undesirable aspects of the capture devices define the difference between the capture device and the pinhole camera model, A first syntax element that associates each region of each image frame with a view, A second syntactic element that associates each view with a capture device, A third syntactic element that associates each capture device with one or more optical sensors, A fourth syntax element that associates each optical sensor with a video component, A device including a device.
5. One or more image frames with multiple video components, Parameters of one or more views, Parameters of one or more capture devices, each capture device having one or more optical sensors, and the parameters of the one or more capture devices include one or more parameters related to one or more undesirable aspects of each of the capture devices, and the parameters related to the undesirable aspects of the capture devices define the difference between the capture device and the pinhole camera model, A first syntax element that associates each region of each image frame with a view, A second syntactic element that associates each view with a capture device, A third syntactic element that associates each capture device with one or more optical sensors, A fourth syntax element that associates each optical sensor with a video component, A data signal having a data signal.
6. At least one capture device, described by the parameters of one or more capture devices, has a first optical sensor and a second optical sensor, The data signal according to claim 5, wherein at least one of one or more parameters related to one or more undesirable aspects of each of the capture devices includes an index indicating that the external camera parameters of the first optical sensor are different from the corresponding external parameters of the second optical sensor.
7. The data signal according to claim 5 or 6, wherein at least one of the parameters of the one or more capture devices includes an external parameter of the light source of each capture device.
8. The data signal according to claim 5 or 6, wherein the fourth syntactic element associates one of the plurality of video components with an infrared image sensor.
9. The data signal according to claim 5 or 6, wherein one of the plurality of video components indicates a depth confidence value for the depth estimation process.
10. The data signal according to claim 5 or 6, wherein at least one of the parameters of the one or more capture devices is encoded with respect to the view parameters.
11. The data signal according to claim 5 or 6, wherein at least one of the parameters of the at least one capture device includes inertial motion unit data.
12. The parameter of at least one of the one or more capture devices includes one or more acquired artifact indicators, and the one or more acquired artifact indicators of the capture device are Having a sensor with a rolling shutter, Having transcoded video data, Having nominal (uncalibrated) camera parameters, Lack of synchronization between optical sensors in the capture device, Lack of synchronization with other capture devices, The data signal according to claim 5 or 6, which is one or more of the following.
13. At least one of the at least one capture devices includes one or more acquired artifact models, and the acquired artifact models are Lens aberrations, Sensor temperature, Sensor usage period, Production batch ID, Roll shutter timing, A data signal according to claim 5 or 6, comprising one or more of the following.
14. At least one of the parameters of the at least one of the capture devices includes an index of the distance detection principle of each capture device, The aforementioned distance detection principle is, Passive stereo, Active stereo, structured light, Flight time, Light detection and ranging (LiDAR), Any of the above combinations, The data signal according to claim 5 or 6.
15. A method for decoding a data signal, The steps include receiving one or more image frames having multiple video components, Steps include receiving parameters for one or more views, A step of receiving parameters from one or more capture devices, each capture device having one or more optical sensors, and the parameters of the one or more capture devices including one or more parameters related to one or more undesirable aspects of each of the capture devices, the parameters related to the undesirable aspects of the capture devices defining the difference between the capture device and a pinhole camera model, The steps include receiving a first syntactic element that associates each region of each image frame with a view, The steps include receiving a second syntactic element that associates each view with a capture device, The steps include receiving a third syntactic element that associates each capture device with one or more optical sensors, The steps include receiving a fourth syntactic element that associates each optical sensor with a video component, A method of having.
16. The method according to claim 15, further comprising the step of compensating for at least one of the aforementioned undesirable aspects.
17. A method for encoding a data signal, The steps include sending one or more image frames that have multiple video components, The steps include sending parameters for one or more views, A step of transmitting parameters for one or more capture devices, each capture device having one or more optical sensors, and the parameters for the one or more capture devices including one or more parameters relating to one or more undesirable aspects of each of the capture devices, the parameters relating to the undesirable aspects of the capture devices defining the difference between the capture device and a pinhole camera model, The steps include sending a first syntax element that associates each region of each image frame with a view, The steps include sending a second syntactic element that associates each view with a capture device, The steps include transmitting a third syntactic element that associates each capture device with one or more optical sensors, The steps include sending a fourth syntactic element that associates each optical sensor with a video component, A method of having.
18. A computer program that runs on a computer and causes the computer to perform the method described in any one of claims 15 to 17.
19. A system having a processor configured to perform the method described in any one of claims 15 to 17.