VIDEO CODING FOR STRUCTURE-FROM-MOTION (SfM)
By adapting video encoding quality based on motion and blur analysis, and adjusting GOP sizes, the method optimizes bit rate for SfM applications, ensuring high-quality frames are encoded, thereby improving the efficiency and quality of 3D reconstruction.
Patent Information
- Application Number
- PCT/EP2024/050965
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
Existing video compression solutions are not suitable for videos intended for structure-from-motion (SfM) applications, as they do not effectively prioritize encoding based on the needs of machine vision tasks, leading to inefficient bit usage and potential degradation of 3D reconstruction quality.
A method for encoding videos that adapts the quality of individual pictures based on motion and blur analysis, using an analyzer to guide the encoder to encode certain pictures at higher quality and others at lower quality, and adjusting Group of Picture (GOP) sizes to optimize bit rate without affecting the SfM end result.
This approach allows for significant bit rate savings while maintaining the quality of the 3D model produced by SfM methods, ensuring that critical frames for feature extraction are encoded at high quality, thus enhancing the efficiency of machine vision tasks.
Smart Images

Figure EP2024050965_24072025_PF_FP_ABST
Abstract
Description
VIDEO CODING FOR STRUCTURE-FROM-MOTION (SfM) TECHNICAL FIELD
[0001] Disclosed are embodiments related to video coding. 5 BACKGROUND
[0002] Video Compression
[0003] As more and more video is being produced, the target receiver of these videos has changed. Previously, video was primarily consumed by humans, with the consequence that both compression standards and encoders were optimized for the human visual system. Nowadays, the 10 focus has shifted more and more towards machines (e.g., algorithms) analyzing video content. With machines evaluating content that is produced by other machines, humans are no longer in the loop, i.e., there is no need to optimize standards or encoders towards preserving the optimal quality for humans.
[0004] Therefore, if the encoder knows that the produced video stream will be primarily 15 used by other machines, it can optimize the encoding towards features that are more important to machines. This topic is referred to “video coding for machines (VCM)”. VCM solutions can rely on encoding video using new codecs targeting machine consumption and / or using existing codecs with machine consumption in mind.
[0005] The current state-of-the-art video compression standard is known as “Versatile 20 Video Coding (VVC)”. VVC has been developed in a collaborative effort between ITU-T SG16 and ISO / IEC JTC 1 / SC29 (MPEG), concretely by the Joint Video Expert Team (JVET). This team has also started an investigation into VCM, with the focus being on developing guidelines on optimizing encoders and receiving systems.
[0006] An important part of a video bitstream is metadata, which is data in the bitstream 25 that contain side information and does not affect the values of the decoded video samples. Metadata can take different form, with one example being Supplemental Enhancement Information (SEI) messages. Examples of different SEI messages are the annotated regions SEI, which can be used to signal information such as bounding boxes and labels, and the decodedpicture hash SEI, which can be used to signal the hash value of a decoded picture so that the decoder can verify its reconstruction.
[0007] A video sequence (or video for short) is made up of a sequence of pictures (which are also known as “images” or “frames”). A video bitstream (or just “bitstream” for short) 5 comprises pictures in compressed format, so-called coded pictures. The decoding order is the order of the coded picture in the bitstream as well as the order in which the coded pictures are decoded by a decoder. The decoder decodes coded pictures into decoded pictures (which are also referred to as “reconstructed pictures”). The output order of decoded pictures is conveyed in the bitstream, and that is the order in which decoded pictures are output from the decoder. If the 10 decoded pictures are displayed, they are commonly displayed in output order. The decoding order and output order may or may not be the same order. When the orders differ, the decoder needs to buffer some decoded pictures to ensure that all pictures are output in output order. In some video coding formats, such as H.264, HEVC and VVC, the output order is conveyed in the bitstream using picture order count (POC) information. Each coded picture is associated with a 15 POC value, and the decoded pictures are output in increasing POC value order.
[0008] In general, the goal for the encoder is to find an encoding of the data with a low distortion. For traditional video encoding, summed square error is often used as a distortion metric. If the luminance (brightness) of the original image is denoted ^^^^^^(^^, ^^), and theluminance (brightness) of the encoded and then decoded image is denoted ^^ௗ^^(^^, ^^), the SSE is20 equal to ேି^ ெି^ =^ ^ ^^ − ^^ ^^ଶ ^^^^ ^.
[0009] to encode the same picture. As a simple example, one way to encode the picture may to divide it into 16x16 sample blocks and encode each block with the help of previously encoded block. Another way 25 may be to divide it into smaller blocks, such as 8x8 sample blocks.
[0010] When two or more ways to encode a picture is available, it is possible to select which one is better using rate-distortion optimization. This is typically done by calculating a cost for each coding choice, where the cost is based on both the distortion (SSE) and the rate (thenumber of bits that it takes to encode the picture). As an example, if ^^^^௫ (^^, ^^) is the decodedpicture obtained when encoding using 16x16 blocks, and ^^ ௫଼(^^, ^^) is the decoded pictureobtained when encoding using 8x8 blocks, we get the two distortions: ^^^^^^^^௫^^ = ^^^^^^ ^^^^^௫^^(^^, ^^), ^^^^^^(^^, ^^)^for the 16x16 case, and^^^^^^଼௫଼ = ^^^^^^ ^^^ ௫଼(^^, ^^), ^^^^^^(^^, ^^)^for the 8x8 case. Assume the picture with 16x16 blocksand ^^ ௫଼ bits to encode thethe following two costs arecalculated: ^^^^^^^^^^௫^^(^^) = ^^^^^^^^௫^^ + ^^ ^^^^௫^^^^^^^^^^଼௫଼(^^) = ^^^^^^଼௫଼ + ^^ ^^ ௫଼
[0011] The best way to encode the picture is then the one associated with the smallestcost. Regardless of the value of ^^, it is easy to see that if the rates are the same (^^^^௫^ = ^^ ௫଼), itis better to take the one with smaller distortion. Likewise, if the distortions are the same(^^^^^^^^௫^^ = ^^^^^^଼௫଼), it is better to take the one with the smaller rate. The value of ^^ is insteadimportant when one way of encoding has both a higher distortion and a higher rate. If it is more important with a low bit rate than a low distortion, a high ^^ should be selected, since bit rate will influence the cost more than if a smaller ^^ is chosen. Likewise, if it is more important with a low distortion, a low ^^ should be selected, since bit rate will influence the cost less than if a higher ^^ is chosen. Typically, in video coding, a lambda is chosen based on the quantization parameter (QP) used. A high QP will give a high lambda and vice versa.
[0012] This rate / distortion way of encoding is often carried out on individual blocks. Assume that we have an 8x8 block of samples that we want to encode. We can then often encode it using prediction from a neighboring block that has already been encoded, such as the block above. This is an example of intra prediction. It is also often possible to encode the block using prediction from a block in the same position in a previously encoded picture. This is an example of inter prediction. To decide which one to select, rate / distortion optimization can be used here: ^^^^^^^^^^௧^^(^^) = ^^^^^^^^௧^^ + ^^ ^^^^௧^^^^^^^^^^^^௧^^(^^) = ^^^^^^^^௧^^ + ^^ ^^^^௧^^
[0013] Here, the number of bits ^^^^௧^^may be due to selection of which block to predict from (above, left), the angle of prediction (straight down or 45 degrees) etc. Likewise, the number of bits ^^^^௧^^may be due to selecting which previously encoded picture to predict from (one frame back or two frames back) and where in that frame (two pixels to the left, one pixel 5 up). The sum squared errors ^^^^^^^^௧^^and ^^^^^^^^௧^^are now the corresponding prediction errors calculated in the 8x8 area.
[0014] Video compression for machines
[0015] One common way to approach encoding videos for machine vision purposes is the “analyze-then-compress” paradigm, in which a video is first analyzed and then the information 10 obtained from this analysis is then used to guide the encoding of the video. For example, the pictures of the video are fed into an analyzer, which can, for example, implement an object detection algorithm. This algorithm produces guiding information, which can for example be a list of objects, enumerating objects in each picture and indicating where in each picture objects are. The same pictures are entered into the encoder which also takes the guiding information as 15 input. In the example, based on the positions of the objects, the encoder adjusts the quantization parameter (QP) values to use for the coding tree units (CTU) of the picture. The QP value can be seen as a proxy for the quality of a video or picture, with a low QP value corresponding to a high bit rate and high quality, and a high QP value corresponding to a low bit rate and low quality.
[0016] This approach has been very successful for compression of video where the 20 machine vision task at the consumer end is object detection or tracking. The analyzer can then find where objects (or object-like regions) are, and make the encoder encode these regions with higher quality, and all other regions of the picture with lower quality. In this way, the overall bit rate can be decreased without sacrificing recognition ability of the objects in the picture, since these have been coded at full fidelity (see, e.g., reference [1]). 25
[0017] SLAM and SfM
[0018] Simultaneous localization and mapping (SLAM) and structure-from-motion (SfM) are computer vision techniques where a 3D model is constructed from a set of pictures of a scene taken with a single video camera, as opposed to having several cameras taking simultaneous pictures of the scene. Such techniques may include the following steps:
[0019] (1) Picture culling: Some pictures from the video may be discarded due to motion blur, or from being too similar to previous pictures (note that no 3D information can be gleaned from the video if neither the camera nor the object moves); picture culling methods are typically heuristic and hand-tuned for specific SLAM and SfM methods; 5
[0020] (2) Feature detection: Feature points that are easy to find across images (such as the corner of a table) are identified in the pictures; the detected features are also described using a feature descriptor (a wide range of feature detection and feature description methods are available in the art);
[0021] (3) Correspondence finding: Feature points from different pictures are matched, 10 e.g., the same table corner is identified as feature point 137 in picture 1 and feature point 536 in picture 2 (a wide range of feature matching methods are available in the art); and
[0022] (4) 3D reconstruction / mapping. Photogrammetric algorithms are used to reconstruct the 3D positions of the feature points, often these methods include outlier rejection (mismatched points) using, for example, Random sample consensus (RANSAC). 15
[0023] An example of a complete SfM pipeline is described in reference [2].
[0024] SLAM and SfM are typically referring to two variations of the same theme. In visual SLAM, the objective is to perform online / real-time localization of a device while also creating a map of the environment from the incoming video. Thus, a SLAM algorithm can typically only use pictures from the video up to the current point in time, and not future pictures. 20 In SfM, the objective is typically to create a high-quality 3D model of a scene or an object from a number of images captured from a scene, thus being typically an offline or non-real-time process. It may operate on a set of pictures that have already been culled, so the first point ‘picture culling’ above may already have taken place. Furthermore, the algorithm is typically allowed to use all of the pictures at once, so unseen “future” pictures may not be an issue. In this 25 disclosure we will describe video compression that can be applied to both of these techniques, and we will use the SLAM and SfM terms interchangeably. SUMMARY
[0025] Certain challenges presently exist. For instance, one problem with previous solutions for compressing video is that they may not be suitable for video that is going to be usedfor SfM. As an example, if the task is to create a 3D reconstruction of a forest, it is not going to help to try to identify objects such as trees and code them well, since the entire scene is full of trees.
[0026] Accordingly, in one aspect there is provided a method for encoding a video 5 produced by a camera, the video comprising an ordered set of pictures, the ordered set of pictures comprising a first picture and a second picture. The method comprises obtaining information about the first picture. The method also comprises using the obtained information about the first picture, determining whether a high-quality encoding condition is satisfied for the first picture. The method also comprises, as a result of determining that the high-quality encoding condition is 10 satisfied for the first picture, causing an encoder to encode the first picture at a first quality level or encoding the first picture at the first quality level. The method further comprises obtaining information about the second picture; using the obtained information about the second picture, determining whether the high-quality encoding condition is satisfied for the second picture; and, as a result of determining that the high-quality encoding condition is not satisfied for the second 15 picture, causing the encoder to encode the second picture at a second quality level that is less than the first quality level, causing the encoder to refrain from encoding the second picture, encoding the second picture at the second quality level, or refraining from encoding the second picture. The information about the first picture comprises: motion information indicating an amount by which the camera has moved relative to an object-of-interest, OOI, in the first picture 20 since a prior point in time; and / or blur value indicating a degree to which the first picture is blurry.
[0027] In another aspect there is provided a method for encoding a video produced by a camera, the video comprising an ordered set of pictures. The method comprises analyzing a plurality of the pictures in the set of pictures. The method also comprises, based on the analysis 25 of the plurality of pictures, performing a group of picture (GOP) size setting process, wherein the GOP size setting process includes: setting a GOP size to a first value for a first group of the plurality of pictures; and setting the GOP size to a second value for a second group of the plurality of pictures, wherein the second value is different than the first value. The plurality of pictures comprises a first picture and a second picture and analyzing the plurality of the pictures 30 comprises: locating an OOI in the first picture and locating the OOI in the second picture.
[0028] In another aspect there is provided a method for producing a 3D model of an OOI. The method comprises obtaining a reconstructed picture that was reconstructed using a coded version of the picture, wherein the coded version of the picture was produced using a selected level of quantization and the reconstructed picture includes the OOI. The method also comprises 5 obtaining information identifying the selected level of quantization. The method further comprises, based on the information identifying the selected level of quantization, determining whether or not to use the reconstructed picture in a modeling process for producing the 3D model of the OOI.
[0029] In another aspect there is provided a computer program comprising instructions 10 which when executed by processing circuitry of an apparatus causes the apparatus to perform any of the methods disclosed herein. In one embodiment, there is provided a carrier containing the computer program wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium. In another aspect there is provided an apparatus that is configured to perform the methods disclosed herein. The apparatus may include memory 15 and processing circuitry coupled to the memory.
[0030] An advantage of the embodiments disclosed herein is that they enable varying the quality of a video at a sub-picture level, which saves bit rate without affecting the SfM end result. BRIEF DESCRIPTION OF THE DRAWINGS 20
[0031] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.
[0032] FIG.1 illustrates a system according to an embodiment.
[0033] FIG.2 is a block diagram of an encoder according to an embodiment.
[0034] FIG.3 is a block diagram of a decoder according to an embodiment. 25
[0035] FIG.4 is a flowchart illustrating a process according to an embodiment.
[0036] FIG.5 is a flowchart illustrating a process according to an embodiment.
[0037] FIG.6 is a flowchart illustrating a process according to an embodiment.
[0038] FIG.7 is a block diagram of an apparatus according to an embodiment.DETAILED DESCRIPTION
[0039] FIG.1 illustrates a system 100 according to an embodiment. System 100 includes an encoder 102 and a decoder 104. In some embodiments, encoder 102 is in communication with decoder 104 via a network 110 (e.g., the Internet or other network). 5 Encoder 102 encodes a source video 101, produced by camera 199, into a bitstream comprising an encoded video and may transmit the bitstream to decoder 104 via network 110. In some embodiments, encoder 102 is not in communication with decoder 104, and, in such an embodiment, rather than transmitting bitstream to decoder 104, the bitstream is stored in a data storage unit 190 and decoder 104 retrieves the bitstream from data storage unit 190. Decoder 104 10 decodes the pictures included in the encoded video to produce video data for display and / or further image processing (e.g. a machine vision task, such as generating a 3D of an object-of- interest (OOI) in the pictures of the video). Accordingly, decoder 104 may be part of a device 103 having an image processor 105 and / or a display 106. The image processor 105 may perform machine vision tasks on the decoded pictures. One such machine vision task may belocating an 15 object in the picture and creating a 3D model of the object. The device 103 may be a mobile device, a set-top device, a head-mounted display, or any other device.
[0040] Additionally, as shown in FIG.1, an analyzer 190 is in communication with the encoder 102. Analyzer 190 functions to analyze the pictures of the video and, based on the analysis, provide information to encoder 102 that will influence how encoder 102 will encode the 20 pictures. While in the embodiment shown in FIG.1, analyzer is shown as being separate from encoder 102, in some embodiments, analyzer 190 is a component of encoder 102.
[0041] FIG.2 illustrates functional components of encoder 102 according to some embodiments. It should be noted that encoders may be implemented differently so implementation other than this specific example can be used. Encoder 102 employs a subtractor 25 241 to produce a residual block which is the difference in sample values between an input block and a prediction block (i.e., the output of a selector 251, which is either an inter prediction block output by an inter predictor 250 (a.k.a., motion compensator) or an intra prediction block output by an intra predictor 249). Then a forward transform 242 is performed on the residual block to produce a transformed block comprising transform coefficients. A quantization unit 243 30 quantizes the transform coefficients based on a quantization parameter (QP) value (e.g., a QPvalue obtained based on a picture QP value for the picture in which the input block is a part and a block specific QP offset value for the input block), thereby producing quantized transform coefficients which are then encoded into the bitstream by encoder 244 (e.g., an entropy encoder) and the bitstream with the encoded transform coefficients is output from encoder 102. Next, 5 encoder 102 uses the quantized transform coefficients to produce a reconstructed block. This is done by first applying inverse quantization 245 and inverse transform 246 to the transform coefficients to produce a reconstructed residual block and using an adder 247 to add the prediction block to the reconstructed residual block, thereby producing the reconstructed block, which is stored in the reconstruction picture buffer (RPB) 266. Loop filtering by a loop filter 10 (LF) stage 267 is applied and the final decoded picture is stored in a decoded picture buffer (DPB) 268, where it can then be used by the inter predictor 250 to produce an inter prediction block for the next picture to be processed. LF stage 267 may include three sub-stages: i) a deblocking filter, ii) a sample adaptive offset (SAO) filter, and iii) an Adaptive Loop Filter (ALF). 15
[0042] FIG.3 illustrates functional components of decoder 104 according to some embodiments. It should be noted that decoder 104 may be implemented differently so implementations other than this specific example can be used. Decoder 104 includes a decoder module 361 (e.g., an entropy decoder) that decodes from the bitstream quantized transform coefficient values of a block. Decoder 104 also includes a reconstruction stage 398 in which 20 the quantized transform coefficient values are subject to an inverse quantization process 362 and inverse transform process 363 to produce a residual block. This residual block is input to adder 364 that adds the residual block and a prediction block output from selector 390 to form a reconstructed block. Selector 390 either selects to output an inter prediction block or an intra prediction block. The reconstructed block is stored in a RPB 365. The inter prediction block is 25 generated by the inter prediction module 350 and the intra prediction block is generated by the intra prediction module 369. Following the reconstruction stage 398, a loop filter stage 367 applies loop filtering to the reconstructed picture or block and the final decoded, reconstructed picture may be stored in a decoded picture buffer (DPB) 368 and output to image processor (IP) 105 (also may be output to display 106). Pictures are stored in the DPB for two primary 30 reasons: 1) to wait for picture output and 2) to be used for reference when decoding future pictures.
[0043] As described above, a challenge presently exists because existing solutions for compressing video may not be suitable for video that is going to be used for SfM. As an example, if the task is to create a 3D reconstruction of a forest, it is not going to help to try to identify objects such as trees and code them well, since the entire scene is full of trees. 5
[0044] Accordingly, when a video is to be used by image processor 105 to perform a machine vision task, such as, for example, generate a 3D model of an OOI capture in pictures of the video, analyzer 190 is employed to provide information to encoder 102 so that encoder can compress at high-quality the minimum number of pictures needed by image processor 105. By knowing what the image processor 105 is going to need in high quality, and what can be coded at 10 lower quality, it is possible to save bits without lowering the quality of the resulting 3D model. In short, this disclosure discloses unique ways of adapting the encoder so that the resulting bit rate is lower without negatively affecting the performance of machine vision (e.g., SfM) methods that are applied on the decoded video by image processor 105. As an example, analyzer 190 may determine which pictures will not be good to extract features from and cause encoder 102 to 15 encode these pictures at lower quality using fewer bits. In one embodiment, the analyzer performs a step similar or equivalent to the picture culling step that the SfM method performs at the decoder side.
[0045] As an example, if the video consists of 100 frames, but the frame rate is so high so that sufficient motion in the scene only happens every 10th frame, then the analyzer can 20 identify this and guide the encoder to encode frame 0 at high quality, frame 1 through 9 at low quality, frame 10 at high quality, frame 11 through 19 at low quality and so on. This way, the bit rate can be lowered significantly without affecting the end result since the picture culling process in the SfM method after the decoder will anyway cull frame 1 through 9, 11 through 19 etc.
[0046] In one example, the following steps are performed: 25
[0047] (1) first, a set of consecutive pictures are determined to have the properties that one picture in the set provides sufficient information for an SfM method. The property may be that there is low motion among the pictures in the set;
[0048] (2) second, one picture in the set is selected to be encoded at high quality; and
[0049] (3) third, the one picture is encoded at high quality and the other pictures in the set are encoded at low quality.
[0050] As another example, the analyzer may estimate the 3D motion between pictures in the video and guide the encoder to keep a low quality until sufficient motion has taken place. 5
[0051] Blur
[0052] As another example, the analyzer may estimate the amount of blur (due to for example motion or the camera not being focused). If the amount of blur exceeds a threshold, the encoder may be guided to encode these pictures at a lower quality than other pictures. In one embodiment, this blur threshold may be set adaptively, so that if the entire sequence is blurry, the 10 threshold will be lowered until at least some images pass the threshold.
[0053] GOP Structure
[0054] In some embodiments, it can be advantageous to align the lowering of quality with the group-of-picture (GOP) structure of the transmitted video. As an example, consider the random-access configuration used in JVET in the common test conditions. This structure 15 inherently has very different quality for the different pictures. Pictures whose POC is divisible by 16 (such as 16, 32 and 64) are encoded with a very high quality (e.g., with low QP), whereas pictures with an odd POC are encoded with a low quality (e.g., a higher QP). Assume that all subsequent pictures after POC 0 have been deemed to not have enough 3D motion and / or too much blur, but that picture 63 is the first with enough motion or enough clarity. Then a naïve 20 implementation may encode picture 63 with a very high quality. However, due to the GOP structure, picture 63 will not be used for prediction, and therefore it will be wasteful to spend a lot of bits on this picture. In this case, it may be desirable to encode picture 63 at low quality and instead encode picture 64 at high quality, which anyway will be encoded with a high quality due to its POC. 25
[0055] Shorten GOP size
[0056] In another embodiment, the GOP structure may be shortened in order to match frames selected for high-quality encoding (e.g., favored by the picture culling stage). As an example, given a base GOP size of 64, frames 0, 64 and 128 will be encoded at the highest quality; frames 32 and 96 are of the second highest quality; frames 16, 48, 80 and 112 of thethird highest quality; frames 8, 24, 50, 56, 72, 88, 104 and 120 of the fourth highest quality; and so forth. Assume that frames 0, 64 and 72 are selected for high quality encoding (they are not blurry and have sufficient motion), but that frames 73 through 128 are all blurry. In this scenario, it may be advantageous to start with a GOP of size 64, implicitly giving lots of bits to frame 0 5 and 64, and follow that with a GOP of size 8, giving a lot of bits to frame 72. By varying the GOP size this way, it is possible for the encoder to make sure that pictures favored by the picture culling stage in the decoder are encoded in high quality. After frame 72, the encoder may go back to a GOP structure of 64 frames, i.e., placing a lot of bits again on frame 136. Alternatively, it could use another GOP of 8 frames, one of 16 frames and one of 32 frames so that frame 128 is 10 emphasized again. This varying of the GOP size is also easily detectable by the decoder by examining the bitstream.
[0057] Adapt GOP Size based on Relative Motion between Camera and OOI
[0058] In another embodiment the GOP structure may be shortened in order to adapt to the motion of the camera 199. As an example, in one video, the camera is circling a stationary 15 OOI at a constant speed. Assume that the speed is such that the camera has not moved enough until frame 64 frames. In this case, a constant GOP size of 64 may be desirable, since this gives good quality at frames 0, 64, 128, 192, 256. Further assume, however, that in another video the camera movement speeds up so that it is twice as fast after frame 128. In this case it may be desirable to have a GOP size of 64 in the beginning and a GOP size of 32 in the end, so that the 20 quality is good at frames 0, 64, 128 (where the GOP size is 64) and then at frame 160, 192, 224 and 256. Hence the GOP size can vary with how much the camera baseline changes between pictures.
[0059] Adapt GOP size based on Frame Rate
[0060] In one variant, the GOP size varies depending on the frame rate of the encoded 25 video. As an example, an encoded video of 1 fps may be encoded with forward prediction only in which the output order and decoding order is the same order. An encoded video of 5 fps may use a GOP size of 2 where every second picture P is output before a picture A, wherein picture P is decoded after picture A. An encoded video of higher fps may use an even larger GOP size of for example 4 or 8. In other words, the higher the encoded frame rate, the larger or longer the 30 GOP size is used. In applications where the video is dual-purpose, i.e., is to be used both for SfMand for human consumption, it may be desirable to interpolate the video at the decoder to a constant frame rate of, say, 30 frames per second, possibly using a neural network.
[0061] Adaptive GOP Structure
[0062] In another embodiment, the encoder first analyzes the video, and then based on 5 the analysis creates an adaptive GOP structure so that pictures that are unlikely to be culled are encoded at high quality whereas pictures that will likely not be used in the reconstruction will be coded at low quality. The GOP structure may vary from one GOP to the next and may have odd sizes such as 9, 12 or 17 frames. The video is then encoded according to this dynamic GOP structure. 10
[0063] Use QP at decoder side
[0064] In another embodiment, image processor 105 uses the quantization parameter (QP) as an additional input for selecting which pictures to use for the reconstruction. As an example, if the encoder has encoded the video with a fixed QP per picture, the picture culling algorithm can, in addition to favor the preserving of pictures that are non-blurry and images that 15 have sufficient motion, also favor pictures that have a low QP, i.e., high quality.
[0065] If the encoder has encoded the video with a varying QP over the picture, image processor 105 can favor pictures that have a low average QP. The average QP can be calculated by determining the QP for every luma sample in the picture and then summing all these QP and dividing by the number of samples in the picture. In yet another embodiment, image processor 20 105 can favor pictures with a high bit rate. The decoder may determine the number of bits or bytes for a decoded picture, and favor pictures with a high number.
[0066] SEI Message
[0067] In another embodiment, the analyzer will guide the encoder to signal to image processor 105 which pictures have been selected by analyzer to be encoded with high quality. 25 This signaling can for instance be done using an SEI message. In yet another embodiment, image processor 105 may not get any such side information but may anyway be able to cull away the right pictures since they have such much lower quality, due to the encoder having lowered the bit rate for them.
[0068] Motion vs. Texture
[0069] In one embodiment, an emphasis is put on preserving the correct motion of objects in the scene, but less emphasis on preserving the correct texture, such as the structure of a painted wall. The motivation for this is that for SfM / SLAM, the end result is a set of 3D points. To get an accurate result, i.e., accurate positions of the resulting 3D points, it is important that the SfM / SLAM method is able to identify certain key points, such as a corner of an object, in more than one picture. It is also important that the movements of these key points are correctly represented in the decoded pictures.
[0070] In order to correctly identify key points, it is important to preserve the texture of the picture around the key point, i.e., it is important to preserve the sample values around a key point with high fidelity. Preserving the texture is costly in terms of bits, since this means that transform coefficients (and also other variables such as intra prediction angle, block partitioning information etc.) need to be sent with high fidelity. In order to correctly identify movement, it is important that the motion of pixels is preserved between frames. This also costs bits since motion information such as motion vectors may need to be sent with higher fidelity (e.g., 1 / 8thresolution rather than 1 / 4thresolution) to preserve the motion better.
[0071] When the compressed material is going to be used for SfM / SLAM at the decoder side, there is a trade-off between how many bits to spend on texture and how many bits to spend on motion. As an example, if the texture quality is already good enough to reliably identify the same key point in two different pictures, spending extra bits on increasing texture quality is not going to increase the quality of the SfM output. Likewise, if motion information is not of sufficient accuracy, the quality of the SfM output may suffer. Therefore, in one embodiment a trade-off between motion information and texture information is allowed.
[0072] For instance, in one embodiment, the way to calculate the cost function is changed. Instead of using: ^^^^^^^^^^௧^^(^^) = ^^^^^^^^௧^^ + ^^ ^^^^௧^^^^one may use ^^^^^^^^^^௧^^(^^) = ^^^^^^^^௧^^ + ^^ ^^^^௧^^^^ .
[0073] This has the effect of making it cheaper for the encoder to spend bits on inter modes, i.e., ^^^^௧^^can be twice as high and the inter way of encoding will still be selected. This in turn means that it is possible to encode motion vectors at a higher resolution, changing the trade-off between texture information and motion information. It should be noted that the value 5 0.5 is just an example, and that other values are possible. In summary, this has the effect of making the motion vectors “cheaper” than other bits. This will make the encoder spend more bits on motion vectors, and hence preserve the motion better, at the expense of texture.
[0074] This may be done sporadically and not continuously, so that the SLAM or SfM algorithm still receives good quality information about the texture in the static regions. 10
[0075] In one example, the rate-distortion calculation is done in the encoder by evaluating a number of different modes to encode a block B and selecting the mode of encoding B that minimizes the value of a cost function. The cost function may be in the form of SSE+^^*r, where ^^ is a derived rate-distortion trade-off value for the block, SSE is a distortion value for the decoded or reconstructed block B, and r is a bit count value of the encoded block B.15
[0076] In one embodiment, a criteria ^^ + (^^^ ∗ ^^^) + (^^^ ∗ ^^^) is used instead whereD is the distortion value for the decoded or reconstructed block B, ^^^and ^^^are different derived rate-distortion trade-off values with ^^^being greaterbm is a bit count value of the bits representing the motion information for the coded block, and bo is a bit count value of the bits representing non-motion information for the coded block. 20
[0077] Feature points
[0078] In another embodiment, the analyzer is doing a feature analysis on the input video. The outcome is signaled to the encoder, which then proceeds to encode areas containing feature points in higher quality than areas not containing feature points. This may lower the bit rate and provide similar or better end SfM results if the feature detection method in the SfM 25 method on the decoder side chooses feature point that are primarily positioned in the areas that were coded in higher quality. In one variant, the analyzer determines areas that are important based on the feature analysis and provides spatial information of the areas to the encoder.
[0079] FIG.4 is a flow chart illustrating a process 400, according to an embodiment, for encoding a video produced by camera 199, the video comprising an ordered set of pictures, theordered set of pictures comprising a first picture and a second picture. Process 400 may begin in step s402.
[0080] Step s402 comprises obtaining information about the first picture.
[0081] Step s404 comprises using the obtained information about the first picture, 5 determining whether a high-quality encoding condition is satisfied for the first picture.
[0082] Step s406 comprises as a result of determining that the high-quality encoding condition is satisfied for the first picture, causing an encoder to encode the first picture at a first quality level or encoding the first picture at the first quality level.
[0083] Step s406 comprises obtaining information about the second picture. 10
[0084] Step s406 comprises using the obtained information about the second picture, determining whether the high-quality encoding condition is satisfied for the second picture.
[0085] Step s408 comprises, as a result of determining that the high-quality encoding condition is not satisfied for the second picture, causing the encoder to encode the second picture at a second quality level that is less than the first quality level, causing the encoder to refrain 15 from encoding the second picture, encoding the second picture at the second quality level, or refraining from encoding the second picture.
[0086] The information about the first picture comprises: motion information indicating an amount by which the camera has moved relative to an object-of-interest, OOI, in the first picture since a prior point in time and / or blur value indicating a degree to which the first picture 20 is blurry.
[0087] In some embodiments, a group-of-picture, GOP, size and a first picture order count, POC, value for the first picture is also used to determine whether the high-quality encoding condition is satisfied for the first picture.
[0088] In some embodiments, the method further comprises: analyzing a plurality of the 25 pictures in the set of pictures; and, based on the analysis of the plurality of pictures: setting a group of picture, GOP, size to a first value for a first group of the plurality of pictures; and setting the GOP size to a second value for a second group of the plurality of pictures, wherein the second value is different than the first value. The first group of pictures and second group of pictures are disjoint.
[0089] In some embodiments, the method further comprises causing the encoder to send a message indicating a subset of the ordered set of pictures or sending the message indicating the subset of the ordered set of pictures, and each picture included in the indicated subset of pictures was encoded at the first quality level. 5
[0090] In some embodiments, obtaining information about the first picture comprises: locating the OOI in a preceding picture that precedes the first picture in the ordered set of pictures; locating the OOI in the first picture; and obtaining the motion information based on the location of the OOI in the preceding picture and the location of the OOI in the first picture.
[0091] In some embodiments, obtaining information about the first picture comprises: 10 determining an orientation of the OOI in a preceding picture that precedes the first picture in the ordered set of pictures; determining an orientation the OOI in the first picture; and obtaining the motion information based on the orientation of the OOI in the preceding picture and the orientation of the OOI in the first picture.
[0092] In some embodiments, the information about the first picture comprises the blur15 value indicating a degree to which the first picture is blurry, and determining whether a high- quality encoding condition is satisfied for the first picture comprises determining whether the blue value is less than a blur threshold.
[0093] In some embodiments, the method further comprises setting the blur threshold based on an analysis of the blurriness of one or more of the pictures in the ordered set of pictures. 20
[0094] FIG.5 is a flow chart illustrating a process 500, according to an embodiment, for encoding a video produced by camera 199, the video comprising an ordered set of pictures, the ordered set of pictures comprising a first picture and a second picture. Process 500 may begin in step s502.
[0095] Step s502 comprises analyzing a plurality of the pictures in the set of pictures. 25
[0096] Step s504 comprises, based on the analysis of the plurality of pictures, performing a group of picture (GOP) size setting process, wherein the GOP size setting process comprises setting a GOP size to a first value for a first group of the plurality of pictures and setting the GOP size to a second value for a second group of the plurality of pictures, wherein the second value is different than the first value. The plurality of pictures comprises a first picture and a secondpicture and analyzing the plurality of the pictures comprises locating an OOI in the first picture and locating the OOI in the second picture.
[0097] In some embodiments, analyzing the plurality of the pictures further comprises: determining an orientation of the OOI in the first picture; determining an orientation the OOI in 5 the second picture; and obtaining motion information based on the orientation of the OOI in the first picture and the orientation of the OOI in the second picture. In some embodiments, performing the GOP size setting process based on the analysis of the plurality of pictures comprises performing the GOP size setting process using the obtained motion information.
[0098] In some embodiments, performing the GOP size setting process based on the 10 analysis comprises performing the GOP size setting process using the motion information.
[0099] FIG.6 is a flow chart illustrating a process 600, according to an embodiment, for producing a 3D model of an OOI. Process 600 may begin in step s602.
[0100] Step s602 comprises obtaining a reconstructed picture that was reconstructed using a coded version of the picture, wherein the coded version of the picture was produced 15 using a selected level of quantization and the reconstructed picture includes the OOI.
[0101] Step s604 comprises obtaining information identifying the selected level of quantization. In some embodiments, obtaining the information identifying the selected level of quantization comprises obtaining the information from a bitstream comprising the coded version of the picture. 20
[0102] Step s604 comprises, based on the information identifying the selected level of quantization, determining whether or not to use the reconstructed picture in a modeling process for producing the 3D model of the OOI.
[0103] FIG.7 is a block diagram of an apparatus 700 for implementing encoder 102 or decoder 104 according to some embodiments. When apparatus 700 implements encoder 102, 25 apparatus 700 may be referred to as an encoder apparatus, when apparatus 700 implements decoder 104, apparatus 700 may be referred to as a decoder apparatus. As shown in FIG.7, apparatus 700 may comprise: processing circuitry (PC) 702, which may include one or more processors (P) 755 (e.g., one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gatearrays (FPGAs), and the like), which processors may be co-located in a single housing or in a single data center or may be geographically distributed (i.e., encoder apparatus 700 may be a distributed computing apparatus); at least one network interface 748 (e.g., a physical interface or air interface) comprising a transmitter (Tx) 745 and a receiver (Rx) 747 for enabling apparatus 5 700 to transmit data to and receive data from other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network) to which network interface 748 is connected (physically or wirelessly) (e.g., network interface 748 may be coupled to an antenna arrangement comprising one or more antennas for enabling encoder apparatus 700 to wirelessly transmit / receive data); and a storage unit (a.k.a., “data storage system”) 708, which may include one or more non- 10 volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 702 includes a programmable processor, a computer readable storage medium (CRSM) 742 may be provided. CRSM 742 may store a computer program (CP) 743 comprising computer readable instructions (CRI) 744. CRSM 742 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, 15 flash memory), and the like. In some embodiments, the CRI 744 of computer program 743 is configured such that when executed by PC 702, the CRI causes encoder apparatus 700 to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, encoder apparatus 700 may be configured to perform steps described herein without the need for code. That is, for example, PC 702 may consist merely of one or more 20 ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.
[0104] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above-described exemplary 25 embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
[0105] As used herein transmitting a message “to” or “toward” an intended recipient encompasses transmitting the message directly to the intended recipient or transmitting the 30 message indirectly to the intended recipient (i.e., one or more other nodes are used to relay the message from the source node to the intended recipient). Likewise, as used herein receiving amessage “from” a sender encompasses receiving the message directly from the sender or indirectly from the sender (i.e., one or more nodes are used to relay the message from the sender to the receiving node). Further, as used herein “a” means “at least one” or “one or more.”
[0106] Additionally, while the processes described above and illustrated in the drawings 5 are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.
Claims
CLAIMS 1. A method (400) for encoding a video produced by a camera (199), the video comprising an ordered set of pictures, the ordered set of pictures comprising a first picture and a 5 second picture, the method comprising: obtaining (s402) information about the first picture; using (s404) the obtained information about the first picture, determining whether a high- quality encoding condition is satisfied for the first picture; as a result of determining that the high-quality encoding condition is satisfied for the first picture, causing (s406) an encoder to encode the first picture at a first quality level or encoding the first picture at the first quality level; obtaining (s408) information about the second picture; using (s410) the obtained information about the second picture, determining whether the high-quality encoding condition is satisfied for the second picture; and as a result of determining that the high-quality encoding condition is not satisfied for the second picture, causing the encoder to encode the second picture at a second quality level that is less than the first quality level, causing the encoder to refrain from encoding the second picture, encoding the second picture at the second quality level, or refraining from encoding the second picture (s412), wherein the information about the first picture comprises: motion information indicating an amount by which the camera has moved relative to an object-of-interest, OOI, in the first picture since a prior point in time; and / or a blur value indicating a degree to which the first picture is blurry.
2. The method of claim 1, wherein a group-of-picture, GOP, size and a first picture order count, POC, value for the first picture is also used to determine whether the high-quality encoding condition is satisfied for the first picture.
3. The method of claim 1 or 2, wherein the method further comprises: analyzing a plurality of the pictures in the ordered set of pictures; and based on the analysis of the plurality of pictures:setting a group of picture, GOP, size to a first value for a first group of the plurality of pictures; and setting the GOP size to a second value for a second group of the plurality of pictures, wherein 5 the second value is different than the first value.
4. The method of any one of claims 1-3, wherein the method further comprises causing the encoder to send a message indicating a subset of the ordered set of pictures or sending the message indicating the subset of the ordered set of pictures, and each picture included in the indicated subset of pictures was encoded at the first quality level.
5. The method of claim 4, wherein the message is a Supplemental Enhancement Information, SEI, message.
6. The method of any one of claims 1-5, wherein obtaining information about the first picture comprises: locating the OOI in a preceding picture that precedes the first picture in the ordered set of pictures; locating the OOI in the first picture; and obtaining the motion information based on the location of the OOI in the preceding picture and the location of the OOI in the first picture.
7. The method of any one of claims 1-6, wherein obtaining information about the first picture comprises: determining an orientation of the OOI in a preceding picture that precedes the first picture in the ordered set of pictures; determining an orientation of the OOI in the first picture; and obtaining the motion information based on the orientation of the OOI in the preceding picture and the orientation of the OOI in the first picture.
8. The method of any one of claims 1-7, wherein the information about the first picture comprises the blur value indicating a degree to which the first picture is blurry, and determining whether a high-quality encoding condition is satisfied for the first picture 5 comprises determining whether the blue value is less than a blur threshold.
9. The method of claim 8, wherein the method further comprises setting the blur threshold based on an analysis of the blurriness of one or more of the pictures in the ordered set of pictures.
10. A method (500) for encoding a video produced by a camera (199), the video comprising an ordered set of pictures, the method comprising: analyzing (s502) a plurality of the pictures in the ordered set of pictures; and based on the analysis of the plurality of pictures, performing (s504) a group of picture, GOP, size setting process comprising: setting a GOP size to a first value for a first group of the plurality of pictures; and setting the GOP size to a second value for a second group of the plurality of pictures, wherein the second value is different than the first value, further wherein the plurality of pictures comprises a first picture and a second picture, and analyzing the plurality of the pictures comprises: locating an object-of-interest, OOI, in the first picture; and locating the OOI in the second picture.
11. The method of claim 10, wherein analyzing the plurality of the pictures further comprises: determining an orientation of the OOI in the first picture; determining an orientation the OOI in the second picture; and obtaining motion information based on the orientation of the OOI in the first picture and the orientation of the OOI in the second picture, and performing the GOP size setting process based on the analysis comprises performing the GOP size setting process using the obtained motion information.
12. The method of claim 11, wherein performing the GOP size setting process based on the analysis comprises performing the GOP size setting process using the motion information.
13. A method (600) for producing a three-dimensional, 3D, model of an object-of- 5 interest, OOI, the method comprising: obtaining (s602) a reconstructed picture that was reconstructed using a coded version of the picture, wherein the coded version of the picture was produced using a selected level of quantization and the reconstructed picture includes the OOI; obtaining (s604) information identifying the selected level of quantization; based on the information identifying the selected level of quantization, determining (s605) whether or not to use the reconstructed picture in a modeling process for producing the 3D model of the OOI.
14. The method of claim 13, wherein obtaining the information identifying the selected level of quantization comprises obtaining the information from a bitstream comprising the coded version of the picture.
15. A computer program (743) comprising instructions (744) which when executed by processing circuitry (702) of an apparatus causes the apparatus to perform the method of any one of claims 1-14.
16. A carrier containing the computer program of claim 15, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (742).
17. An apparatus (700) for encoding a video produced by a camera (199), the video comprising an ordered set of pictures, the ordered set of pictures comprising a first picture and a second picture, wherein the apparatus is configured to perform a method (400) comprising: obtaining (s402) information about the first picture; using (s404) the obtained information about the first picture, determining whether a high- quality encoding condition is satisfied for the first picture;as a result of determining that the high-quality encoding condition is satisfied for the first picture, causing (s406) an encoder to encode the first picture at a first quality level or encoding the first picture at the first quality level; obtaining (s408) information about the second picture; 5 using (s410) the obtained information about the second picture, determining whether the high-quality encoding condition is satisfied for the second picture; and as a result of determining that the high-quality encoding condition is not satisfied for the second picture, causing the encoder to encode the second picture at a second quality level that is less than the first quality level, causing the encoder to refrain from encoding the second picture, encoding the second picture at the second quality level, or refraining from encoding the second picture (s412), wherein the information about the first picture comprises: motion information indicating an amount by which the camera has moved relative to an object-of-interest, OOI, in the first picture since a prior point in time; and / or a blur value indicating a degree to which the first picture is blurry.
18. The apparatus of claim 17, wherein a group-of-picture, GOP, size and a first picture order count, POC, value for the first picture is also used to determine whether the high-quality encoding condition is satisfied for the first picture.
19. The apparatus of claim 17 or 18, wherein the method further comprises: analyzing a plurality of the pictures in the ordered set of pictures; and based on the analysis of the plurality of pictures: setting a group of picture, GOP, size to a first value for a first group of the plurality of pictures; and setting the GOP size to a second value for a second group of the plurality of pictures, wherein the second value is different than the first value.
20. The apparatus of any one of claims 17-19, whereinthe method further comprises causing the encoder to send a message indicating a subset of the ordered set of pictures or sending the message indicating the subset of the ordered set of pictures, and each picture included in the indicated subset of pictures was encoded at the first quality 5 level.
21. The apparatus of claim 20, wherein the message is a Supplemental Enhancement Information, SEI, message.
22. The apparatus of any one of claims 17-21, wherein obtaining information about the first picture comprises: locating the OOI in a preceding picture that precedes the first picture in the ordered set of pictures; locating the OOI in the first picture; and obtaining the motion information based on the location of the OOI in the preceding picture and the location of the OOI in the first picture.
23. The apparatus of any one of claims 12-22, wherein obtaining information about the first picture comprises: determining an orientation of the OOI in a preceding picture that precedes the first picture in the ordered set of pictures; determining an orientation of the OOI in the first picture; and obtaining the motion information based on the orientation of the OOI in the preceding picture and the orientation of the OOI in the first picture.
24. The apparatus of any one of claims 17-23, wherein the information about the first picture comprises the blur value indicating a degree to which the first picture is blurry, and determining whether a high-quality encoding condition is satisfied for the first picture comprises determining whether the blue value is less than a blur threshold.
25. The apparatus of claim 24, wherein the method further comprises setting the blur threshold based on an analysis of the blurriness of one or more of the pictures in the ordered set of pictures. 5 26. An apparatus (700) for encoding a video produced by a camera, the video comprising an ordered set of pictures, wherein the apparatus is configured to perform a method (500) comprising: analyzing (s502) a plurality of the pictures in the ordered set of pictures; and based on the analysis of the plurality of pictures, performing (s504) a group of picture, GOP, size setting process comprising: setting a GOP size to a first value for a first group of the plurality of pictures; and setting the GOP size to a second value for a second group of the plurality of pictures, wherein the second value is different than the first value, further wherein the plurality of pictures comprises a first picture and a second picture, and analyzing the plurality of the pictures comprises: locating an object-of-interest, OOI, in the first picture; and locating the OOI in the second picture.
27. The apparatus of claim 26, wherein analyzing the plurality of the pictures further comprises: determining an orientation of the OOI in the first picture; determining an orientation the OOI in the second picture; and obtaining motion information based on the orientation of the OOI in the first picture and the orientation of the OOI in the second picture, and performing the GOP size setting process based on the analysis comprises performing the GOP size setting process using the obtained motion information.
28. The apparatus of claim 26, wherein performing the GOP size setting process based on the analysis comprises performing the GOP size setting process using the motion information.
29. An apparatus (700) for producing a three-dimensional, 3D, model of an object-of- interest, OOI, the apparatus being configured to perform a method (660) comprising: obtaining (s602) a reconstructed picture that was reconstructed using a coded version of the picture, wherein the coded version of the picture was produced using a selected level of 5 quantization and the reconstructed picture includes the OOI; obtaining (s604) information identifying the selected level of quantization; based on the information identifying the selected level of quantization, determining (s605) whether or not to use the reconstructed picture in a modeling process for producing the 3D model of the OOI.
30. The apparatus of claim 29, wherein obtaining the information identifying the selected level of quantization comprises obtaining the information from a bitstream comprising the coded version of the picture.
Citation Information
Patent Citations
Method and encoder system for encoding video
EP3021579B1