Methods for encoding video streams

By processing only the region of interest at full resolution and applying downscaled video processing, the method addresses the high computational and energy demands of high-resolution video, achieving efficient and environmentally friendly encoding.

JP7867938B2Active Publication Date: 2026-06-01AXIS

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
AXIS
Filing Date
2022-10-14
Publication Date
2026-06-01

Smart Images

  • Figure 0007867938000001
    Figure 0007867938000001
  • Figure 0007867938000002
    Figure 0007867938000002
  • Figure 0007867938000003
    Figure 0007867938000003
Patent Text Reader

Abstract

To provide a method for encoding a video stream in a more energy and computational power efficient way.SOLUTION: A method (400) for encoding a video stream comprises: acquiring (S402) pixel data of the video stream having a first resolution; extracting (S406) a crop corresponding to a region of interest from the pixel data of the video stream, the crop having the first resolution; down-scaling (S408) the pixel data of the video stream into a down-scaled video stream having a second resolution lower than the first resolution; up-scaling (S414) the processed down-scaled video stream into an up-scaled video stream having the first resolution; merging (S416) the processed crop and the up-scaled video stream into a merged video stream; and encoding (S420) the merged video stream.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video technology. More specifically, the present invention relates to a method for encoding video streams. The present invention further relates to a video encoding device.

Background Art

[0002] There is an increasing demand for high-resolution video and cameras. It can be a wide-angle camera such as a fisheye, panoramic camera or simply a high-resolution camera. These types of cameras require a high throughput of data, which often requires a high-performance video processing circuit that is power hungry. At the same time, there is a demand for cameras that are environmentally friendly with low energy consumption. Therefore, the use of these high-performance video processing circuits for more mainstream surveillance cameras conflicts with the need to make the technology greener.

[0003] Therefore, there is a need for improvement in providing high-resolution video while keeping energy consumption low at the same time.

[0004] U.S. Patent Application No. 2021105488 discloses a video encoding method. The method includes steps of acquiring a video frame, selecting one or more regions of interest in the video frame, encoding the region of interest or each region of interest at a first resolution, and encoding a base layer, where the base layer includes at least a portion of the video frame that is not included in the region of interest or each region of interest at a second resolution. The first resolution is higher than the second resolution.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

[0006] In view of the above, an object of the present invention is to provide a method for encoding a video stream in a more energy- and computationally power-efficient manner.

[0007] It has been found that downscaling image data that does not represent the region of interest very early in the image processing pipeline, such as immediately after capture, can save computational power and therefore energy consumption. To meet the high-resolution requirements of video, only the region of interest is processed at full resolution.

[0008] According to a first aspect, a method for encoding a video stream is presented. The method includes: obtaining pixel data of a video stream having a first resolution; extracting a crop corresponding to a region of interest from the pixel data of the video stream, wherein the crop has a first resolution; downscaling the pixel data of the video stream to a downscaled video stream having a second resolution lower than the first resolution; processing the downscaled video stream through one or more video processing operations; processing the crop through one or more video processing operations; upscaling the processed downscaled video stream to an upscaled video stream having a first resolution; merging the processed crop and the upscaled video stream to form a merged video stream; and encoding the merged video stream. The one or more video processing operations include at least one of a noise reduction function, a rotation function, an image sensor correction function, an image scaling function, a gamma correction function, an image enhancement function, a color space conversion function, and a chroma subsampling function.

[0009] The term "pixel data" is used herein to mean that the data of a video stream contains information about the pixel values ​​of the video stream.

[0010] The term "resolution" as used herein refers to the number of pixels in each dimension of a video stream, namely, width and height. In light of this, the terms "upscaling" and "downscaling" refer to increasing and decreasing the resolution, respectively.

[0011] The term "region of interest" is used herein to mean a region of a video stream of particular interest. A region of interest may be based, for example, on movement in the scene, or on an object detected in the scene or present in a certain area of ​​the scene.

[0012] It should be noted that the downscaled video stream and crop are preferably processed through the same video processing operation. Such processing simplifies subsequent merging of the processed crop and the upscaled video stream.

[0013] It should be noted that the merged video stream has a higher resolution than the second resolution. Preferably, the merged video stream has the first resolution.

[0014] As mentioned above, a possible related advantage is that only the portion of the video stream representing the region of interest is processed at full resolution. The remainder of the video stream is processed in a downscaled format, thus facilitating a reduction in the computational load on the device processing the video stream. After processing, the downscaled portion of the video stream is upscaled to match the resolution of the crop. This allows the now upscaled video stream to be merged with the processed crop to form a merged video stream. At least one or more portions of the merged video stream belonging to the region of interest will have the original first resolution. Furthermore, this method facilitates saving memory bandwidth by processing the downscaled video stream.

[0015] This method may further include encoding the processed downscaled video stream using a scalable video encoding scheme. The act of encoding the merged video stream may include using the encoded downscaled video stream as a base layer. In other words, the encoded downscaled video stream can be used in encoding a high-resolution merged video stream. The possible associated benefits provided by this are that more computationally and power-efficient encoding of the merged video stream is facilitated. According to this, the high-resolution information of the crop is encoded based on a scaled-up version of the low-resolution video, and for the combined high-resolution video, the reuse of all non-ROI areas, as well as good base predictions for the high-resolution video and the motion vectors of the high-resolution video, may be possible. The final result is a single video that is much smaller than the full high-resolution video, and that single video is full resolution, from which the low-resolution video can be easily extracted.

[0016] The region of interest can be a sub-part of the scene shown by the video stream. In other words, the region of interest can be a subsection of the video stream. A sub-part can be a single area of ​​the scene shown by the video stream. A sub-part can consist of several discontinuous areas of the scene shown by the video stream.

[0017] This method may further include identifying regions of interest in a scene represented by a video stream.

[0018] The act of identifying a region of interest can be performed by processing the acquired pixel data of a video stream through an attention model.

[0019] The term "attention model" as used herein refers to a model for identifying regions of interest from pixel data of a video stream.

[0020] The attention model may be based on an object detection model and at least one of the following: an object classification model, a motion detection model, a user-defined input, or a combination thereof.

[0021] The act of acquiring pixel data from a video stream may include capturing pixel data using an image sensor.

[0022] One or more video processing operations may be at least one of the following: noise reduction, transformation, rotation, privacy mask, image sensor correction, image scaling, gamma correction, image enhancement, color space conversion, chroma subsampling, compression, data storage, and data transmission.

[0023] According to a second embodiment, a video encoding device for encoding a video stream is presented. The video encoding device includes a circuit configured to perform: an acquisition function configured to acquire pixel data of a video stream having a first resolution; an extraction function configured to extract a crop corresponding to a region of interest from the pixel data of the video stream, wherein the crop has a first resolution; a downscaling function configured to downscale the pixel data of the video stream to a downscaled video stream having a second resolution lower than the first resolution; a video processing function configured to process the downscaled video stream through one or more video processing operations of a video processing pipeline, and to process the crop through one or more video processing operations of a video processing pipeline; an upscaling function configured to upscale the processed downscaled video stream to an upscaled video stream having a first resolution; a merging function configured to merge the processed crop and the upscaled video stream to form a merged video stream; and a first encoding function configured to encode the merged video stream. One or more video processing operations include at least one of the following: noise reduction, rotation, image sensor correction, image scaling, gamma correction, image enhancement, color space conversion, and chroma subsampling.

[0024] The circuit may be further configured to perform a second encoding function configured to encode the processed downscaled video stream using a scalable video encoding scheme. The first encoding function may be configured to encode the merged video stream using the encoded processed downscaled video stream as a base layer.

[0025] The circuit may be further configured to perform a region of interest identification function configured to identify a region of interest in a scene represented by a video stream.

[0026] The video processing function may be configured to process both the downscaled video stream and the crop through the same one or more video processing operations of the video processing pipeline.

[0027] The video processing operations of the video processing pipeline may be implemented as a computer software portion running on a general-purpose processor or on a graphics processing unit, a field-programmable gate array, a fixed-function application-specific integrated circuit, or an analog circuit.

[0028] The above-described features and advantages of the first aspect also apply to the second aspect when applicable. For the sake of avoiding excessive repetition, reference is made to the above.

[0029] According to a third aspect, there is provided a non-transitory computer-readable recording medium recording program code configured to implement the method according to the first aspect when executed in a device having processing capabilities.

[0030] The above-described features and advantages of the first and second aspects also apply to the third aspect when applicable. For the sake of avoiding excessive repetition, reference is made to the above.

[0031] Still other objects, features, aspects, and advantages of the present invention will become apparent from the following detailed description of the invention and from the drawings. The same features and advantages described with respect to one aspect are applicable to other aspects unless otherwise specified.

[0032] Next, the above and other aspects of the concept of the present invention will be described in more detail with reference to the accompanying drawings illustrating variations of the invention. The drawings should not be considered to limit the invention to any particular variation, but rather they are used to illustrate and understand the concept of the present invention.

[0033] As shown in the figure, the sizes of the layers and regions are exaggerated for illustrative purposes and are therefore provided to illustrate the schematic structure of the variant of the concept of the present invention. Similar reference numbers refer to similar elements throughout. [Brief explanation of the drawing]

[0034] [Figure 1] This is a schematic diagram illustrating a system for encoding video streams. [Figure 2] As an example, this is a schematic diagram showing a camera equipped with a video encoding device according to the present invention. [Figure 3] As an example, this is a schematic diagram illustrating a video encoding device according to the present invention. [Figure 4] This is a flowchart showing the steps of a method for encoding a video stream according to the present invention. [Modes for carrying out the invention]

[0035] Next, the concept of the present invention will be more fully described below with reference to the accompanying drawings, which show currently preferred variations of the concept of the present invention. However, the concept of the present invention can be implemented in many different forms and should not be construed as being limited to the variations described herein, but rather these variations are provided for completeness and completeness and will fully convey the scope of the concept of the present invention to those skilled in the art.

[0036] Next, a method for encoding a video stream, as well as a video encoding device, will be described with reference to Figures 1 to 4.

[0037] Figure 1 schematically shows a system 100 comprising a video encoding device 300 according to the present invention. The video encoding device 300 will be further described with reference to Figure 3. Figure 1 presents an example of how the video encoding device 300 may be implemented. The video encoding device 300 is connected to a video source 102. The video source 102 provides a video stream to the video encoding device 300. The video source 102 may be an image sensor. The video source 102 may be an image rendering engine.

[0038] The video encoding device 300 is connected to the video processing pipeline 104. The video processing pipeline 104 is configured to perform one or more video processing operations 106A to D. One or more video processing operations 106A to D include one or more of the following: noise reduction, transformation, rotation, privacy masking, image sensor correction, image scaling, gamma correction, image enhancement, color space conversion, chroma subsampling, compression, data storage, and data transmission. The image sensor correction function may include a Bayer filter. The color space conversion function may include conversion between different formats, such as RGB, YUV, and YCbCr. Although the video processing pipeline 104 is shown as being external to the video encoding device 300 in this example, the video processing pipeline 104 may be provided as an integral part of the video encoding device 300. In other words, the video processing pipeline 104 may be part of the video encoding device 300. The connection between the video source 102, the video encoding device 300, and the video processing pipeline 104 is preferably a wired connection.

[0039] Figure 2 shows a camera 200 as an example. The camera 200 comprises a video encoding device 300 according to the present invention. The system 100, as described with respect to Figure 1, can be implemented in a single device, such as the camera 200 shown herein. The camera 200 may be a monitoring or surveillance camera. As will be readily apparent to those skilled in the art, the camera 200 comprises other components necessary for the camera 200 to operate, such as a power supply unit and memory. These are omitted for illustrative purposes.

[0040] Figure 3 schematically shows the video encoding device 300 as described with respect to Figures 1 and 2. The video encoding device 300 includes circuit 302.

[0041] Circuit 302 is any type of circuit 302 comprising a processing unit. Circuit 302 may physically comprise a single circuit module. Alternatively, circuit 302 may be distributed across several circuit modules. The video encoding device 300 may further comprise a transceiver 304 and a memory 306. Circuit 302 is communicatively connected to the transceiver 304 and the memory 306.

[0042] Circuit 302 is configured to provide overall control over the functions and operation of the video encoding device 300. Circuit 302 includes a processing unit, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The processing unit may be configured to execute program code stored in memory 306 in order to perform the functions and operations of circuit 302.

[0043] The transceiver 304 is configured to enable circuit 302 to communicate with other devices. The transceiver 304 is configured to exchange data with circuit 302 over a data bus. Accompanying control lines and an address bus between the transceiver 304 and circuit 302 may also be present.

[0044] Memory 306 may be one or more of a buffer, flash memory, hard drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or another suitable device. In a typical configuration, memory includes non-volatile memory for long-term data storage and volatile memory that serves as system memory for the video encoding device 300. Memory 306 is configured to exchange data with circuit 302 over a data bus. Accompanying control lines and an address bus between memory 306 and circuit 302 may also be present.

[0045] The functions and operations of the video encoding device 300 may be embodied in the form of executable logic routines (e.g., lines of code, software programs, etc.) that are stored in the video encoding device 300's non-temporary computer-readable recording medium (e.g., memory) and executed by the circuit 302 (e.g., using a processor). Furthermore, the functions and operations of the circuit 302 may be a standalone software application or form part of a software application that performs additional tasks related to the circuit 302. The functions and operations described may be considered as being configured to be performed by the corresponding device, such as the method described below with respect to Figure 4. Also, while the functions and operations described may be implemented in software, such functions may also be performed via dedicated hardware or firmware, or some combination of hardware, firmware and / or software. The following functions may be stored in the non-temporary computer-readable recording medium.

[0046] Circuit 302 is configured to perform an acquisition function 308 which is configured to acquire pixel data of a video stream having a first resolution.

[0047] Circuit 302 is further configured to perform an extraction function 310 configured to extract a crop corresponding to a region of interest from the pixel data of the video stream. Thus, the crop contains pixel data representing the region of interest, but does not necessarily contain all the pixel data of each image frame in the video stream. The crop is cut out from the pixel data of the image frames in the video stream. The crop has a first resolution.

[0048] Circuit 302 is further configured to perform a downscaling function 312 configured to downscale the pixel data of a video stream to a downscaled video stream. The downscaled video stream has a second resolution, which is lower than the first resolution.

[0049] The circuit 302 is further configured to execute a video processing function 314 configured to process a downscaled video stream through one or more video processing operations 106A to D of the video processing pipeline 104. The video processing function 314 is further configured to process cropping through one or more video processing operations 106A to D of the video processing pipeline 104. The downscaled video stream and cropping can be processed simultaneously, i.e., in parallel, through the video processing pipeline 104. The downscaled video stream and cropping can be processed sequentially, i.e., alternately, through the video processing pipeline 104.

[0050] Some of the video processing operations 106A to D may be dependent on each other. Therefore, some of the video processing operations 106A to D must be executed sequentially. Some of the video processing operations 106A to D may not be dependent on each other. Therefore, some of the video processing operations 106A to D can be executed in parallel.

[0051] Circuit 302 is further configured to perform an upscaling function 318 configured to upscale the processed downscaled video stream into an upscaled video stream. Preferably, the upscaled video stream has a first resolution.

[0052] Circuit 302 is further configured to perform a merging function 320 configured to merge the processed cropped and upscaled video streams into a merged video stream, preferably having a first resolution.

[0053] Circuit 302 is further configured to perform a first encoding function 322 configured to encode the merged video stream.

[0054] Circuit 302 may be configured to perform a second encoding function 324 configured to encode a downscaled video stream using a scalable video encoding scheme, such as SVC for H.264, SHVC or LCEVC for H.265, or SVT-AV1. The first encoding function 322 may be configured to encode a merged video stream using the encoded processed downscaled video stream as a base layer. The operation of the second encoding function 324 may be incorporated into the first encoding function 322.

[0055] Circuit 302 may be configured to perform a region of interest identification function 316 configured to identify a region of interest in a scene shown by a video stream. Alternatively, or in combination, the region of interest may be identified using external data. By this specification, the term “external data” means data that is not related to the video stream. For example, an external motion sensor may be communicably connected to the video encoding device 300 to provide information about where and / or when motion is detected in the area shown by the video stream. An example of such an external motion sensor is a radar sensor.

[0056] The video processing function 314 is preferably configured to process both the downscaled video stream and the cropped video stream through the same one or more video processing operations 106A to D of the video processing pipeline 104.

[0057] Figure 4 is a flowchart showing the steps of Method 400 for encoding a video stream. Different steps are described in more detail below. Although shown in a specific order, the steps of Method 400 can be performed in any preferred order, in parallel, and multiple times. Blocks with dashed outlines should be interpreted as optional steps of Method 400.

[0058] In S402, pixel data of the video stream is acquired, and the video stream has a first resolution. Acquiring pixel data of the video stream may include capturing pixel data by an image sensor. The video stream comprises a series of image frames.

[0059] In S406, a crop corresponding to the region of interest is extracted from the pixel data of the video stream, and the crop has a first resolution. The region of interest is a sub-part of the scene shown by the video stream. Therefore, the crop contains pixel data representing the region of interest, but does not necessarily contain all the pixel data of each image frame in the video stream. The crop is cut out from the pixel data of the image frames in the video stream.

[0060] In S408, the pixel data of the video stream is downscaled to a downscaled video stream. The downscaled video stream has a second resolution, which is lower than the first resolution. The pixel data of the downscaled video stream contains less data than the pixel data of the original video stream. Therefore, the downscaled video stream can be processed using fewer computing resources.

[0061] In S410, the downscaled video stream is processed through one or more video processing operations 106A to D. In S412, cropping is processed through one or more video processing operations 106A to D. It should be noted that the downscaled video stream and cropping are preferably processed by the same video processing operations. One or more video processing operations 106A to D include one or more of the following: noise reduction, conversion, rotation, privacy mask, image sensor correction, image scaling, gamma correction, image enhancement, color space conversion, chroma subsampling, compression, data storage, and data transmission. Processing of the downscaled video stream and cropping should generally be interpreted as operations performed on the captured pixel data before encoding.

[0062] In S414, the processed downscaled video stream is upscaled to an upscaled video stream. The upscaled video stream preferably has a first resolution. Upscaling the video stream allows the upscaled video stream to be merged with a high-resolution video stream, which preferably has a first resolution.

[0063] In S416, the processed crop and the upscaled video stream are merged to form a merged video stream. In other words, the processed crop is combined with the upscaled video stream to form a single video stream, which is the merged video stream. The merged video stream is therefore a video stream in which the pixel data of areas in the video stream that represent a region of interest contains a relatively large amount of information, while the pixel data of less important areas contains a relatively small amount of information.

[0064] In S420, the merged video streams are encoded. The merged video streams can be encoded using standard encoding schemes, such as H.262 (MPEG-2 Part 2), MPEG-4 Part 2, H.264 (MPEG-4 Part 10), HEVC (H.265), Theora, RealVideo RV40, VP9, ​​and AV1.

[0065] In S418, the downscaled video stream may be encoded using a scalable video encoding scheme. In S420, the act of encoding the merged video stream may include using the encoded downscaled video stream as the base layer.

[0066] This method may include identifying a region of interest in the scene shown by the video stream in S404. In other words, the region of interest is determined from what is shown by the video stream. Identifying the region of interest in S404 may be done by processing the acquired pixel data of the video stream through an attention model. In other words, the attention model is configured to identify the region of interest by analyzing the pixel data of the video stream. The attention model may be based on an object detection model and at least one of an object classification model, a motion detection model, a user-defined input, or a combination thereof. An example of object detection and / or object classification may be that several objects may be marked as of interest. For example, the presence of people may be of particular interest. In such a case, the region of the video stream showing people may be treated as a region of interest. More specifically, people's faces may be of particular interest. In that case, only faces need to be processed at full resolution. An example of a motion detection model may be when an area of ​​the video stream showing moving objects should be treated as a region of interest. An example of a user-defined input may be that the user marks a sub-part of the scene shown by the video stream as of interest. In the case of a surveillance camera, this could be a doorway or walkway in the scene shown by the video stream. Alternatively, or in combination, the region of interest may be identified based on external data. By this specification, the term “external data” means data that is not related to the video stream. For example, the region of interest may be identified from the sensor signals of an external motion sensor, such as a radar sensor.

[0067] In addition, the variations of the disclosed variations can be understood and realized by those skilled in the art when performing the claimed invention, based on the drawings, this disclosure, and the accompanying study of the claims.

Claims

1. A method (400) for encoding a video stream, wherein the method (400) is Acquiring the pixel data of the aforementioned video stream (S402), Extracting a crop corresponding to the region of interest from the aforementioned pixel data (S406), wherein the crop has a first resolution (S406), The pixel data is downscaled to a downscaled video stream having a second resolution lower than the first resolution (S408), Processing the downscaled video stream having the second resolution through one or more video processing operations (106A to D) comprising at least one of noise reduction, rotation, image sensor correction, image scaling, gamma correction, image enhancement, color space conversion, and chroma subsampling functions (S410), Processing the crop having the first resolution through one or more video processing operations (106A to D) (S412), wherein the processing of pixel data having the first resolution through one or more video processing operations is performed only on the crop having the first resolution, and the processing of the remaining pixel data other than the crop is limited to processing the downscaled video stream having the second resolution, Upscaling the processed downscaled video stream having the second resolution to an upscaled video stream having the first resolution (S414), Merging the processed crop having the first resolution and the upscaled video stream having the first resolution into a merged video stream (S416), Encoding the merged video stream (S420) Method (400), including.

2. Encode the processed downscaled video stream having the second resolution using a scalable video encoding scheme (S418) It further includes, The method according to claim 1 (400), wherein encoding the merged video stream having the first resolution (S420) includes using the encoded downscaled video stream having the second resolution as a base layer.

3. The method according to claim 1 (400), wherein the region of interest is a sub-part of the scene shown by the video stream.

4. The method according to claim 1 (400), further comprising identifying the region of interest in the scene shown by the video stream (S404).

5. The method according to claim 4 (400), wherein identifying the region of interest (S404) is performed by processing the acquired pixel data of the video stream through an attention model.

6. The method according to claim 5 (400), wherein the attention model is based on an object detection model and at least one of an object classification model, a motion detection model, a user-defined input, or a combination thereof.

7. The method according to claim 1 (400), wherein acquiring the pixel data of the video stream includes capturing the pixel data with an image sensor.

8. A video encoding device (300) for encoding a video stream, wherein the video encoding device (300) An acquisition function (308) configured to acquire pixel data of the video stream having a first resolution, An extraction function (310) configured to extract a crop corresponding to a region of interest from the aforementioned pixel data, wherein the crop has the first resolution, A downscaling function (312) configured to downscale the pixel data to a downscaled video stream having a second resolution lower than the first resolution, A video processing function (314) configured to process the downscaled video stream having a second resolution through one or more video processing operations (106A to D) of a video processing pipeline (104), and configured to process the crop having a first resolution through one or more video processing operations (106A to D) of the video processing pipeline (104), wherein one or more video processing operations (106A to D) comprises at least one of noise reduction, rotation, image sensor correction, image scaling, gamma correction, image enhancement, color space conversion, and chroma subsampling functions, and processing of the pixel data having the first resolution through one or more video processing operations is performed only on the crop having the first resolution, and processing of the remaining pixel data other than the crop is limited to processing the downscaled video stream having a second resolution, An upscaling function (318) configured to upscale the processed downscaled video stream to an upscaled video stream having the first resolution, A merging function (320) configured to merge the processed crop having the first resolution and the upscaled video stream having the first resolution into a merged video stream, A first encoding function (322) configured to encode the merged video stream and A video encoding device (300) comprising a circuit (302) configured to perform the following.

9. The circuit (302) is A second encoding function (324) configured to encode the processed downscaled video stream having the second resolution using a scalable video encoding scheme. Further configured to perform, The video encoding device (300) according to claim 8, wherein the first encoding function (322) is configured to encode the merged video stream having the first resolution using the encoded processed downscaled video stream having the second resolution as a base layer.

10. The video encoding device (300) according to claim 8 or 9, wherein the circuit (302) is further configured to perform a region of interest identification function (316) configured to identify the region of interest in a scene shown by the video stream.

11. The video encoding device (300) according to claim 8 or 9, wherein the video processing function (314) is configured to process both the downscaled video stream having the second resolution and the crop having the first resolution through the same one or more video processing operations (106A to D) of the video processing pipeline (104).

12. A non-temporary computer-readable recording medium that stores program code configured to perform the method (400) according to any one of claims 1 to 7 when executed on a device having processing capabilities.