Area of ​​Interest Coding for VCM

A hybrid video codec with spatial and temporal ROI encoding methods addresses the inefficiency of existing codecs for machine vision, enhancing encoding efficiency and compression for machine and hybrid human/machine vision tasks.

JP7893979B2Active Publication Date: 2026-07-22TENCENT AMERICA LLC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2023-09-29
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing video codecs are primarily designed for human consumption and lack efficiency in encoding videos for machine vision tasks such as object detection, segmentation, and tracking, necessitating the development of specialized codecs for machine and hybrid human/machine vision.

Method used

A hybrid video codec combining legacy and learning-based codecs, with spatial and temporal region of interest (ROI) encoding methods to prioritize relevant objects, using bounding boxes and selective frame sampling for efficient encoding.

Benefits of technology

Enhances encoding efficiency for machine vision tasks by reducing bitrate and improving compression, allowing for effective machine and hybrid human/machine video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893979000009
    Figure 0007893979000009
  • Figure 0007893979000010
    Figure 0007893979000010
  • Figure 0007893979000011
    Figure 0007893979000011
Patent Text Reader

Abstract

A technique for encoding video for machine vision and human / machine hybrid vision includes receiving image data. The technique may also include detecting a plurality of bounding boxes associated with a plurality of objects of interest in a frame of the image data and detecting a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes. The technique may then include encoding the frame-level bounding box using a first bit rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , , , , ,

[0004] ,

[0001] Cross - reference to Related Applications

[0001] This application claims priority from U.S. Provisional Application No. 63 / 412,367, filed on September 30, 2022, and U.S. Application No. 18 / 477,189, filed on September 28, 2023, the entire disclosures of which are incorporated herein by reference.

[0002]

[0002] This disclosure relates to video coding for machine vision. Specifically, techniques for encoding video for machine vision and human / machine hybrid vision are disclosed.

Background Art

[0003]

[0003] Conventionally, videos or images have been consumed by humans for various uses such as entertainment, education, etc. Thus, video coding or image coding often utilizes the characteristics of the human visual system for better compression efficiency while maintaining good subjective quality.

[0004]

[0004] In recent years, with the increase in machine learning applications, many intelligent platforms, together with a large number of sensors, are using videos for machine vision tasks such as object detection, segmentation, or tracking. How to encode videos or images for consumption by machine tasks is an interesting and difficult problem, which has led to the introduction of research on video coding for machine (VCM). To achieve this goal, the international standard group MPEG created an ad - hoc group, "Video Coding for Machine (VCM)", to standardize the related techniques for better interoperability between different devices.

[0005]

[0005] Existing video codecs are primarily for human consumption. However, an increasing amount of video is being consumed by machines for machine vision tasks such as object detection, instance segmentation, and object tracking. It is important to develop video codecs that efficiently encode video for machine vision or hybrid machine / human vision. [Overview of the project]

[0006]

[0006] The following provides a simplified overview of one or more embodiments of the present disclosure in order to provide a basic understanding of such embodiments. The overview of the present invention is not a broad overview of all intended embodiments, nor does it identify the main or important elements of all embodiments, nor does it define the scope of any or all embodiments. Its sole purpose is to present in a simplified form some concepts of one or more embodiments of the present disclosure as a prelude to the more detailed descriptions to be presented later.

[0007]

[0007] Methods, apparatus, and non-temporary computer-readable media relating to encoding video for machine vision and human / machine hybrid vision.

[0008]

[0008] Methods for encoding video for machine vision and human / machine hybrid vision may be provided. The method may be performed by one or more processors and may include receiving image data, detecting a plurality of bounding boxes relating to a plurality of objects of interest in a frame of the image data, detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, and encoding the frame-level bounding box using a first bitrate.

[0009]

[0009] An apparatus for encoding video for machine vision and human / machine hybrid vision may be provided. The apparatus may include at least one memory configured to store program code and at least one processor configured to access the program code. The at least one processor may be configured to operate as instructed by the program code, and the program code includes a first receive code configured to cause at least one processor to receive image data; a first detection code configured to cause at least one processor to detect a plurality of bounding boxes relating to a plurality of objects of interest in a frame of image data; a second detection code configured to cause at least one processor to detect a frame-level bounding box for a frame based on the coordinates of the plurality of bounding boxes; and a first encoding code configured to cause at least one processor to encode the frame-level bounding box using a first bitrate.

[0010]

[0010] A non-temporary computer-readable medium storing computer instructions is provided, which, when the computer instructions are executed by at least one processor for encoding video for machine vision and human / machine hybrid vision, causes at least one processor to receive image data, detect a plurality of bounding boxes relating to a plurality of objects of interest in a frame of the image data, detect a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, and encode the frame-level bounding box using a first bitrate.

[0011]

[0011] Additional embodiments are described below, partially apparent therefrom, and / or may be grasped by the practice of the embodiments presented in this disclosure.

[0012]

[0012] The above and other features and aspects of the embodiments of the present disclosure will become apparent from the following description made in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0013] [Figure 1]

[0013] This is a diagram of an exemplary network device according to various embodiments of the present disclosure. [Figure 2]

[0014] This figure shows the architecture of the disclosed hybrid video codec according to one embodiment of the present disclosure. [Figure 3]

[0015] This is a diagram of video coding for a mechanical system according to various embodiments of the present disclosure. [Figure 4A]

[0016] This is a diagram of a spatial region of interest bounding box according to various embodiments of the present disclosure. [Figure 4B]

[0017] This figure shows the calculation of the bounding box of the region of interest according to various embodiments of the present disclosure. [Figure 5]

[0018] This figure shows the calculation of the region of interest bounding box for an intra-period n according to various embodiments of this disclosure. [Figure 6]

[0019] This is a flowchart illustrating an exemplary process for determining a region of interest in a hybrid video codec according to various embodiments of the present disclosure. [Modes for carrying out the invention]

[0014]

[0020] A detailed description of exemplary embodiments follows with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0015]

[0021] The above disclosures are illustrative and descriptive, but are not intended to be exhaustive or to limit implementations to the exact forms disclosed. Modifications and variations may be possible in light of the above disclosures or may be derived from the practice of implementations. Furthermore, one or more features or components of some embodiments may be incorporated into or combined with some embodiments (or one or more features of some embodiments). In addition, it should be understood that in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be changed.

[0016]

[0022] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The specific control hardware or software code used to implement these systems and / or methods is not limiting to the implementation form. Therefore, it should be understood that the operation and behavior of the systems and / or methods are described herein independently of any specific software code, and that software and hardware may be designed to implement the systems and / or methods based on the descriptions herein.

[0017]

[0023] Certain combinations of features are described in the claims and / or disclosed herein, but these combinations do not limit the disclosure of possible implementations. In practice, many of these features can be combined in ways not described in detail in the claims and / or disclosed herein. Each dependent claim described below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim combined with any other claims in that claim set.

[0018]

[0024] Any element, act, or instruction used in this specification should not be construed as important or essential unless explicitly described as such. Also, the articles "a" and "an" used in this specification are intended to include one or more items and may be used interchangeably with "one or more". When only one item is intended, the term "one" or a similar expression is used. Also, terms such as "has", "have", "having", "include", "including", etc. used in this specification shall be considered open-ended terms. Furthermore, the phrase "based on" shall mean "at least partially based on" unless otherwise specified. Additionally, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0019]

[0025] References throughout this specification to "some embodiments", "an embodiment", or similar phrases mean that the particular features, structures, or characteristics described in relation to the indicated embodiments are included in some embodiments of the solution. Thus, phrases such as "in some embodiments" and "in an embodiment" throughout this specification, and similar phrases, do not necessarily, but may, all refer to the same embodiment.

[0020]

[0026] Furthermore, the features, advantages, and characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. In light of the description herein, those skilled in the art will recognize that the disclosure can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of the disclosure may be recognized in some embodiments.

[0021]

[0027] The disclosed methods can be used separately or combined in any order. Further, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0022]

[0028] Embodiments of the present disclosure are directed to video coding for machines. Specifically, methods for encoding video for machine vision and human / machine hybrid vision are disclosed. Conventional video coders are designed for human consumption. In some embodiments, conventional video coders can be combined with a learning-based coder to form a hybrid coder such that video can be efficiently coded for machine vision and hybrid human and machine vision.

[0023]

[0029] Figure 1 shows an exemplary device for performing translation services. Device 100 can correspond to any type of known computer, server, or data processing device. For example, device 100 may comprise a processor, personal computer (PC), printed circuit board (PCB) with computing devices, minicomputer, mainframe computer, microcomputer, telephone computing device, wired / wireless computing device (e.g., smartphone, personal digital assistant (PDA)), laptop, tablet, smart device, or any other similar operating device.

[0024]

[0030] In some embodiments, as shown in Figure 1, the device 100 may include a set of components such as a processor 120, memory 130, storage components 140, input components 150, output components 160, and a communication interface 170.

[0025]

[0031] Bus 110 may comprise one or more components that enable communication between a set of components of device 100. For example, bus 110 could be a communication bus, a crossover bar, a network, etc. Although bus 110 is shown as a single line in Figure 1, bus 110 may be implemented using multiple (two or more) connections between the set of components of device 100. This disclosure is not limited to this extent.

[0026]

[0032] Device 100 may comprise one or more processors, such as processor 120. Processor 120 may be implemented as hardware, firmware, and / or a combination of hardware and software. For example, processor 120 may comprise a central processing unit (CPU), graphics processing unit (GPU), acceleration unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), general-purpose single-chip or multi-chip processor, or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the operations described herein. A general-purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. Processor 120 may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors working with a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be carried out by circuits specific to a given operation.

[0027]

[0033] The processor 120 can control the overall operation of device 100 and / or a set of components of device 100 (for example, memory 130, storage component 140, input component 150, output component 160, and communication interface 170).

[0028]

[0034] Device 100 may further comprise memory 130. In some embodiments, memory 130 may comprise random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic memory, optical memory, and / or other types of dynamic or static storage devices. Memory 130 may store information and / or instructions for use (e.g., execution) by processor 120.

[0029]

[0035] The storage component 140 of device 100 may store information and / or computer-readable instructions and / or code relating to the operation and use of device 100. For example, the storage component 140 may include, together with a corresponding drive, a hard disk (e.g., magnetic disk, optical disk, magneto-optical disk, and / or solid-state disk), a compact disk (CD), a digital multipurpose disk (DVD), a Universal Serial Bus (USB) flash drive, a Personal Computer Memory Card International Association (PCMCIA) card, a floppy disk, a cartridge, a magnetic tape, and / or another type of non-temporary computer-readable medium.

[0030]

[0036] Device 100 may further comprise input components 150. The input components 150 may include one or more components that enable Device 100 to receive information, such as via user input (e.g., a touchscreen, keyboard, keypad, mouse, stylus, button, switch, microphone, camera, etc.). Alternatively or additionally, the input components 150 may include sensors for detecting information (e.g., a Global Positioning System (GPS) component, accelerometer, gyroscope, actuator, etc.).

[0031]

[0037] The output component 160 of device 100 may include one or more components that can provide output information from device 100 (for example, a display, a liquid crystal display (LCD), a light-emitting diode (LED), an organic light-emitting diode (OLED), a haptic feedback device, a speaker, etc.).

[0032]

[0038] Device 100 may further comprise a communication interface 170. The communication interface 170 may include receiver components, transmitter components, and / or transceiver components. The communication interface 170 may enable device 100 to establish connections with and / or transmit communications with other devices (e.g., a server, another device). Communications may be carried out via wired connections, wireless connections, or a combination of wired and wireless connections. The communication interface 170 may enable device 100 to receive information from and / or provide information to another device. In some embodiments, the communication interface 170 may provide communication with another device over a network, such as a local area network (LAN), wide area network (WAN), metropolitan area network (MAN), private network, ad hoc network, intranet, internet, fiber optic-based network, cellular network (e.g., fifth-generation (5G) network, long-term evolution (LTE) network, third-generation (3G) network, code division multiple access (CDMA) network, etc.), public land mobile network (PLMN), telephone network (e.g., public switched telephone network (PSTN)), and / or a combination of these or other types of networks. Alternatively or additionally, the communication interface 170 may provide communication with another device over a device-to-device (D2D) communication link, such as FlashLinQ, WiMedia, Bluetooth®, ZigBee®, Wi-Fi®, LTE, 5G, etc. In other embodiments, the communication interface 170 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, and the like.

[0033]

[0039] Device 100 is included in the core network 240 and may perform one or more processes described herein. Device 100 may perform operations based on the processor 120 executing computer-readable instructions and / or code that can be stored in a non-temporary computer-readable medium such as memory 130 and / or storage component 140. The computer-readable medium may refer to a non-temporary memory device. The memory device may include a memory space within a single physical storage device and / or a memory space spread across multiple physical storage devices.

[0034]

[0040] Computer-readable instructions and / or code may be read into memory 130 and / or storage component 140 from another computer-readable medium or from another device via the communication interface 170. When the computer-readable instructions and / or code stored in memory 130 and / or storage component 140 are executed by the processor 120, or when they have been executed, the device 100 may perform one or more processes described herein.

[0035]

[0041] Alternatively or in addition, hardwired circuitry may be used instead of or in combination with software instructions to carry out one or more processes described herein. Therefore, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0036]

[0042] The number and arrangement of components shown in Figure 1 are provided as an example. In practice, there may be additional components other than those shown in Figure 1, fewer components, different components, or components arranged differently. Furthermore, two or more components shown in Figure 1 may be implemented within a single component, or a single component shown in Figure 1 may be implemented as multiple distributed components. Additionally or alternatively, a set of (one or more) components shown in Figure 1 may perform one or more operations that are described as being performed by another set of components shown in Figure 1.

[0037]

[0043] Figure 2 is a block diagram of a view of one embodiment of the hybrid video codec 200. The hybrid video codec 200 may include a legacy codec 220 and a learning-based codec 230. The input 201 to the hybrid codec may be video or an image, since the image may be treated as a special type of video (e.g., video with one image). In Figure 2, the legacy video codec 220 may be used to compress the input video 201 at different scales (e.g., original resolution or downsampled), and the downsampling ratio of the downsampling module 210 may be fixed and known to both the encoder 221 and the decoder 223, or the downsampling ratio may be user-defined, e.g., 100% (e.g., no downsampling), 50%, 25%, etc., and may be sent as metadata in the bitstream 224 to inform the decoder 222. The legacy video codec may be an image codec such as VVC, HEVC, H264, or JPEG, JPEG2000. The downsampling module 210 may be a classical image downsampler or a learning-based image downsampler. The decoded downsampled video 203 (e.g., “low-resolution video 203” in Figure 2) may be upsampled to the original resolution of the video (e.g., “high-resolution video 204”) that can be used for human vision using the upsampling module 250. The upsampling module 250 may be a classical image upsampler or a learning-based image upsampler such as a learning-based super-resolution module.

[0038]

[0044] In some embodiments, the hybrid video codec 200 may also employ a learning-based video codec 230 to compress the downsampled video 202.

[0039]

[0045] In the encoder, a reconstructed video 203 is also generated and can be upsampled to the original input resolution. The upsampled reconstructed video 205 can then be subtracted from the input video to generate a residual video signal 202, which can be fed to the learning-based codec 230 in Figure 2. The upsampling module 240 of the hybrid video codec 200 may be the same as the upsampling module 250 (after low resolution). Video 203 can be decoded in the decoder. The output of the residual decoder 238 can be added on top of the high-resolution video 204 to form a reconstructed video 205, which can be used for machine vision tasks.

[0040]

[0046] Figure 3 shows an embodiment of the architecture for a video coding machine (VCM), such as a hybrid video codec 200. The sensor output 300 follows a video coding path 311, passes through a VCM encoder 310, and reaches a VCM decoder 320 where video decoding 321 takes place. Another path goes to feature extraction 312, then to feature transformation 313, and finally to feature coding 314 and feature decoding 322. The output of the VCM decoder 320 is primarily for machine consumption, i.e., machine vision 305. In some cases, it may also be used for human vision 306. One or more machine tasks for understanding video content are then performed.

[0041]

[0047] As is well known, video content often contains a lot of information, such as moving objects in the foreground and static scenes in the background. Classical video codecs often consider temporal or specific redundancy to compress the content. In video coding for machines, machine vision tasks are often object detection, instance segmentation, or object tracking. In these types of tasks, objects such as people, cars, and bicycles are of primary interest, while other information such as trees, grass, and the sky are of little interest. It would be beneficial to consider these types of observations to further reduce the information that needs to be transmitted or stored. Therefore, region of interest coding is disclosed in this application.

[0042]

[0048] The proposed methods may be used separately or in any combination in any order. Furthermore, each of the methods (or embodiments), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-temporary computer-readable medium. In this disclosure, pictures, images, and frames are interchangeable.

[0043]

[0049] According to the embodiment, two types of ROI methods are described, for example, spatial region of interest (ROI) and temporal ROI.

[0044]

[0050] Spatial region of interest coding for image coding

[0045]

[0051] According to one embodiment, on the encoder side, an object detector may be used to detect the bounding boxes of all relevant objects in a frame / image (image or frame may be used interchangeably in this disclosure), and then a region of interest bounding box may be calculated such that it includes all the bounding boxes of the relevant objects. Figure 4A is an illustrative diagram of a region of interest bounding box determined for a frame 400 that encapsulates several objects of interest.

[0046]

[0052] As shown in Figure 4A, pedestrians and people riding bicycles or motorcycles are marked with their relative dark bounding boxes. The outer white box is the ROI bounding box, which contains all the bounding boxes of people in the scene.

[0047]

[0053] According to one embodiment, for a frame, the bounding boxes of all related objects are such that for k=0, ..., N,

[0048]

number

[0049]

number

[0050]

number

[0051]

number

[0052]

[0054] In the above equation,

[0053]

number

[0054]

number

[0055]

number

[0056]

number

[0057]

[0055] The calculation of the ROI bounding box is shown in Figure 4B. As shown in Figure 4B, the black boxes represent the related objects (e.g., 465-1, 465-2), and the dotted box 460 represents the box that exactly contains the related objects. The dashed box 455 is the ROI bounding with margins added.

[0058]

[0056] According to one embodiment, after the calculation of the ROI bounding box (e.g., 455), only the content within the ROI bounding box may be encoded to obtain the image bitstream in order to reduce the bitrate for transmission or storage. For example, a first higher bitrate may be used to encode only the content within the ROI box.

[0059]

[0057] In the same or another embodiment, frames without an ROI bounding box may be encoded using a low bitrate configuration, such as a high quantization process size or reduced scale, if their content may be of interest in the future. For example, for coding frames that simply do not have an ROI bounding box, the ROI region may be filled with a constant value, such as 0 or 128.

[0060]

[0058] In the same or other embodiments, the entire frame may be encoded using a low bitrate configuration, such as a high quantization process size or a reduced scale.

[0061]

[0059] According to one embodiment, at the decoder side, a picture or frame containing an ROI bounding box may be decoded and used alone, or it may be combined with frames that do not have an ROI portion to form a final picture frame. In some embodiments, a padding method may be used to extend the boundaries of the picture or frame to the same size as the original picture (which has an area outside the ROI bounding box). In embodiments where it is necessary to restore the picture to its original size, the original size and ROI bounding may be sent as metadata in the bitstream.

[0062]

[0060] Accordingly, as an exemplary embodiment, a method for spatial ROI disclosed herein may include receiving image data from a sensor, comprising a plurality of frames. A plurality of bounding boxes relating to a plurality of objects of interest in the frame may then be detected for each frame in the image data. A frame-level bounding box for a frame may be detected based on the coordinates of the plurality of bounding boxes. The frame-level bounding box may be encoded using a first bitrate. In embodiments, a portion of a frame not included in the frame-level bounding box may be encoded using a second bitrate, the second bitrate being lower than the first bitrate. In embodiments, in response to encoding the frame-level bounding box using the first bitrate, the frame, or at least the remainder of the frame, may be encoded using the second bitrate. In some embodiments, a portion of a frame included in the frame-level bounding box may be filled using a constant value before encoding the frame using the second bitrate.

[0063]

[0061] Coding of spatial regions of interest for video coding

[0064]

[0062] In some embodiments, already coded frames may be used as a reference to predict the current frame. In such embodiments, ROI bounding box information for each frame may be carried as metadata in the bitstream so that the decoder can place the decoded ROI content in the correct position in the original image to facilitate motion compensation.

[0065]

[0063] According to one embodiment, a video sequence can be divided into multiple intra periods. For example, the intra periods may be at intervals of about 1 or 2 seconds, which may correspond to 32 or 64 frames for a video with a frame rate of 30 frames / second. According to one embodiment, for one intra period, ROI bounding boxes can be calculated for all frames. A common ROI bounding box size can be calculated that includes all individual ROI bounding boxes with a certain margin.

[0066]

[0064] Accordingly, according to one embodiment, the method disclosed herein may include determining a common intra-period level bounding box for an intra-period comprising a plurality of temporally consecutive frames based on a plurality of frame-level bounding boxes for the plurality of temporally consecutive frames within the intra-period. The method may also include signaling the top-left and bottom-right coordinates of the common intra-period level bounding box as metadata in the bitstream. In some embodiments, the intra-period may comprise a predetermined number of temporally consecutive frames.

[0067]

[0065] As described above, certain information related to a frame can be signaled as metadata. For example, the original size of a frame and the size of the frame-level bounding box can be signaled as metadata. When signaling the size of the frame-level bounding box, the top-left and bottom-right coordinates of the frame-level bounding box can be signaled.

[0068]

[0066] Figure 5 shows a common ROI bounding box for the intra-period, i.e., the black boxes (505-1, 505-1, ... 505-N). By using a common ROI bounding box for the intra-period, only the common bounding box information, i.e., its top-left and bottom-right coordinates, is sent in the bitstream as metadata.

[0069]

[0067] According to one embodiment, only content within a common ROI bounding for each frame is sent in the bitstream or as metadata for each intra-period or for a specific number of frames. In such an embodiment, all frames in an intra-period have the same resolution so that regular motion prediction / compensation can be performed. Thus, in the embodiment, the frames of image data to be analyzed for the ROI may be selected based on a predetermined sample rate or one frame may be selected for every predetermined number of frames of image data. In the embodiment, the predetermined number of frames may vary based on the complexity of the scene. In the embodiment, some additional information related to the frames, such as a predetermined sample rate or a predetermined number of frames, may be signaled in the bitstream as metadata.

[0070]

[0068] In the same or different embodiments, the original frame that does not have content in the ROI region may also be coded and sent in a low bitrate configuration, i.e., a large quantization process or a lower scale, etc. In embodiments, the ROI region may be filled with a constant value, such as 0 or 128.

[0071]

[0069] In the same or another embodiment, the original frame may also be coded and sent in a low bitrate configuration, i.e., a large quantization process or a lower scale, etc.

[0072]

[0070] According to one embodiment, in the decoder, a picture or frame containing an ROI bounding box may be decoded and used alone or combined with frames that do not have an ROI portion to form a final picture. In the embodiment, a padding method may be used to extend the boundaries of the picture to the same size as the original picture (which has an area outside the ROI bounding box). In embodiments where video restoration at the original resolution is required, the original video resolution may be sent as metadata.

[0073]

[0071] Time-of-interest coding for video coding

[0074]

[0072] As is known, consecutive video frames can share a lot of common information. For example, consecutive frames may contain the same object with slightly different positions or shapes. This observation can be used in video coding for machines to reduce transmission rate or memory. The encoder may utilize an analysis module to determine the motion characteristics of the video segment.

[0075]

[0073] In the same or another embodiment, video frames may be downsampled in time, i.e., only one frame out of every N frames may be selected for encoding and transmission / storage. N may be set as a fixed value or it may be a variable. For example, N may be small for dynamic scenes, e.g., 2, or relatively large for relatively static content, e.g., 4 or 8.

[0076]

[0074] In the same or another embodiment, instead of subsampling the video evenly throughout the time domain, portions of the video may be selected for coding and transmission / storage if an action or event occurs in a particular portion. Portions of content may be ignored without coding if there is no action / event in a particular portion. The remainder of the content may be sampled evenly.

[0077]

[0075] In the same or other embodiments, the video frames are sampled unevenly. For example, for the first 12 frames, frame numbers 0, 3, 7, 8, and 11 are selected for encoding.

[0078]

[0076] Metadata such as the sample rate N may be sent in the bitstream to indicate timing information of the coded content. If different parts of the video use the same sample rate N, the sample rate N is sent only at the beginning of the bitstream for that part. For unevenly sampled video, metadata such as frame number or timestamp may be sent in the bitstream.

[0079]

[0077] In situations where it is necessary to restore video frames with the original frame rate, duplicated pictures may be inserted into the decoded video. For example, if only frames with frame numbers 0, 3, 7, 8, and 1 are sent, they may be indicated as f0, f3, f7, f8, and f11, respectively. The decoder may then duplicate those frames as, for example, f0, f0, f0, f3, f3, f3, f3, f7, f8, f8, f8, f11, etc., in order to obtain 12 frames.

[0080]

[0078] In embodiments, in order to further reduce the bitrate for transmission or storage, embodiments may combine a spatial ROI method and a temporal ROI method as disclosed herein in video coding for machines.

[0081]

[0079] Figure 6 shows a process 600 for encoding video for machine vision and human / machine hybrid vision according to one embodiment.

[0082]

[0080] As shown in Figure 6, image data may be received during operation 605. In this embodiment, the image data may be received via a network or from a sensor.

[0083]

[0081] In operation 610, multiple bounding boxes associated with multiple objects of interest may be determined within a single frame of the image data.

[0084]

[0082] In the embodiment, frames may be selected based on a predetermined sample rate, or one frame may be selected for every predetermined number of frames of image data. In the embodiment, the predetermined number of frames may vary based on the complexity of the scene.

[0085]

[0083] In operation 615, a frame-level bounding box for a frame may be detected based on the coordinates of multiple bounding boxes.

[0086]

[0084] In operation 620, the frame-level bounding box may be encoded using a first bitrate.

[0087]

[0085] In operation 625, sections of a frame not included in the frame-level bounding box may be encoded using a second bitrate. In embodiments, the second bitrate may be lower than the first bitrate. In embodiments, in response to encoding the frame-level bounding box using the first bitrate and / or in response to encoding the frame using the second bitrate, sections of a frame included in the frame-level bounding box may be filled using a constant value before encoding the frame using the second bitrate.

[0088]

[0086] In the embodiment, for an intra-period comprising multiple temporally consecutive frames, a common intra-period level bounding box may be determined based on multiple frame-level bounding boxes for the multiple temporally consecutive frames within the intra-period, and the top-left and bottom-right coordinates of the common intra-period level bounding box may be signaled in the bitstream as metadata. An intra-period consists of a predetermined number of temporally consecutive frames.

[0089]

[0087] Other metadata that can be signaled in the bitstream may include the original size of the frame, the size of the frame-level bounding box, the top-left and bottom-right coordinates of the frame-level bounding box, a predetermined sample rate, or the predetermined number of frames mentioned above.

[0090]

[0088] The above disclosures are illustrative and explanatory, but are not intended to be exhaustive or to limit implementations to the exact forms disclosed. Modifications and variations may be possible in light of the above disclosures or may be obtained from the practice of implementations.

[0091]

[0089] It should be understood that the specific order or hierarchy of blocks in the process / flowchart disclosed herein is illustrative of an exemplary technique. It should be understood that the specific order or hierarchy of blocks in the process / flowchart may be rearranged based on design preferences. Furthermore, some blocks may be combined or omitted. The appended method claims present elements of various blocks in an exemplary order and are not limited to the specific order or hierarchy presented.

[0092]

[0090] Some embodiments may relate to systems, methods, and / or computer-readable media at a level that integrates any possible technical details. Furthermore, one or more of the above-described components may be implemented as instructions that are stored in a computer-readable medium and are executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include (one or more) computer-readable non-temporary storage media having computer-readable program instructions thereon for causing a processor to perform an action.

[0093]

[0091] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any preferred combination of the above. A non-exhaustive list of more specific examples of computer-readable storage mediums includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves on which instructions are recorded, and any preferred combination of the above. The computer-readable storage media used herein should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0094]

[0092] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fibers, wireless transmitters, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer-readable program instructions from the network and forwards those computer-readable program instructions for storage in a computer-readable storage medium within each computing / processing device.

[0095]

[0093] The computer-readable program code / instructions for performing the operation may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions may run entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing state information of computer-readable program instructions to personalize the electronic circuit in order to perform an action or operation.

[0096]

[0094] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device for creating machines, and as a result, instructions executed via the processor of the computer or other programmable data processing device create means for implementing operations specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can guide computers, programmable data processing devices, and / or other devices to operate in a particular manner, and as a result, the computer-readable storage medium on which the instructions are stored comprises a product containing instructions that implement modes of operations specified in one or more blocks of a flowchart and / or block diagram.

[0097]

[0095] Computer-readable program instructions may also be loaded onto a computer, other programmable device, or other device to cause a series of operations to be performed on the computer, other programmable data processing device, or other device in order to create a computer implementation process, and as a result, the instructions executed on the computer, other programmable device, or other device implement the operations specified in one or more blocks of a flowchart and / or block diagram.

[0098]

[0096] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical operation. Methods, computer systems, and computer-readable media may include additional blocks, fewer blocks, different blocks, or blocks arranged differently from those shown in the figures. In some alternative implementations, the operations within blocks may be performed out of order in the figures. For example, two blocks shown consecutively may be executed effectively or substantially simultaneously, or blocks may sometimes be executed in reverse order, depending on the functionality involved. Each block in a block diagram and / or flowchart, as well as combinations of blocks in a block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs a specified operation or executes it in combination with dedicated hardware and computer instructions.

[0099]

[0097] It will be apparent that the systems and / or methods described herein may be implemented in different forms, such as hardware, firmware, or a combination of hardware and software. The specific control hardware or software code used to implement these systems and / or methods is not limited to the implementation form. Therefore, it should be understood that the operation and behavior of the systems and / or methods are described herein independently of any specific software code, and that software and hardware may be designed to implement the systems and / or methods based on the descriptions herein.

Claims

1. A method for encoding video for machine vision and human / machine hybrid vision, wherein the method is performed by one or more processors, Receiving image data and To detect multiple bounding boxes associated with multiple objects of interest in the frame of the aforementioned image data, Detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, Encoding the frame-level bounding box using the first bitrate and Includes, The method described above is For an intra-period that includes multiple temporally consecutive frames, a common intra-period level bounding box is determined based on multiple frame-level bounding boxes for the multiple temporally consecutive frames within the intra-period. The top-left and bottom-right coordinates of the common intra-period level bounding box are signaled as metadata in the bitstream. Methods that further include this.

2. The method according to claim 1, wherein the intra period includes a predetermined number of temporally consecutive frames.

3. A method for encoding video for machine vision and human / machine hybrid vision, wherein the method is performed by one or more processors, Receiving image data and To detect multiple bounding boxes associated with multiple objects of interest in the frame of the aforementioned image data, Detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, Encoding the frame-level bounding box using the first bitrate and Includes, The method described above is Signaling the original size of the frame and the size of the frame-level bounding box as metadata in the bitstream. Methods that further include this.

4. The method according to claim 3, wherein signaling the size of the frame-level bounding box includes signaling the top-left and bottom-right coordinates of the frame-level bounding box.

5. The method described above is Encode the portion of the frame that is not included in the frame-level bounding box using a second bitrate. It further includes, The method according to claim 1, wherein the second bitrate is lower than the first bitrate.

6. A method for encoding video for machine vision and human / machine hybrid vision, wherein the method is performed by one or more processors, Receiving image data and To detect multiple bounding boxes associated with multiple objects of interest in the frame of the aforementioned image data, Detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, Encoding the frame-level bounding box using the first bitrate and Includes, The method described above is Encoding the frame using a second bitrate in response to encoding the frame-level bounding box using the first bitrate. It further includes, A method in which a section of the frame contained within the frame-level bounding box is filled with a constant value before encoding the frame using the second bitrate.

7. The aforementioned method, Encoding the frame using a second bitrate It further includes, The method according to claim 1, wherein the second bitrate is lower than the first bitrate.

8. The method according to claim 1, wherein the frames of the image data are selected based on a predetermined sample rate, or one frame is selected for each predetermined number of frames of the image data.

9. A method for encoding video for machine vision and human / machine hybrid vision, wherein the method is performed by one or more processors, Receiving image data and To detect multiple bounding boxes associated with multiple objects of interest in the frame of the aforementioned image data, Detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, Encoding the frame-level bounding box using the first bitrate and Includes, The frames of the image data are selected based on a predetermined sample rate, or one frame is selected for every predetermined number of frames of the image data. A method wherein the predetermined number of frames varies based on the complexity of the scene.

10. The method described above is Signaling the predetermined sample rate or predetermined number of frames as metadata in the bitstream. The method according to claim 9, further comprising:

11. An apparatus for encoding video for machine vision and human / machine hybrid vision, wherein the apparatus is At least one memory configured to store program code, At least one processor configured to access the program code and operate as instructed by the program code, The program code is provided with, A first receive code configured to cause at least one processor to receive image data, A first detection code configured to cause at least one processor to detect a plurality of bounding boxes associated with a plurality of objects of interest in the frame of the image data, A second detection code configured to cause at least one processor to detect a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, A first encoding code configured to cause at least one processor to encode the frame-level bounding box using a first bitrate and Includes, The aforementioned program code A first determination code configured to cause at least one processor to determine a common intra-period level bounding box for an intra-period including a plurality of temporally consecutive frames, based on a plurality of frame-level bounding boxes for the plurality of temporally consecutive frames during the intra-period, A first signaling code configured to cause at least one processor to signal the top-left and bottom-right coordinates of the common intra-period level bounding box as metadata in the bitstream, and A device that further includes the following.

12. An apparatus for encoding video for machine vision and human / machine hybrid vision, wherein the apparatus is At least one memory configured to store program code, At least one processor configured to access the program code and operate as instructed by the program code, The program code is provided with, A first receive code configured to cause at least one processor to receive image data, A first detection code configured to cause at least one processor to detect a plurality of bounding boxes associated with a plurality of objects of interest in the frame of the image data, A second detection code configured to cause at least one processor to detect a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, A first encoding code configured to cause at least one processor to encode the frame-level bounding box using a first bitrate and Includes, The aforementioned program code A second signaling code configured to cause at least one of the processors to signal the original size of the frame and the size of the frame-level bounding box as metadata in the bitstream. A device that further includes the following.

13. The aforementioned program code A second encoding code configured to cause at least one processor to encode the portion of the frame not included in the frame-level bounding box using a second bitrate. It further includes, The apparatus according to claim 11, wherein the second bitrate is lower than the first bitrate.

14. An apparatus for encoding video for machine vision and human / machine hybrid vision, wherein the apparatus is At least one memory configured to store program code, At least one processor configured to access the program code and operate as instructed by the program code, The program code is provided with, A first receive code configured to cause at least one processor to receive image data, A first detection code configured to cause at least one processor to detect a plurality of bounding boxes associated with a plurality of objects of interest in the frame of the image data, A second detection code configured to cause at least one processor to detect a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, A first encoding code configured to cause at least one processor to encode the frame-level bounding box using a first bitrate and Includes, The aforementioned program code A third encoding code configured to cause at least one processor to encode the frame using a second bitrate in response to encoding the frame-level bounding box using the first bitrate. It further includes, An apparatus in which a section of the frame contained within the frame-level bounding box is filled with a constant value before the frame is encoded using the second bitrate.

15. The apparatus according to claim 11, wherein the frames of the image data are selected based on a predetermined sample rate, or one frame is selected for each predetermined number of frames of the image data.

16. A non-temporary computer-readable medium storing computer instructions, wherein when the computer instructions are executed by at least one processor to encode video for machine vision and human / machine hybrid vision, the at least one processor, Receiving image data and To detect multiple bounding boxes associated with multiple objects of interest in the frame of the aforementioned image data, Detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, Encoding the frame-level bounding box using the first bitrate and Have them do it, The aforementioned at least one processor, For an intra-period that includes multiple temporally consecutive frames, a common intra-period level bounding box is determined based on multiple frame-level bounding boxes for the multiple temporally consecutive frames within the intra-period. The top-left and bottom-right coordinates of the common intra-period level bounding box are signaled as metadata in the bitstream. A non-temporary computer-readable medium that can perform further operations.

17. A non-temporary computer-readable medium storing computer instructions, wherein when the computer instructions are executed by at least one processor to encode video for machine vision and human / machine hybrid vision, the at least one processor shall Receiving image data and To detect multiple bounding boxes associated with multiple objects of interest in the frame of the aforementioned image data, Detecting a frame-level bounding box for the frame based on the coordinates of the plurality of bounding boxes, Encoding the frame-level bounding box using the first bitrate and Have them do it, The aforementioned at least one processor, Signaling the original size of the frame and the size of the frame-level bounding box as metadata in the bitstream. A non-temporary computer-readable medium that can perform further operations.

18. A program that causes at least one processor to perform the method according to any one of claims 1 to 10.