Region of interest coding for VCM
The method addresses the inefficiency of existing codecs by using bounding box detection and region of interest coding to enhance video encoding for machine and hybrid human/machine vision, improving efficiency and performance in object detection and tracking.
Patent Information
- Application Number
- JP2025517758
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-28
- Filing Date
- 2023-09-29
- Publication Date
- 2025-09-29
- Estimated Expiration
- 2043-09-29
AI Technical Summary
Existing video codecs are primarily designed for human consumption and fail to efficiently encode video for machine vision tasks such as object detection, segmentation, and tracking, necessitating the development of codecs tailored for machine and hybrid human/machine vision.
A method and apparatus for encoding video that involves detecting bounding boxes for objects of interest, calculating a frame-level bounding box, and encoding using a first bit rate, with spatial and temporal region of interest coding techniques to reduce bit rate, incorporating learning-based codecs with legacy codecs to form hybrid video codecs.
Enhances video encoding efficiency for machine and hybrid human/machine vision by reducing bit rate and improving performance in object detection and tracking tasks.
Smart Images

Figure 2025532199000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority from U.S. Provisional Application No. 63 / 412,367, filed September 30, 2022, and U.S. Application No. 18 / 477,189, filed September 28, 2023, in the United States Patent and Trademark Office, the entire disclosures of which are incorporated herein by reference.
[0002]
[0002] This disclosure relates to video coding for machine vision. In particular, techniques for encoding video for machine vision and human / machine hybrid vision are disclosed. [Background technology]
[0003]
[0003] Traditionally, videos or images are consumed by humans for a variety of uses, such as entertainment, education, etc. Therefore, video coding or image coding often exploits the characteristics of the human visual system for better compression efficiency while maintaining good subjective quality.
[0004]
[0004] In recent years, with the increase in machine learning applications, many intelligent platforms, together with a large number of sensors, utilize video for machine vision tasks, such as object detection, segmentation, or tracking. How to encode video or images for consumption by machine tasks has become an interesting and challenging problem, which has led to the introduction of Video Coding for Machines (VCM) research. To achieve this goal, the international standards group MPEG created an ad hoc group, "Video Coding for Machines (VCM)," to standardize related techniques for better interoperability between different devices.
[0005]
[0005] Existing video codecs are primarily intended for human consumption. However, an increasing amount of video is being consumed by machines for machine vision tasks such as object detection, instance segmentation, and object tracking. It is important to develop video codecs that efficiently encode video for machine vision or hybrid machine / human vision. Summary of the Invention
[0006] The following presents a simplified summary of one or more embodiments of the present disclosure in order to provide a basic understanding of such embodiments. This summary is not an extensive overview of all contemplated embodiments, nor is it intended to identify key or critical elements of all embodiments or to delineate the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of the present disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0007]
[0007] Methods, apparatus, and non-transitory computer-readable media relating to encoding video for machine vision and hybrid human / machine vision.
[0008] A method for encoding video for machine vision and human / machine hybrid vision may be provided. The method may be executed by one or more processors and may include receiving image data, detecting a plurality of bounding boxes associated with a plurality of objects of interest in a frame of the image data, detecting a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes, and encoding the frame-level bounding box using a first bit rate.
[0009] An apparatus for encoding video for machine vision and human / machine hybrid vision may be provided. The apparatus may include at least one memory configured to store program code and at least one processor configured to access the program code. The at least one processor may be configured to operate as instructed by the program code, the program code including first receiving code configured to cause the at least one processor to receive image data, first detection code configured to cause the at least one processor to detect a plurality of bounding boxes associated with a plurality of objects of interest in a frame of the image data, second detection code configured to cause the at least one processor to detect a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes, and first encoding code configured to cause the at least one processor to encode the frame-level bounding box using a first bit rate.
[0010]
[0010] A non-transitory computer-readable medium may be provided that stores computer instructions that, when executed by at least one processor for encoding video for machine vision and human / machine hybrid vision, cause the at least one processor to receive image data, detect a plurality of bounding boxes associated with a plurality of objects of interest in a frame of the image data, detect a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes, and encode the frame-level bounding box using a first bit rate.
[0011]
[0011] Additional embodiments will be set forth in the description that follows, and in part will be obvious from the description, and / or may be learned by practice of presented embodiments of the present disclosure.
[0012]
[0012] The above and other features and aspects of embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0013] [Figure 1]
[0013] FIG. 1 is a diagram of an exemplary network device according to various embodiments of the present disclosure. [Figure 2]
[0014] FIG. 1 illustrates an architecture of the disclosed hybrid video codec, according to one embodiment of the present disclosure. [Figure 3]
[0015] FIG. 1 is a diagram of video coding for a machine system according to various embodiments of the present disclosure. [Figure 4A]
[0016] FIG. 2 is a diagram of a spatial region of interest bounding box, according to various embodiments of the present disclosure. [Figure 4B]
[0017] 10A-10C illustrate region of interest bounding box calculations according to various embodiments of the present disclosure. [Figure 5]
[0018] FIG. 10 illustrates region of interest bounding box calculation for n intra-periods according to various embodiments of the present disclosure. [Figure 6]
[0019] 1 is a flowchart of an exemplary process for determining a region of interest in a hybrid video codec, according to various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014]
[0020] The following detailed description of the exemplary embodiments refers to the accompanying drawings, in which the same reference numbers in different drawings may identify the same or similar elements.
[0015]
[0021] The above disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations. Furthermore, one or more features or components of some embodiments may be incorporated into or combined with other embodiments (or one or more features of some embodiments). Furthermore, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be permuted.
[0016]
[0022] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods does not limit the implementation. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0017]
[0023] Although particular combinations of features are recited in the claims and / or disclosed herein, these combinations do not limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or disclosed herein. Although each dependent claim set forth below may depend directly on only one claim, the disclosure of possible implementations includes each dependent claim in combination with every other claim in the claim set.
[0018]
[0024] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Where only one item is intended, the term "one" or similar phrases are used. Also, as used herein, terms such as "has," "have," "having," "include," and "including" are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based at least in part on," unless otherwise specified. Furthermore, phrases such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include A only, B only, or both A and B.
[0019]
[0025] References throughout this specification to "some embodiments," "an embodiment," or similar phrases mean that a particular feature, structure, or characteristic described with respect to the illustrated embodiment is included in some embodiments of the solution. Thus, throughout this specification, the phrases "in some embodiments," "in an embodiment," and similar phrases may, but do not necessarily, all refer to the same embodiment.
[0020]
[0026] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. In light of the description herein, those skilled in the art will recognize that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other cases, additional features and advantages may be recognized in some embodiments, which may not be present in all embodiments of the present disclosure.
[0021]
[0027] The disclosed methods may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0022]
[0028] Embodiments of the present disclosure are directed to video coding for machines. In particular, methods are disclosed for encoding video for machine vision and hybrid human / machine vision. Traditional video codecs are designed for human consumption. In some embodiments, traditional video codecs can be combined with learning-based codecs to form hybrid codecs such that video can be efficiently coded for machine vision and hybrid human and machine vision.
[0023]
[0029] 1 is a diagram of an exemplary device for implementing a translation service. Device 100 may correspond to any type of known computer, server, or data processing device. For example, device 100 may comprise a processor, a personal computer (PC), a printed circuit board (PCB) with computing device, a minicomputer, a mainframe computer, a microcomputer, a telephone computing device, a wired / wireless computing device (e.g., a smartphone, a personal digital assistant (PDA)), a laptop, a tablet, a smart device, or any other similar operating device.
[0024]
[0030] In some embodiments, as shown in FIG. 1 , device 100 may include a set of components such as a processor 120, a memory 130, a storage component 140, an input component 150, an output component 160, and a communication interface 170.
[0025]
[0031] Bus 110 may comprise one or more components that enable communication between a set of components of device 100. For example, bus 110 may be a communication bus, a crossover bar, a network, etc. Although bus 110 is depicted in FIG. 1 as a single line, bus 110 may be implemented using multiple (two or more) connections between a set of components of device 100. The disclosure is not limited in this respect.
[0026]
[0032] Device 100 may comprise one or more processors, such as processor 120. Processor 120 may be implemented in hardware, firmware, and / or a combination of hardware and software. For example, processor 120 may comprise a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field programmable gate array (FPGA), application specific integrated circuit (ASIC), general-purpose single-chip or multi-chip processor, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the operations described herein. A general-purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. Processor 120 may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. In some embodiments, particular processes and methods may be performed by circuitry that is specific to a given operation.
[0027]
[0033] The processor 120 may control the overall operation of the device 100 and / or a set of components of the device 100 (e.g., the memory 130, the storage components 140, the input components 150, the output components 160, and the communication interface 170).
[0028]
[0034] Device 100 may further comprise memory 130. In some embodiments, memory 130 may comprise random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic memory, optical memory, and / or another type of dynamic or static storage device. Memory 130 may store information and / or instructions for use (e.g., execution) by processor 120.
[0029]
[0035] Storage component 140 of device 100 may store information and / or computer-readable instructions and / or code related to the operation and use of device 100. For example, storage component 140 may include a hard disk (e.g., a magnetic disk, optical disk, magneto-optical disk, and / or solid-state disk), a compact disk (CD), a digital versatile disk (DVD), a universal serial bus (USB) flash drive, a Personal Computer Memory Card International Association (PCMCIA) card, a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.
[0030]
[0036] Device 100 may further comprise input component 150. Input component 150 may include one or more components that enable device 100 to receive information, such as via user input (e.g., a touchscreen, a keyboard, a keypad, a mouse, a stylus, a button, a switch, a microphone, a camera, etc.). Alternatively or additionally, input component 150 may include sensors for detecting information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.).
[0031]
[0037] Output components 160 of device 100 may include one or more components (e.g., a display, a liquid crystal display (LCD), a light emitting diode (LED), an organic light emitting diode (OLED), a haptic feedback device, a speaker, etc.) that may provide output information from device 100.
[0032]
[0038] Device 100 may further comprise a communication interface 170. Communication interface 170 may include a receiver component, a transmitter component, and / or a transceiver component. Communication interface 170 may enable device 100 to establish a connection and / or transfer communications with other devices (e.g., a server, another device). The communications may be implemented via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 170 may enable device 100 to receive information from and / or provide information to another device. In some embodiments, communication interface 170 may provide for communication with another device over a network, such as a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, the Internet, a fiber optic-based network, a cellular network (e.g., a fifth generation (5G) network, a long term evolution (LTE) network, a third generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a telephone network (e.g., a public switched telephone network (PSTN)), etc., and / or a combination of these or other types of networks. Alternatively or additionally, communication interface 170 may provide for communication with another device over a device-to-device (D2D) communication link, such as FlashLinQ, WiMedia, Bluetooth, ZigBee, Wi-Fi, LTE, 5G, etc. In other embodiments, communication interface 170 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, or the like.
[0033]
[0039] Device 100 may be included in core network 240 and may perform one or more processes described herein. Device 100 may perform operations based on processor 120 executing computer-readable instructions and / or code, which may be stored by a non-transitory computer-readable medium, such as memory 130 and / or storage component 140. A computer-readable medium may refer to a non-transitory memory device. A memory device may include memory space within a single physical storage device and / or memory space spread across multiple physical storage devices.
[0034]
[0040] Computer-readable instructions and / or code may be loaded into memory 130 and / or storage component 140 from another computer-readable medium or from another device via communications interface 170. The computer-readable instructions and / or code stored in memory 130 and / or storage component 140, when executed by processor 120, may cause device 100 to perform one or more processes described herein.
[0035]
[0041] Alternatively, or in addition, hardwired circuitry may be used in place of or in combination with software instructions to implement one or more processes described herein. Thus, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.
[0036]
[0042] The number and arrangement of components shown in Figure 1 are provided as an example. In practice, there may be additional, fewer, different, or differently arranged components than those shown in Figure 1. Furthermore, two or more components shown in Figure 1 may be implemented within a single component, or a single component shown in Figure 1 may be implemented as multiple distributed components. Additionally or alternatively, a set of components shown in Figure 1 may perform one or more operations described as being performed by another set of components shown in Figure 1.
[0037]
[0043] Figure 2 is a block diagram of one embodiment of a hybrid video codec 200. The hybrid video codec 200 may include a legacy codec 220 and a learning-based codec 230. The input 201 to the hybrid codec may be video or may be an image, since an image may be treated as a special type of video (e.g., a video with one image). In Figure 2, the legacy video codec 220 may be used to compress the input video 201 at a different scale (e.g., original resolution or downsampled), and the downsampling ratio of the downsampling module 210 may be fixed and known in both the encoder 221 and the decoder 223, or the downsampling ratio may be user-defined, e.g., 100% (e.g., no downsampling is performed), 50%, 25%, etc., and may be sent as metadata in the bitstream 224 to inform the decoder 222. The legacy video codec may be an image codec such as VVC, HEVC, H264, or JPEG, JPEG2000. The downsampling module 210 may be a classical image downsampler or a learning-based image downsampler. The decoded downsampled video 203 (e.g., "low-resolution video 203" in FIG. 2) may be upsampled to the original resolution of the video (e.g., "high-resolution video 204") that can be used for human vision using the upsampling module 250. The upsampling module 250 may be a classical image upsampler or a learning-based image upsampler, such as a learning-based super-resolution module.
[0038]
[0044] In some embodiments, the hybrid video codec 200 may also employ a learning-based video codec 230 to compress the downsampled video 202 .
[0039]
[0045] In the encoder, a reconstructed video 203 may also be generated and upsampled to the original input resolution. The upsampled reconstructed video 205 may then be subtracted from the input video to generate a residual video signal 202, which may be fed to the learning-based codec 230 of FIG. 2. The upsampling module 240 of the hybrid video codec 200 may be the same as the upsampling module 250 (after the low resolution). The video 203 may be decoded in the decoder. The output of the residual decoder 238 may be added on top of the high-resolution video 204 to form a reconstructed video 205, which may be used for machine vision tasks.
[0040]
[0046] 3 shows an embodiment of an architecture for a video coding machine (VCM), such as hybrid video codec 200. Sensor output 300 follows a video encoding path 311, through a VCM encoder 310, to a VCM decoder 320 where video decoding 321 occurs. Another path is to feature extraction 312, to feature transformation 313, to feature encoding 314, and to feature decoding 322. The output of VCM decoder 320 is primarily for machine consumption, i.e., machine vision 305. In some cases, it may also be used for human vision 306. One or more machine tasks for understanding the video content are then performed.
[0041]
[0047] As is known, video content often contains a lot of information, for example, moving objects in the foreground and static scenes in the background. Classical video codecs often consider temporal or specific redundancies to compress content. In video coding for machines, machine vision tasks are often object detection, instance segmentation, or object tracking. In these types of tasks, objects such as people, cars, and bicycles are of primary interest, while other information such as trees, grass, and sky is of little interest. It would be beneficial to consider these types of observations to further reduce the information that needs to be transmitted or stored. Therefore, region of interest coding is disclosed in this application.
[0042]
[0048] The proposed methods may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. In this disclosure, picture, image, and frame are interchangeable.
[0043]
[0049] According to an embodiment, two types of ROI methods are described, eg, spatial region of interest (ROI) and temporal ROI.
[0044]
[0050] Spatial Region of Interest Coding for Image Coding
[0045]
[0051] According to one embodiment, at the encoder side, an object detector may be used to detect the bounding boxes of all relevant objects in a frame / image (image or frame may be used interchangeably in this disclosure), and then a region of interest bounding box may be calculated to contain all of the bounding boxes of the relevant objects. Figure 4A is an example diagram of a region of interest bounding box determined for a frame 400 that encapsulates multiple objects of interest.
[0046]
[0052] As shown in Figure 4A, pedestrians and people riding bicycles or motorcycles are marked with bounding boxes of relatively dark boxes. The outer white box is the ROI bounding box that contains the bounding boxes of all people in the scene.
[0047]
[0053] According to one embodiment, for a frame, the bounding boxes of all involved objects are
[0048]
number
[0049]
number
[0050]
number
[0051]
number
[0052]
[0054] In the above equation,
[0053]
number
[0054]
number
[0055]
number
[0056]
number
[0057] The calculation of the ROI bounding box is shown in Figure 4B. As shown in Figure 4B, the black boxes indicate the objects of interest (e.g., 465-1, 465-2), and the dotted box 460 indicates the box that exactly contains the object of interest. The dashed box 455 is the ROI bounding with margins added.
[0058] According to one embodiment, after calculating the ROI bounding box (e.g., 455), in order to reduce the bit rate for transmission or storage, only the content within the ROI bounding box may be encoded to obtain a bit stream of the image. For example, a first higher bit rate may be used to encode only the content within the ROI box.
[0059] In the same or another embodiment, frames without ROI bounding boxes may be coded using low bitrate configurations, such as high quantization step sizes or reduced scales, if their content may be of interest in the future. As an example, for coding frames that simply do not have ROI bounding boxes, the ROI region may be filled with a constant value, such as 0 or 128.
[0060]
[0058] In the same or other embodiments, the entire frame may be encoded using a low bitrate configuration, such as a high quantization step size or a reduced scale.
[0061] According to one embodiment, at the decoder side, a picture or frame containing an ROI bounding box can be decoded and used alone or can be combined with frames without ROI portions to form a final picture frame. In some embodiments, a padding method can be used to extend the picture or frame boundary to the same size of the original picture (with the area outside the ROI bounding box). In embodiments where the original size of the picture needs to be restored, the original size, ROI bounding, can be sent as metadata in the bitstream.
[0062]
[0060] Thus, as an example embodiment, a method for spatial ROI disclosed herein may include receiving image data including multiple frames from a sensor. Then, for a frame in the image data, multiple bounding boxes associated with multiple objects of interest in the frame may be detected. A frame-level bounding box for the frame may be detected based on coordinates of the multiple bounding boxes. The frame-level bounding box may be encoded using a first bit rate. In embodiments, a section of the frame not included in the frame-level bounding box may be encoded using a second bit rate, the second bit rate being lower than the first bit rate. In embodiments, in response to encoding the frame-level bounding box using the first bit rate, the frame, or at least the remainder of the frame, may be encoded using a second bit rate. In some embodiments, the section of the frame included in the frame-level bounding box may be filled using a constant value prior to encoding the frame using the second bit rate.
[0063] Spatial Region of Interest Coding for Video Coding
[0064] In embodiments, already coded frames may be used as references to predict the current frame. In such embodiments, ROI bounding box information for each frame may be carried as metadata in the bitstream to enable the decoder to place the decoded ROI content in the correct position in the original image to facilitate motion compensation.
[0065]
[0063] According to one embodiment, a video sequence can be divided into multiple intra-periods. As an example, an intra-period can be approximately 1 second or 2 seconds apart, which can correspond to 32 frames or 64 frames for a video with a frame rate of 30 frames / second. According to one embodiment, for one intra-period, ROI bounding boxes for all frames can be calculated. A common ROI bounding box size can be calculated, which includes all individual ROI bounding boxes with a certain margin.
[0066]
[0064] Thus, according to one embodiment, a method disclosed herein may include, for an intra-period including a plurality of temporally consecutive frames, determining a common intra-period level bounding box based on a plurality of frame-level bounding boxes for the plurality of temporally consecutive frames in the intra-period. The method may also include signaling the upper-left and lower-right coordinates of the common intra-period level bounding box as metadata in the bitstream. In some embodiments, an intra-period may include a predetermined number of temporally consecutive frames.
[0067]
[0065] As mentioned above, certain information related to a frame can be signaled as metadata. As an example, the original size of the frame and the size of the frame-level bounding box can be signaled as metadata. When signaling the size of the frame-level bounding box, the upper-left coordinate and the lower-right coordinate of the frame-level bounding box can be signaled.
[0068]
[0066] Figure 5 shows a common ROI bounding box for an intra period, i.e., a black box (505-1, 505-1, ... 505-N). Due to the use of a common ROI bounding box for an intra period, only the common bounding box information, i.e., its top-left and bottom-right coordinates, are sent in the bitstream as metadata.
[0069] According to one embodiment, only the content within the common ROI bounding for each frame is sent in the bitstream or as metadata for each intra period or for a specific number of frames. In such an embodiment, regular motion prediction / compensation can be performed because all frames in an intra period have the same resolution. Thus, in an embodiment, the frames of image data to be analyzed for the ROI can be selected based on a predetermined sample rate or one for every predetermined number of frames of image data. In an embodiment, the predetermined number of frames varies based on the complexity of the scene. In an embodiment, some additional information related to the frame, such as the predetermined sample rate or the predetermined number of frames, can be signaled in the bitstream as metadata.
[0070] In the same or another embodiment, the original frame without content in the ROI region may also be coded and sent in a lower bitrate configuration, i.e., with a larger quantization step or a lower scale, etc. In an embodiment, the ROI region may be filled with a constant value, such as 0 or 128.
[0071]
[0069] In the same or another embodiment, the original frame may also be coded and sent in a lower bit rate configuration, ie, with a larger quantization step or lower scale, etc.
[0072] According to one embodiment, at a decoder, a picture or frame containing an ROI bounding box can be decoded and used alone or can be combined with a frame without an ROI portion to form a final picture. In an embodiment, a padding method can be used to extend the picture boundary to the same size as the original picture (with the area outside the ROI bounding box). In an embodiment where video reconstruction at the original resolution is required, the original video resolution can be sent as metadata.
[0073] Temporal Region of Interest Coding for Video Coding
[0074] As is known, consecutive video frames may share a lot of common information. For example, consecutive frames may contain the same object with slightly different positions or shapes. This observation can be exploited in video coding for machines to reduce transmission rates or storage. The encoder may utilize an analysis module to determine motion characteristics of a video segment.
[0075] In the same or another embodiment, video frames may be temporally downsampled, i.e., only one frame out of every N frames may be selected for encoding and transmission / storage. N may be set as a fixed value, or it may be variable. For example, N may be small, e.g., 2, for dynamic scenes, or relatively large, e.g., 4 or 8, for relatively static content.
[0076] In the same or another embodiment, instead of uniformly subsampling the video in the time domain, portions of the video may be selected for coding and transmission / storage if an action or event occurs in a particular portion. Portions of the content may be ignored without coding since there is no action / event in the particular portion. The remainder of the content may be uniformly sampled.
[0077] In the same or other embodiments, the video frames are non-uniformly sampled, for example, for the first 12 frames, frame numbers 0, 3, 7, 8, 11 are chosen for encoding.
[0078]
[0076] Metadata such as sample rate N may be sent in the bitstream to signal timing information of the coded content. If a portion of the video uses the same sample rate N, the sample rate N is sent only at the beginning of the bitstream for that portion. For unevenly sampled video, metadata such as frame numbers or timestamps may be sent in the bitstream.
[0079] In situations where video frames with the original frame rate need to be restored, duplicated pictures may be inserted into the decoded video. For example, if only frames with frame numbers 0, 3, 7, 8, and 1 are sent, they may be denoted as f0, f3, f7, f8, and f11, respectively. The decoder may then duplicate those frames as, for example, f0, f0, f0, f3, f3, f3, f3, f7, f8, f8, f8, f11, etc., to obtain 12 frames.
[0080]
[0078] In embodiments, to further reduce bit rates for transmission or storage, embodiments may combine spatial and temporal ROI methods as disclosed herein in video coding for machines.
[0081] FIG. 6 shows a process 600 for encoding video for machine vision and hybrid human / machine vision, according to one embodiment.
[0082] 6, image data may be received in operation 605. In an embodiment, the image data may be received over a network or may be received from a sensor.
[0083]
[0081] In operation 610, multiple bounding boxes associated with multiple objects of interest may be determined in one frame of image data.
[0084] In embodiments, the frames may be selected based on a predetermined sample rate or may be selected one for every predetermined number of frames of image data, in embodiments, the predetermined number of frames may vary based on the complexity of the scene.
[0085]
[0083] In operation 615, a frame-level bounding box may be detected for the frame based on the coordinates of the multiple bounding boxes.
[0086]
[0084] At operation 620, the frame-level bounding box may be encoded using a first bit rate.
[0087] At operation 625, a section of the frame not included in the frame-level bounding box may be encoded using a second bit rate. In an embodiment, the second bit rate may be lower than the first bit rate. In an embodiment, in response to encoding the frame-level bounding box using the first bit rate and / or in response to encoding the frame using the second bit rate, a section of the frame included in the frame-level bounding box may be filled using a constant value prior to encoding the frame using the second bit rate.
[0088] In an embodiment, for an intra-period that includes multiple temporally consecutive frames, a common intra-period level bounding box may be determined based on multiple frame-level bounding boxes for the multiple temporally consecutive frames in the intra-period, and the upper-left and lower-right coordinates of the common intra-period level bounding box may be signaled in the bitstream as metadata. An intra-period consists of a predetermined number of temporally consecutive frames.
[0089]
[0087] Other metadata that may be signaled in the bitstream may include the original size of the frame, the size of the frame-level bounding box, the top-left and bottom-right coordinates of the frame-level bounding box, a predetermined sample rate or the predetermined number of frames mentioned above.
[0090]
[0088] The above disclosure provides illustration and description, but is not intended to be exhaustive or to limit implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations.
[0091] It should be understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed herein is illustrative of example approaches. Based on design preferences, it should be understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Additionally, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in an example order and are not limited to the specific order or hierarchy presented.
[0092]
[0090] Some embodiments may relate to a system, a method, and / or a computer-readable medium at a level that integrates any possible technical details. Furthermore, one or more of the above components described above may be implemented as instructions stored in a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include one or more computer-readable non-transitory storage media having computer-readable program instructions thereon for causing the processor to perform operations.
[0093]
[0091] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves with instructions recorded thereon, and any suitable combination of the above. As used herein, computer-readable storage media should not be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating in a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted through wires.
[0094] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0095] The computer-readable program code / instructions for performing operations may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects or operations.
[0096]
[0094] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executing via the processor of the computer or other programmable data processing apparatus create means for implementing the operations specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0097]
[0095] Computer-readable program instructions may also be loaded into a computer, other programmable apparatus, or other device to cause a series of operations to be performed on the computer, other programmable data processing apparatus, or other device to create a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0098] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, comprising one or more executable instructions for implementing a specified logical operation(s). The methods, computer systems, and computer-readable media may include additional, fewer, different, or differently arranged blocks than those shown in the figures. In some alternative implementations, the operations in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may be executed virtually concurrently or substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on the functionality involved. Each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified operations or executes a combination of dedicated hardware and computer instructions.
[0099] It will be apparent that the systems and / or methods described herein may be implemented in different forms, such as hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods does not limit the implementation. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
Claims
1. 1. A method for encoding video for machine vision and hybrid human / machine vision, the method being performed by one or more processors, the method comprising: receiving image data; Detecting a plurality of bounding boxes associated with a plurality of objects of interest in the frame of image data; Detecting a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes; encoding the frame-level bounding box using a first bit rate; A method comprising:
2. The method comprises: For an intra-period including a plurality of temporally consecutive frames, determining a common intra-period level bounding box based on a plurality of frame-level bounding boxes for the plurality of temporally consecutive frames in the intra-period; signaling the top-left and bottom-right coordinates of said common intra-period level bounding box as metadata in the bitstream; The method of claim 1 further comprising:
3. The method of claim 2 , wherein the intra-period comprises a predetermined number of temporally consecutive frames.
4. The method comprises: signaling the original size of the frame and the size of the frame-level bounding box as metadata in the bitstream. The method of claim 1 further comprising:
5. The method of claim 4 , wherein signaling the size of the frame-level bounding box comprises signaling a top-left coordinate and a bottom-right coordinate of the frame-level bounding box.
6. The method comprises: encoding a section of the frame not included in the frame-level bounding box using a second bit rate; further comprising The method of claim 1 , wherein the second bit rate is lower than the first bit rate.
7. The method comprises: encoding the frame using a second bit rate in response to encoding the frame-level bounding box using the first bit rate. further comprising 2. The method of claim 1, wherein the section of the frame that is included in the frame-level bounding box is filled using a constant value prior to encoding the frame using the second bit rate.
8. The method comprises: encoding the frame using a second bit rate; further comprising The method of claim 1 , wherein the second bit rate is lower than the first bit rate.
9. The method of claim 1 , wherein the frames of the image data are selected based on a predetermined sample rate or are selected one out of every predetermined number of frames of the image data.
10. The method of claim 9 , wherein the predetermined number of frames varies based on scene complexity.
11. The method comprises: signaling the predetermined sample rate or the predetermined number of frames as metadata in a bitstream. The method of claim 10 further comprising:
12. 1. An apparatus for encoding video for machine vision and human / machine hybrid vision, said apparatus comprising: at least one memory configured to store program code; at least one processor configured to access said program code and to operate as instructed by said program code; wherein the program code: a first receiving code configured to cause the at least one processor to receive image data; first detection code configured to cause the at least one processor to detect a plurality of bounding boxes associated with a plurality of objects of interest in the frame of image data; second detection code configured to cause the at least one processor to detect a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes; first encoding code configured to cause the at least one processor to encode the frame-level bounding box using a first bit rate; and Including, Device.
13. The program code is first determination code configured to cause the at least one processor to determine, for an intra-period including a plurality of temporally consecutive frames, a common intra-period level bounding box based on a plurality of frame-level bounding boxes for the plurality of temporally consecutive frames in the intra-period; first signaling code configured to cause the at least one processor to signal top-left and bottom-right coordinates of the common intra-period level bounding box as metadata in a bitstream; and The apparatus of claim 12 further comprising:
14. The program code: second signaling code configured to cause the at least one processor to signal the original size of the frame and the size of the frame-level bounding box as metadata in a bitstream; The apparatus of claim 12 further comprising:
15. The program code: second encoding code configured to cause the at least one processor to encode sections of the frame not included in the frame-level bounding box using a second bit rate; further comprising The apparatus of claim 12 , wherein the second bit rate is lower than the first bit rate.
16. The program code: third encoding code configured to cause the at least one processor to encode the frame using a second bit rate in response to encoding the frame-level bounding box using the first bit rate. further comprising 13. The apparatus of claim 12, wherein the section of the frame included in the frame-level bounding box is filled using a constant value prior to encoding the frame using the second bit rate.
17. The apparatus of claim 12 , wherein the frames of the image data are selected based on a predetermined sample rate or are selected one out of every predetermined number of frames of the image data.
18. 1. A non-transitory computer-readable medium having stored thereon computer instructions, the computer instructions, when executed by at least one processor for encoding video for machine vision and hybrid human-machine vision, causing the at least one processor to: receiving image data; Detecting a plurality of bounding boxes associated with a plurality of objects of interest in the frame of image data; Detecting a frame-level bounding box for the frame based on coordinates of the plurality of bounding boxes; encoding the frame-level bounding box using a first bit rate; A non-transitory computer-readable medium for causing
19. the at least one processor: For an intra-period including a plurality of temporally consecutive frames, determining a common intra-period level bounding box based on a plurality of frame-level bounding boxes for the plurality of temporally consecutive frames in the intra-period; signaling the top-left and bottom-right coordinates of said common intra-period level bounding box as metadata in the bitstream; 20. The non-transitory computer-readable medium of claim 18, further configured to:
20. the at least one processor: signaling the original size of the frame and the size of the frame-level bounding box as metadata in the bitstream.
20. The non-transitory computer-readable medium of claim 18, further configured to:
21. A program causing at least one processor to carry out the method of any one of claims 1 to 11.
Citation Information
Patent Citations
Image processor
JP2006093784A
Image coding method and image decoding method considering human visual characteristic
JP2007020199A
Patch-based video coding for machines
JP2023521553A
Encoding video frames using generated region of interest maps
US20190007690A1
Patch based video coding for machines
WO2021211884A1