Methods, apparatus, and media for video coding in machine vision and human-machine hybrid vision

CN118235402BActive Publication Date: 2026-09-01TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202380014317.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-09-28
Filing Date
2023-09-29
Publication Date
2026-09-01
Estimated Expiration
2043-09-29

Smart Images

  • Figure CN118235402B_ABST
    Figure CN118235402B_ABST
Patent Text Reader

Abstract

A video coding technique for machine vision and human-machine hybrid vision includes receiving image data. The technique may further include detecting multiple bounding boxes associated with multiple objects of interest in a frame of the image data, and detecting frame-level bounding boxes of the frame based on the coordinates of the multiple bounding boxes. The technique may then include encoding the frame-level bounding boxes using a first bitrate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 412,367, filed September 30, 2022, with the U.S. Patent and Trademark Office, and U.S. Patent Application No. 18 / 477,189, filed September 28, 2023, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure generally relates to video coding for machine vision. More specifically, it relates to video coding for machine vision and human-machine hybrid vision, and more particularly to methods, apparatus, and media for video coding for machine vision and human-machine hybrid vision. Background Technology

[0004] Traditionally, videos or images are used by humans for various purposes, such as entertainment and education. Therefore, video coding or image coding often leverages features of the human visual system to achieve better compression efficiency while maintaining good subjective quality.

[0005] In recent years, with the rise of machine learning applications and the proliferation of sensors, many intelligent platforms have utilized video for machine vision tasks such as object detection, segmentation, or tracking. How to encode video or images for machine tasks has become an interesting and challenging problem, leading to the introduction of research into Video Coding for Machine (VCM). To achieve this goal, the international standards organization MPEG (Moving Picture Experts Group) created an ad hoc organization, "Video Coding for Machine (VCM)," to standardize related technologies and improve interoperability between different devices.

[0006] Existing video codecs are primarily designed for human use. However, an increasing number of videos are being used by machines for machine vision tasks, such as object detection, instance segmentation, and object tracking. Therefore, it is crucial to develop a video codec that effectively encodes video for machine vision or a hybrid machine / human vision approach. Summary of the Invention

[0007] The following is a simplified summary of one or more embodiments of this disclosure to provide a basic understanding of these embodiments. This summary is not a broad overview of all contemplated embodiments and is intended neither to identify essential or critical elements of all embodiments nor to describe the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments of this disclosure in a simplified form as a prelude to the specific embodiments presented thereafter.

[0008] Methods, apparatus, and non-transitory computer-readable media for video coding of machine vision and human-machine hybrid vision.

[0009] A method for video encoding for machine vision and human-machine hybrid vision can be provided. The method can be executed by one or more processors and can include receiving image data; detecting multiple bounding boxes associated with multiple objects of interest in frames of the image data; detecting frame-level bounding boxes of the frames based on the coordinates of the multiple bounding boxes; and encoding the frame-level bounding boxes using a first bitrate.

[0010] An apparatus for video encoding for machine vision and human-machine hybrid vision can be provided. The apparatus may include at least one memory configured to store program code; and at least one processor configured to access the program code. The at least one processor may be configured to operate according to instructions of the program code, the program code including: first receiving code configured to cause the at least one processor to receive image data; first detection code configured to cause the at least one processor to detect multiple bounding boxes associated with multiple objects of interest in a frame of the image data; second detection code configured to cause the at least one processor to detect frame-level bounding boxes of the frame based on the coordinates of the multiple bounding boxes; and first encoding code configured to cause the at least one processor to encode the frame-level bounding boxes using a first bitrate.

[0011] A non-transitory computer-readable medium may be provided having computer instructions stored thereon that, when executed by at least one processor for video encoding of machine vision and human-machine hybrid vision, cause at least one processor to receive image data; detect multiple bounding boxes associated with multiple objects of interest in frames of the image data; detect frame-level bounding boxes of the frames based on the coordinates of the multiple bounding boxes; and encode the frame-level bounding boxes using a first bitrate.

[0012] Additional embodiments will be set forth in the description which follows, and will be apparent in part from the description, and / or may be learned by practice of the embodiments presented in this disclosure. Attached Figure Description

[0013] The above and other features and aspects of embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings, in which: Figure 1 This is a schematic diagram of an exemplary network device according to various embodiments of the present disclosure.

[0014] Figure 2 An architecture of a disclosed hybrid video codec according to embodiments of the present disclosure is shown.

[0015] Figure 3 It is a video encoding of a machine system according to various embodiments of the present disclosure.

[0016] Figure 4A This is a schematic diagram of the spatial region of interest bounding box according to various embodiments of the present disclosure.

[0017] Figure 4B This is a schematic diagram illustrating the calculation of the region of interest bounding box according to various embodiments of the present disclosure.

[0018] Figure 5 This is a schematic diagram illustrating the calculation of the region of interest bounding box for an intra-frame time period according to various embodiments of the present disclosure.

[0019] Figure 6 This is a flowchart of an example process for determining a region of interest in a hybrid video codec according to various embodiments of the present disclosure. Specific Implementation The following detailed description of exemplary embodiments is provided with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0021] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the implementation to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or may be obtained from practice of the implementation. Furthermore, one or more features or components in some embodiments may be incorporated into or combined with some embodiments (or one or more features of some embodiments). Moreover, in the flowcharts and operational descriptions provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.

[0022] It is obvious that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, this document describes the operation and behavior of the systems and / or methods without referring to any specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the specifications herein.

[0023] Even if a specific combination of features is recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible embodiments includes combinations of each dependent claim with each other claim in the claim book.

[0024] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the article “a / an” is intended to include one or more items and may be used interchangeably with “one or more.” The term “one” or similar language is used if intended to refer to only one item. Additionally, as used herein, the terms “has / have / having,” “include / including,” etc., are open-ended terms. Furthermore, unless explicitly stated otherwise, the phrase “based on” means “at least partially based on.” Moreover, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0025] References to "some embodiments," "one embodiment," or similar language throughout this specification mean that a particular feature, structure, or characteristic described in connection with the illustrated embodiments is included in some embodiments of this solution. Therefore, the phrases "some embodiments," "one embodiment," and similar language throughout this specification may, but do not necessarily, refer to the same embodiment.

[0026] Furthermore, in one or more embodiments, the features, advantages, and characteristics described herein can be combined in any suitable manner. Based on the description herein, those skilled in the art will recognize that this disclosure can be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of this disclosure may be recognized in certain embodiments.

[0027] The disclosed methods can be used individually or in any combination in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders can be implemented using processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0028] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0029] Embodiments of this disclosure relate to video coding for machines. Specifically, video coding methods for machine vision and hybrid human-machine vision are disclosed. Conventional video codecs are designed for human use. In some embodiments, conventional video codecs can be combined with learning-based codecs to form hybrid codecs, enabling video to be efficiently encoded for both machine vision and hybrid human-machine vision.

[0030] Figure 1 This is a schematic diagram of an example device used to perform translation services. Device 100 can correspond to any type of known computer, server, or data processing device. For example, device 100 may include a processor, a personal computer (PC), a printed circuit board (PCB) including a computing device, a minicomputer, a mainframe computer, a microcomputer, a telephone computing device, a wired / wireless computing device (e.g., a smartphone, a personal digital assistant (PDA)), a laptop computer, a tablet computer, a smart device, or any other similar operating device.

[0031] In some embodiments, such as Figure 1 As shown, device 100 may include a set of components, such as processor 120, memory 130, storage component 140, input component 150, output component 160 and communication interface 170.

[0032] Bus 110 may include one or more components that allow communication between a set of components of device 100. For example, bus 110 may be a communication bus, a crossbar, a network, etc. Although bus 110 is... Figure 1 While depicted as a single line, bus 110 can be implemented using multiple (two or more) connections between a set of components of device 100. This disclosure is not limited thereto.

[0033] Device 100 may include one or more processors, such as processor 120. Processor 120 may be implemented using hardware, firmware, and / or a combination of hardware and software. For example, processor 120 may include a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), general-purpose single-chip or multi-chip processor, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the operations described herein. A general-purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. Processor 120 may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be performed by circuitry specific to a given operation.

[0034] The processor 120 can control the overall operation of the device 100 and / or a group of components of the device 100 (e.g., memory 130, storage component 140, input component 150, output component 160, and communication interface 170).

[0035] Device 100 may also include memory 130. In some embodiments, memory 130 may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic storage, optical storage, and / or another type of dynamic or static storage device. Memory 130 may store information and / or instructions for use (e.g., execution) by processor 120.

[0036] The storage component 140 of device 100 may store information and / or computer-readable instructions and / or code related to the operation and use of device 100. For example, storage component 140 may include hard disks (e.g., magnetic disks, optical disks, magneto-optical disks, and / or solid-state drives), compact discs (CDs), digital versatile discs (DVDs), universal serial bus (USB) flash drives, Personal Computer Memory Card International Association (PCMCIA) cards, floppy disks, cassette tapes, magnetic tapes, and / or another type of non-transitory computer-readable media, and corresponding drives.

[0037] Device 100 may also include an input component 150. Input component 150 may include one or more components that allow device 100 to receive information, such as via user input (e.g., touchscreen, keyboard, keypad, mouse, stylus, button, switch, microphone, camera, etc.). Alternatively or additionally, input component 150 may include sensors for sensing information (e.g., a global positioning system (GPS) component, accelerometer, gyroscope, actuator, etc.).

[0038] The output component 160 of device 100 may include one or more components that can provide output information from device 100 (e.g., display, liquid crystal display (LCD), light-emitting diode (LED), organic light-emitting diode (OLED), haptic feedback device, speaker, etc.).

[0039] Device 100 may also include a communication interface 170. Communication interface 170 may include a receiver component, a transmitter component, and / or a transceiver component. Communication interface 170 enables device 100 to establish connections and / or transmit communications with other devices (e.g., a server, another device). Communication can be achieved through wired connections, wireless connections, or a combination of wired and wireless connections. Communication interface 170 allows device 100 to receive information from and / or provide information to another device. In some embodiments, the communication interface 170 can provide communication with another device via a network, such as a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic network, a cellular network (e.g., fifth-generation (5G), long-term evolution (LTE), third-generation (3G), code division multiple access (CDMA), etc.), a public land mobile network (PLMN), a telephone network (e.g., a public switched telephone network (PSTN)), and / or a combination of these or other types of networks. Alternatively or additionally, communication interface 170 may provide communication with another device via a device-to-device (D2D) communication link, such as FlashLinQ, WiMedia (Wireless Multimedia), Bluetooth, ZigBee, Wi-Fi (Wireless Fidelity), LTE, 5G, etc. In other embodiments, communication interface 170 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, etc.

[0040] Device 100 may be included in core network 240 and performs one or more processes described herein. Device 100 may perform operations based on processor 120 executing computer-readable instructions and / or code, which may be stored by a non-transitory computer-readable medium such as memory 130 and / or storage component 140. Computer-readable medium may refer to a non-transitory storage device. Storage devices may include storage space within a single physical storage device and / or storage space distributed across multiple physical storage devices.

[0041] Computer-readable instructions and / or code may be read into memory 130 and / or storage component 140 via communication interface 170 from another computer-readable medium or from another device. The computer-readable instructions and / or code stored in memory 130 and / or storage component 140 may, if or when executed by processor 120, cause device 100 to perform one or more of the processes described herein.

[0042] Alternatively or additionally, hard-wired circuitry may be used in place of or in combination with software instructions to perform one or more of the processes described herein. Therefore, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0043] Figure 1 The number and arrangement of the components shown are provided as an example. In reality, with... Figure 1 Compared to the components shown, there may be additional components, fewer components, different components, or components with different arrangements. Furthermore, Figure 1 The two or more components shown can be implemented within a single component, or Figure 1 The single component shown can be implemented as multiple distributed components. Additionally or alternatively, Figure 1 The collection of (one or more) components shown can perform actions described as being performed by Figure 1 One or more operations performed by another set of components shown.

[0044] Figure 2 This is a block diagram of an embodiment of a hybrid video codec 200. The hybrid video codec 200 may include a conventional codec 220 and a learning-based codec 230. The input 201 to the hybrid codec may be video or an image, as an image can be considered a special type of video (e.g., a video with one image). Figure 2In this context, a conventional video codec 220 can be used to compress input video 201 at different ratios (e.g., original resolution or downsampled). The downsampling ratio of the downsampling module 210 can be fixed and known in the encoder 221 and decoder 223, or the downsampling ratio can be defined by the user, such as 100% (e.g., no downsampling), 50%, 25%, etc., and sent as metadata in the bitstream 224 to notify the decoder 222. The conventional video codec can be VVC (Versatile Video Coding), HEVC (High Efficiency Video Coding), H.264, or an image codec such as JPEG (Joint Photographic Experts Group), JPEG2000. The downsampling module 210 can be a classic image downsampler or a learning-based image downsampler. The decoded downsampled video 203 (e.g., Figure 2 The “low-resolution video 203” in the image can be upsampled to the original resolution of the video using the upsampling module 250 (e.g., “high-resolution video 204”), which can be used for human vision. The upsampling module 250 can be a classic image upsampler or a learning-based image upsampler, such as a learning-based super-resolution module.

[0045] In some embodiments, the hybrid video codec 200 may also employ a learning-based video codec 230 to compress the downsampled video 202.

[0046] In the encoder, the reconstructed video 203 can also be generated and upsampled to the original input resolution. The upsampled reconstructed video 205 can then be subtracted from the input video to generate the remaining video signal 202, which can be fed into… Figure 2 In the learning-based codec 230, the upsampling module 240 of the hybrid video codec 200 can be the same as the upsampling module 250 (after the low resolution). Video 203 can be decoded at the decoder. The output of the residual decoder 238 can be added on top of the high-resolution video 204 to form a reconstructed video 205 that can be used for machine vision tasks.

[0047] Figure 3An embodiment of the architecture for a video coding machine (VCM), such as a hybrid video codec 200, is shown. Sensor output 300 travels along the video encoding path through VCM encoder 310 to VCM decoder 320, where it undergoes video decoding 321. Another path is for feature extraction 311, feature transformation 312, feature encoding 313, and feature decoding 322. The output of VCM decoder 320 is primarily used for machine applications, i.e., machine vision 332. In some cases, the output of VCM decoder 320 can also be used for human vision 331. One or more machine tasks for understanding the video content are then performed.

[0048] As is well known, video content typically contains a large amount of information; for example, there might be a moving object in the foreground and a static scene in the background. Classical video codecs often study temporal or spatial redundancy to compress content. In machine video coding, machine vision tasks are typically object detection, instance segmentation, or object tracking. In these types of tasks, the primary focus is on the object (e.g., people, cars, bicycles), while other information (e.g., trees, grass, sky) is largely ignored. Studying these types of observations will help to further reduce the amount of information that needs to be transmitted or stored. Therefore, we disclose region-of-interest coding in this application.

[0049] The proposed methods can be used individually or in any combination in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium. In this disclosure, pictures, images, and frames are interchangeable.

[0050] According to the embodiments, two types of region of interest (ROI) methods are discussed, such as spatial ROI and temporal ROI.

[0051] Spatial region of interest coding for image encoding According to one embodiment, on the encoder side, an object detector can be used to detect the bounding boxes of all objects of interest in a frame / image (in this disclosure, images or frames can be used interchangeably), and then the region of interest bounding boxes can be calculated such that the region of interest bounding boxes contain all the bounding boxes of the objects of interest. Figure 4A This is an exemplary illustration of the bounding box of the region of interest determined for frame 400, which includes multiple objects of interest.

[0052] like Figure 4AAs shown, pedestrians and cyclists or motorcyclists are marked with relatively dark bounding boxes. The outer white box is the ROI bounding box, which contains the bounding boxes of all people in the scene.

[0053] According to one embodiment, for a frame, it is assumed that the bounding boxes of all objects of interest are represented as k=0, …, N, where N is the number of objects of interest in the scene, and These are the x and y coordinates of the top-left corner of the bounding box of the k-th object, while These are the x and y coordinates of the bottom right corner of the bounding box of the k-th object. The coordinates of the ROI bounding box can be calculated as follows:

[0054]

[0055]

[0056]

[0057] in, These are the coordinates of the top-left and bottom-right corners of the ROI bounding box. , , and These are the margins used to expand the ROI bounding box, where W and H are the frame width and height, respectively. Depending on the confidence level of the object detector or use case, the encoder can select different margin parameters, for example... = = = = 16.

[0058] Figure 4B The calculation of the ROI bounding box is shown in the figure. Figure 4B As shown, the black boxes represent objects of interest (e.g., 465-1, 465-2), and the dashed box 460 represents the box that exactly contains the objects of interest. The dashed box 455 is the boundary of the ROI with margins added.

[0059] According to one embodiment, after calculating the ROI bounding box (e.g., 455), in order to reduce the bitrate used for transmission or storage, only the content within the ROI bounding box can be encoded to obtain the image bitstream. For example, a first, higher bitrate can be used to encode only the content within the ROI box.

[0060] In the same or another embodiment, frames without ROI bounding boxes can be encoded using a low bitrate configuration (e.g., a high quantization stride or a reduced scale) in case these contents may be of interest in the future. As an example, for encoding only frames without ROI bounding boxes, the ROI regions can be filled with constant values ​​(e.g., 0 or 128, etc.).

[0061] In the same or other embodiments, the entire frame can be encoded using a low bitrate configuration (e.g., a high quantization step size or a reduced scale).

[0062] According to one embodiment, on the decoder side, an image or frame containing an ROI bounding box can be decoded and used independently, or an image or frame containing an ROI bounding box can be placed together with a frame without an ROI portion to form a final image frame. In some embodiments, a padding method can be used to extend the boundaries of an image or frame to the same size as the original image (the area outside the ROI bounding box). In embodiments where it is necessary to restore the original size image, the original size and ROI boundaries can be sent as metadata in the bitstream.

[0063] Therefore, as an exemplary embodiment, the method for spatial ROI disclosed herein may include receiving image data comprising multiple frames from a sensor. Then, for each frame in the image data, multiple bounding boxes associated with multiple objects of interest within the frame may be detected. The frame-level bounding boxes of the frame may be detected based on the coordinates of the multiple bounding boxes. The frame-level bounding boxes may be encoded using a first bitrate. In an embodiment, a second bitrate, less than the first bitrate, may be used to encode portions of the frame not included in the frame-level bounding boxes. In an embodiment, in response to encoding the frame-level bounding boxes using the first bitrate, the second bitrate may be used to encode the frame or at least the remainder of the frame. In some embodiments, a constant value may be used to fill in portions of the frame included in the frame-level bounding boxes before encoding the frame using the second bitrate.

[0064] Spatial region of interest coding for video encoding In one embodiment, encoded frames can be used as a reference to predict the current frame. In such an embodiment, the ROI bounding box information for each frame can be carried as metadata in the bitstream, enabling the decoder to place the decoded ROI content into the correct position in the original image for motion compensation.

[0065] According to one embodiment, a video sequence can be segmented into multiple intra-frame segments. For example, an intra-frame segment may be approximately 1 second or 2 seconds, which could correspond to 32 or 64 frames for a video with a frame rate of 30 frames per second. According to one embodiment, for an intra-frame segment, the ROI bounding boxes for all frames can be calculated. The size of the common ROI bounding box, which includes all individual ROI bounding boxes with specific margins, can be calculated.

[0066] Therefore, according to one embodiment, the method disclosed herein may include determining a common intra-frame time period-level bounding box based on multiple frame-level bounding boxes of the multiple time-sequential frames within an intra-frame time period, for an intra-frame time period comprising multiple time-sequential frames. The method may further include writing the upper-left and lower-right corner coordinates of the common intra-frame time period-level bounding box as metadata into the bitstream. In some embodiments, the intra-frame time period may include a predetermined number of time-sequential frames.

[0067] As mentioned above, some information associated with a frame can be written to the bitstream as metadata. For example, the original size of the frame and the size of the frame-level bounding box can be written to the bitstream as metadata. When writing the size of the frame-level bounding box to the bitstream, the coordinates of the top-left and bottom-right corners of the frame-level bounding box can be written to the bitstream.

[0068] Figure 5 The common ROI bounding boxes used for an intra-frame time period are shown, namely the black boxes (505-1, 505-2, ..., 505-N). Since the common ROI bounding boxes are used within an intra-frame time period, only the information of the common bounding box (i.e., the coordinates of the top-left and bottom-right corners of the common bounding box) is sent in the bitstream as metadata.

[0069] According to one embodiment, for each intra-frame time period or for a specific number of frames, only the content within the common ROI boundary of each frame is transmitted in the bitstream or as metadata. In such an embodiment, since all frames within the intra-frame time period have the same resolution, conventional motion prediction / compensation can be performed. Therefore, in this embodiment, the frames of the image data being analyzed for the ROI can be selected based on a predetermined sampling rate, or can be selected from one of each predetermined number of frames in the image data. In this embodiment, the predetermined number of frames varies based on the complexity of the scene. In this embodiment, some additional information associated with the frames, such as the predetermined sampling rate or the predetermined number of frames, can be written as metadata into the bitstream.

[0070] In the same or another embodiment, raw frames with no content in the ROI region can also be encoded and transmitted at a low bitrate configuration (i.e., large step size or lower scale, etc.). In embodiments, the ROI region can be filled with constant values, such as 0 or 128, etc.

[0071] In the same or another embodiment, the original frame may also be encoded and transmitted at a low bit rate configuration (i.e., large step size or lower scale, etc.).

[0072] According to one embodiment, at the decoder, an image or frame containing the ROI bounding box can be decoded and used independently, or an image or frame containing the ROI bounding box can be combined with frames without ROI portions to form the final image. In this embodiment, a padding method can be used to extend the image's boundaries to the same size as the original image (the area outside the ROI bounding box). In embodiments where video restoration requires the original resolution, the original video resolution can be sent as metadata.

[0073] Region of Interest (ROI) coding for video encoding As is well known, consecutive video frames can share a great deal of common information. For example, consecutive frames may contain the same object with slightly different positions or shapes. This observation can be used in video encoding for machines to reduce transmission rates or storage. Encoders can utilize analysis modules to determine the motion characteristics of video segments.

[0074] In the same or another embodiment, video frames can be temporally downsampled, meaning that only one frame out of every N frames can be selected for encoding and transmission / storage. N can be a fixed value or a variable. For example, N can be smaller, such as 2, for dynamic scenes, or larger, such as 4 or 8, for relatively static content.

[0075] In the same or another embodiment, if an action or event occurs in a specific portion, a portion of the video can be selected for encoding and transmission / storage, rather than uniformly subsampling the video in the temporal domain. Since there is no action or event in the specific portion, that portion can be ignored and not encoded. The remaining content can be sampled uniformly.

[0076] In the same or other embodiments, video frames are sampled non-uniformly. For example, for the first 12 frames, we select frame numbers 0, 3, 7, 8, and 11 for encoding.

[0077] To inform the timing information of the encoded content, metadata such as the sampling rate N can be sent in the bitstream. If a portion of the video uses the same sampling rate N, that sampling rate is only sent at the beginning of the bitstream for that portion. For non-uniformly sampled video, metadata such as frame numbers or timestamps can be sent in the bitstream.

[0078] In cases where it is necessary to recover video frames with the original frame rate, copied images can be inserted into the decoded video. For example, if only frames with frame numbers 0, 3, 7, 8, and 1 are sent, and these frames can be represented as f0, f3, f7, f8, and f11 respectively, the decoder can copy these frames to obtain 12 frames, for example: f0, f0, f0, f3, f3, f3, f7, f8, f8, f8, and f11.

[0079] In an embodiment, to further reduce the bitrate used for transmission or storage, the embodiment may combine spatial ROI methods and temporal ROI methods as disclosed herein in video coding for machines.

[0080] Figure 6 A process 600 for video encoding of machine vision and human-machine hybrid vision according to one embodiment is shown.

[0081] like Figure 6 As shown, image data can be received at operation 605. In this embodiment, image data can be received via a network or from a sensor.

[0082] At operation 610, multiple bounding boxes associated with multiple objects of interest can be determined in a frame of image data.

[0083] In one embodiment, the frame may be selected based on a predetermined sampling rate, or the frame may be selected from one of a predetermined number of frames in the image data. In another embodiment, the predetermined number of frames may vary based on the complexity of the scene.

[0084] At operation 615, frame-level bounding boxes of a frame can be detected based on the coordinates of multiple bounding boxes.

[0085] At operation 620, the first bitrate can be used to encode the frame-level bounding box.

[0086] At operation 625, a second bitrate can be used to encode portions of the frame not included in the frame-level bounding box. In an embodiment, the second bitrate may be less than the first bitrate. In an embodiment, in response to encoding the frame-level bounding box using the first bitrate and / or encoding the frame using the second bitrate, portions of the frame included in the frame-level bounding box are padded with constant values ​​before encoding the frame using the second bitrate.

[0087] In this embodiment, for an intra-frame time period comprising multiple time-sequential frames, a common intra-frame time period-level bounding box can be determined based on multiple frame-level bounding boxes of the multiple time-sequential frames within the intra-frame time period; the upper-left and lower-right corner coordinates of the common intra-frame time period-level bounding box are written as metadata into the bitstream. An intra-frame time period consists of a predetermined number of time-sequential frames.

[0088] Other metadata that can be written to the bitstream may include the original size of the frame, the size of the frame-level bounding box, the coordinates of the top left and bottom right corners of the frame-level bounding box, the predetermined sampling rate, or the predetermined number of frames.

[0089] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or may be obtained from practice of the embodiments.

[0090] It should be understood that the specific order or hierarchy of blocks in the process / flowcharts disclosed herein is illustrative of the method. Based on design preferences, it is understood that the specific order or hierarchy of blocks in the process / flowcharts can be rearranged. Furthermore, some blocks can be combined or others omitted. The appended method claims present the elements of various blocks in a sample order and are not intended to limit one to the specific order or hierarchy presented.

[0091] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of technical detail in the integration. Furthermore, one or more of the aforementioned components may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include one or more computer-readable non-transitory storage media having computer-readable program instructions thereon for causing the processor to perform operations.

[0092] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punched cards or raised structures in recesses on which instructions are recorded, and any suitable combination of the foregoing. The computer-readable storage media used herein should not be construed as transient signals, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through optical fibers), or electrical signals transmitted through wires.

[0093] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or to an external computer or external storage device, via a network (such as the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.

[0094] Computer-readable program code / instructions used to perform operations can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and procedural programming languages ​​such as the "C" programming language or similar programming languages. Computer-readable program instructions can execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet provided by an Internet service provider). In some embodiments, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of computer-readable program instructions to personalize the electronic circuits, thereby performing various aspects or operations.

[0095] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the operations indicated in the flowchart and / or block diagram blocks. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture comprising instructions for implementing aspects of the operations indicated in the flowchart and / or block diagram blocks.

[0096] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operations to be performed on the computer, other programmable apparatus or other device, thereby producing a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, implement the operations indicated in the flowchart and / or block diagram blocks.

[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart may represent a module, fragment, or portion of instructions, including one or more executable instructions for implementing a specified logical operation. The method, computer system, and computer-readable medium may include additional blocks, fewer blocks, different blocks, or blocks in different arrangements. In some alternative implementations, the operations within a block may not occur in the order shown in the figures. For example, in fact, two blocks shown consecutively may be executed simultaneously or substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functionality involved in each block illustrated in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts may be implemented by a system based on dedicated hardware that performs the specified operations or executes a combination of dedicated hardware and computer instructions.

[0098] It is evident that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, this document describes the operation and behavior of the systems and / or methods without reference to any specific software code. It should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

Claims

1. A method for video coding for machine vision and human-machine hybrid vision, characterized in that, The method includes: Receive image data; Detect multiple bounding boxes associated with multiple objects of interest in frames of the image data, wherein the frames of the image data are selected based on a predetermined sampling rate, or the frames of the image data are selected from one of a predetermined number of frames of the image data, wherein the predetermined number of frames varies based on the complexity of the scene; The frame-level bounding box of the frame is detected based on the coordinates of the plurality of bounding boxes, wherein the top-left and bottom-right coordinates of the frame-level bounding box are determined based on the top-left and bottom-right coordinates of the plurality of bounding boxes, as well as the width and height of the frame; and The frame-level bounding box is encoded using a first bitrate, and portions of the frame not included in the frame-level bounding box are encoded using a second bitrate, or the frame is encoded using the second bitrate, wherein the second bitrate is less than the first bitrate. For an intra-frame time period comprising multiple temporally ordered frames, a common intra-frame time period-level bounding box is determined based on multiple frame-level bounding boxes of the multiple temporally ordered frames within the intra-frame time period; and The coordinates of the top left and bottom right corners of the common intra-frame time-level bounding box are written into the bitstream as metadata.

2. The method according to claim 1, characterized in that, The intra-frame time period includes a predetermined number of time-sequential frames.

3. The method according to claim 1, characterized in that, The method further includes: The original size of the frame and the size of the frame-level bounding box are written into the bitstream as metadata.

4. The method according to claim 3, characterized in that, Writing the dimensions of the frame-level bounding box into the bitstream includes writing the coordinates of the top-left corner and the bottom-right corner of the frame-level bounding box into the bitstream.

5. The method according to claim 1, characterized in that, The method further includes: In response to encoding the frame-level bounding box using the first bitrate, the frame is encoded using the second bitrate. Specifically, before encoding the frame using the second bitrate, a portion of the frame included in the frame-level bounding box is filled with a constant value.

6. The method according to claim 1, characterized in that, The method further includes: Write the predetermined sampling rate or the predetermined number of frames as metadata into the bitstream.

7. An apparatus for video encoding of machine vision and human-machine hybrid vision, characterized in that, The device includes: At least one memory configured to store program code; and At least one processor is configured to access the program code and operate in accordance with the instructions of the program code to implement the method of any one of claims 1 to 6.

8. A non-transitory computer-readable medium, characterized in that, It stores computer instructions that, when executed by at least one processor for video encoding of machine vision and human / machine hybrid vision, cause the at least one processor to perform the method of any one of claims 1 to 6.

9. A method for storing a bitstream, characterized in that, The method of performing video encoding according to any one of claims 1 to 6 generates a bitstream and stores the bitstream.

10. A method for transmitting a code stream, characterized in that, The method of performing video encoding according to any one of claims 1 to 6 generates a bitstream and transmits the bitstream.

Citation Information

Patent Citations

  • Method and system for transcoding regions of interests in video surveillance

    US20110051808A1

  • Variable rate shading

    US20190172247A1

  • Patch based video coding for machines

    WO2021211884A1