Method and apparatus for encoding video for machine vision
Patent Information
- Application Number
- CN202280016592.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-22
- Filing Date
- 2022-09-28
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-09-28
Smart Images

Figure CN116982313B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 277,517, filed November 9, 2021, and U.S. Patent Application No. 17 / 950,564, filed September 22, 2022, pursuant to 35 U.S. SC § 119, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure relates to video coding for machine vision. Specifically, video coding methods for machine vision and human / machine hybrid vision are disclosed. Background Technology
[0004] Traditionally, videos or images are used for various purposes, such as entertainment and education. Therefore, video coding or image coding often leverages the characteristics of the human visual system to achieve better compression efficiency while maintaining good subjective quality.
[0005] In recent years, with the rise of machine learning applications and the proliferation of sensors, many intelligent platforms have utilized video for machine vision tasks, such as object detection and segmentation. How to encode video or images for machine tasks has become an interesting and challenging problem, leading to the introduction of research into Video Coding (VCM). To achieve this goal, the international standards group MPEG created an ad hoc group, "VCM," to standardize related technologies, thereby enabling better interoperability between different devices.
[0006] Existing video codecs are primarily designed for human use. However, an increasing number of videos are being used by machines for machine vision tasks, such as object detection and instance segmentation. Therefore, it is crucial to develop an efficient video codec for machine vision or human / machine hybrid vision applications. Summary of the Invention
[0007] The following presents a simplified overview of one or more embodiments of this disclosure to provide a basic understanding of these embodiments. This overview is not a comprehensive summary of all contemplated embodiments and is neither intended to identify key or critical elements of all embodiments nor to depict any or all scope of embodiments. Its sole purpose is to present some concepts of one or more embodiments of this disclosure in a simplified form as a prelude to the more detailed description that follows.
[0008] Methods, apparatus, and nonvolatile computer-readable media for encoding video for machine vision and human / machine hybrid vision.
[0009] According to an exemplary embodiment, a method for encoding video for machine vision and human / machine hybrid vision is performed by one or more processors. The method includes: receiving input at a hybrid codec, the input including at least one of video or image data, the hybrid codec including a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision. The method further includes: compressing the input using the first codec, wherein the compression includes downsampling the input using a downsampling module and upsampling the compressed input using an upsampling module to generate a residual signal. The method further includes: quantizing the residual signal to obtain a quantized representation of the input. The method further includes: entropy encoding the quantized representation of the input using one or more convolutional filter modules; and training one or more networks using the entropy-encoded quantized representation.
[0010] According to an exemplary embodiment, an apparatus for encoding video for machine vision and human / machine hybrid vision includes: at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate in accordance with the instructions of the computer program code to perform the method described in the embodiment.
[0011] According to an exemplary embodiment, an apparatus for encoding video for machine vision and human / machine hybrid vision includes: a setup module configured to receive input at a hybrid codec, the input including at least one of video or image data, the hybrid codec including a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision. The apparatus further includes: a compression module configured to compress the input using the first codec, wherein the compression module includes a downsampling unit configured to downsample the input using the downsampling module, and the compression module includes an upsampling unit configured to upsample the compressed input using the upsampling module to generate a residual signal. The apparatus further includes: a quantization module configured to quantize the residual signal to obtain a quantized representation of the input. The apparatus further includes: an entropy encoding module configured to entropy encode the quantized representation of the input using one or more convolutional filter modules. The apparatus further includes: a training module configured to train one or more networks using the entropy-encoded quantized representation.
[0012] According to an exemplary embodiment, a non-volatile computer-readable medium storing computer instructions, when executed by at least one processor, causes the at least one processor to perform a method for encoding video for machine vision and human / machine hybrid vision. The method includes: receiving input at a hybrid codec, the input including at least one of video or image data, the hybrid codec including a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision. The method further includes: compressing the input using the first codec, wherein the compression includes downsampling the input using a downsampling module and upsampling the compressed input using an upsampling module to generate a residual signal. The method further includes: quantizing the residual signal to obtain a quantized representation of the input. The method further includes: entropy encoding the quantized representation of the input using one or more convolutional filter modules. The method further includes: training one or more networks using the entropy-encoded quantized representation.
[0013] Additional embodiments will be set forth in the description which follows, and will be apparent in part from the description, and / or may be learned by practicing the embodiments presented in this disclosure. Attached Figure Description
[0014] The above and other features and aspects of embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings, wherein:
[0015] Figure 1 This is a schematic diagram of an example network device according to various embodiments of the present disclosure.
[0016] Figure 2 An architecture of a hybrid video codec disclosed according to an embodiment of this disclosure is shown.
[0017] Figure 3 This is video encoding for machine systems according to various embodiments of the present disclosure.
[0018] Figure 4 This is a schematic diagram of the architecture of a learning-based image codec according to various embodiments of the present disclosure.
[0019] Figure 5 This is a schematic diagram of the architecture of a learning-based image codec according to various embodiments of the present disclosure.
[0020] Figure 6 This is a flowchart illustrating an example process for training one or more networks for a hybrid video codec according to various embodiments of this disclosure.
[0021] Figure 7 These are examples of learning-based video codecs according to various embodiments of this disclosure. Detailed Implementation
[0022] The following detailed description of the exemplary embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0023] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations can be made based on the foregoing disclosure, or modifications and variations can be obtained from the practice of the embodiments. Furthermore, one or more features or components of some embodiments may be incorporated into or combined with some embodiments (or one or more features of some embodiments). Moreover, it is understood that in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be switched.
[0024] It is evident that the systems and / or methods described herein can be implemented in various forms, including hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation. Therefore, this document describes the operation and behavior of the systems and / or methods without referring to any specific software code—it should be understood that software and hardware can be designed to implement the system and / or method based on the description herein.
[0025] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically recited in the claims and / or disclosed in the specification. While each dependent claim listed below may directly depend on only one claim, the disclosure of possible embodiments includes combinations of each dependent claim in the claim set with each of the other claims.
[0026] Unless explicitly stated otherwise, elements, actions, or instructions used herein should not be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Where only one item is referred to, the term “one” or similar language is used. Furthermore, as used herein, the terms “has,” “have,” “having,” “include,” “including,” etc., are intended to be open-ended terms. Additionally, the phrase “based on” is intended to mean “at least partially based on,” unless explicitly stated otherwise. Furthermore, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.
[0027] References to "some embodiments," "embodiments," or similar language in this specification mean that a particular feature, structure, or characteristic described in connection with the illustrated embodiments is included in some embodiments of this solution. Therefore, the phrases "in some embodiments," "in embodiments," and similar language in this specification may, but do not necessarily, refer to the same embodiments.
[0028] Furthermore, the features, advantages, and characteristics described herein can be combined in one or more embodiments in any suitable manner. Based on the description herein, those skilled in the art will recognize that this disclosure can be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of this disclosure may be recognized in certain embodiments.
[0029] The disclosed methods can be used individually or in combination in any order. Furthermore, each method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-volatile computer-readable medium.
[0030] Embodiments of this disclosure relate to video coding for machines. Specifically, video coding methods for machine vision and human / machine hybrid vision are disclosed. Conventional video codecs are designed for human use. In some embodiments, conventional video codecs can be combined with learning-based codecs to form hybrid codecs, enabling efficient encoding of video for machine vision and human-machine hybrid vision.
[0031] Figure 1This is a schematic diagram of an example device used to perform translation services. Device 100 can correspond to any type of known computer, server, or data processing device. For example, device 100 may include a processor, a personal computer (PC), a printed circuit board (PCB) including a computing device, a minicomputer, a mainframe computer, a microcomputer, a telephone computing device, a wired / wireless computing device (e.g., a smartphone, a personal digital assistant (PDA)), a laptop, a tablet computer, a smart device, or any other similar operating device.
[0032] In some embodiments, such as Figure 1 As shown, device 100 may include a set of components, such as processor 120, memory 130, storage component 140, input component 150, output component 160 and communication interface 170.
[0033] Bus 110 may include one or more components that allow communication between a set of components of device 100. For example, bus 110 may be a communication bus, a cross-over bar, a network, etc. Although bus 110 is... Figure 1 While depicted as a single line, bus 110 can be implemented using multiple (two or more) connections between a set of components of device 100. This disclosure is not limited thereto.
[0034] Device 100 may include one or more processors, such as processor 120. Processor 120 may be implemented as hardware, firmware, and / or a combination of hardware and software. For example, processor 120 may include a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), general-purpose single-chip or multi-chip processor, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the operations described herein. A general-purpose processor may be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. Processor 120 may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be performed by circuitry specific to a given operation.
[0035] The processor 120 can control the overall operation of the device 100 and / or a group of components of the device 100 (e.g., memory 130, storage component 140, input component 150, output component 160, and communication interface 170).
[0036] Device 100 may further include memory 130. In some embodiments, memory 130 may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic memory, optical memory, and / or another type of dynamic or static storage device. Memory 130 may store information and / or instructions for use (e.g., execution) by processor 120.
[0037] The storage component 140 of device 100 may store information and / or computer-readable instructions and / or code relating to the operation and use of device 100. For example, storage component 140 may include hard disks (e.g., magnetic disks, optical disks, magneto-optical disks, and / or solid-state drives), optical disks (CDs), digital versatile disks (DVDs), Universal Serial Bus (USB) flash drives, PCMCIA cards, floppy disks, cassette tapes, magnetic tapes, and / or another type of non-volatile computer-readable media and corresponding drives.
[0038] Device 100 may further include an input component 150. Input component 150 may include one or more components that allow device 100 to receive information, for example, via user input (e.g., a touchscreen, keyboard, keypad, mouse, stylus, button, switch, microphone, camera, etc.). Alternatively or additionally, input component 150 may include sensors for sensing information (e.g., a Global Positioning System (GPS) component, accelerometer, gyroscope, actuator, etc.).
[0039] The output component 160 of device 100 may include one or more components that can provide output information of device 100 (e.g., display, liquid crystal display (LCD), light-emitting diode (LED), organic light-emitting diode (OLED), haptic feedback device, speaker, etc.).
[0040] Device 100 may further include a communication interface 170. Communication interface 170 may include a receiver component, a transmitter component, and / or a transceiver component. Communication interface 170 enables device 100 to establish connections and / or transmit communications with other devices (e.g., a server, another device). Communication can be implemented via wired connections, wireless connections, or a combination of wired and wireless connections. Communication interface 170 may allow device 100 to receive information from and / or provide information to another device. In some embodiments, communication interface 170 may provide communication with another device via a network, such as a local area network (LAN), wide area network (WAN), metropolitan area network (MAN), private network, ad hoc network, intranet, Internet, fiber-based network, cellular network (e.g., fifth-generation (5G) network, long-term evolution (LTE) network, third-generation (3G) network, code division multiple access (CDMA) network, etc.), public land mobile network (PLMN), telephone network (e.g., public switched telephone network (PSTN)), and / or combinations of these networks or other types of networks. Alternatively or additionally, communication interface 170 may provide communication with another device via a device-to-device (D2D) communication link, such as FlashLinQ, WiMedia, Bluetooth, ZigBee, Wi-Fi, LTE, 5G, etc. In other embodiments, communication interface 170 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, etc.
[0041] Device 100 may be included in core network 240 and perform one or more of the processes described herein. Device 100 may perform operations based on processor 120 executing computer-readable instructions and / or code stored in non-volatile computer-readable media (e.g., memory 130 and / or storage component 140). Computer-readable media may refer to non-volatile memory devices. Memory devices may include memory space within a single physical storage device and / or memory space distributed across multiple physical storage devices.
[0042] Computer-readable instructions and / or code can be read from another computer-readable medium or from another device into memory 130 and / or storage component 140 via communication interface 170. The computer-readable instructions and / or code stored in memory 130 and / or storage component 140, if or when executed by processor 120, can cause device 100 to perform one or more of the processes described herein.
[0043] Alternatively or additionally, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more of the processes described herein. Therefore, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.
[0044] supply Figure 1 The number and arrangement of components shown are for illustrative purposes. In practice, with... Figure 1 Compared to the components shown, there can be more components, fewer components, different components, or different arrangements of components. Furthermore, Figure 1 The two or more components shown can be implemented within a single component, or Figure 1 The single component shown can be implemented as multiple distributed components. Additionally or alternatively, Figure 1 The set (one or more) components shown can perform actions described by Figure 1 The other set of components shown performs one or more operations.
[0045] Figure 2 This is a block diagram of an embodiment of a hybrid video codec 200. The hybrid video codec 200 may include a conventional codec 220 and a learning-based codec 230. The input 201 of the hybrid codec can be video or an image, as an image can be considered a special type of video (e.g., a video with one image). Figure 2 In this context, a conventional video codec 220 can be used to compress input video 201 at different scales (e.g., original resolution or downsampled). The downsampling rate of the downsampling module 210 can be fixed and known in both the encoder 221 and decoder 223, or the downsampling rate can be user-defined, such as 100% (e.g., no downsampling), 50%, 25%, etc., and sent as metadata in the bitstream 224 to inform the decoder 222. The conventional video codec can be VVC, HEVC, H.264, or image codecs such as JPEG, JPEG2000. The downsampling module 210 can be a classic image downsampler or a learning-based image downsampler. An upsampling module 250 can be used to upsample the decoded downsampled video 203 (e.g., ...). Figure 2 The upsampling module 250 can be an upsampled version of the video (e.g., "low-resolution video 203") to its original resolution (e.g., "high-resolution video 204"), which can be used for human vision. The upsampling module 250 can be a classic image upsampler or a learning-based image upsampler, such as a learning-based super-resolution module.
[0046] In some embodiments, the hybrid video codec 200 may also employ a learning-based video codec 230 to compress the downsampled video 203.
[0047] In the encoder, a low-resolution video 203 can also be generated and upsampled to the original input resolution. The upsampled reconstructed video 205 can then be subtracted from the input video to generate a residual video signal 202, which can be fed into... Figure 2In the learning-based codec 230, the upsampling module 240 of the hybrid video codec 200 can be the same as the upsampling module 250 (after the low resolution). Video 203 can be decoded at the decoder. The output of the residual decoder 238 can be added on top of the high-resolution video 204 to form a reconstructed video 205 that can be used for machine vision tasks.
[0048] This will be discussed in further detail below. Figure 4 and Figure 5 The network in the example can be trained using images. The publicly available networks, such as... Figure 6 The network in the middle needs to use Figure 2 The residual signal 202 shown is used for retraining.
[0049] During training, the input to the learning-based codec 230 can be the residual signal / image 202, which is also the ground truth. The loss function can utilize rate-distortion loss as follows:
[0050] L 总 =R+λ mse L mse (1)
[0051] In Equation 1, R represents the bitstream cost, which can be the estimated bits per pixel (BPP) value, and L... mse It is the mean square error between residual image 202 and the corresponding reconstructed residual image 205, such as Figure 2 As shown. λ mse It is a positive weighting factor used to balance bit rate cost and compression performance.
[0052] In some embodiments, the rate-distortion loss can be modified as follows:
[0053] L 总 =R+λ ms-ssim L ms-ssim (2)
[0054] Where L ms-ssim The MS-SSIM metric, λ, is calculated using residual image 202 and the corresponding reconstructed residual image 205. ms-ssim It is a weighting factor.
[0055] In some embodiments, task networks can be used in the joint training process. Machine vision task networks, such as the object detection network YOLOv3 or Faster R-CNN, can be added for joint training. Figure 2 The high-resolution video 204 can be an upsampled decoded video. L 检测 This is the detection loss calculated in the object detection network. The total loss function can be expressed as Equation 3 or 4:
[0056] L 总 =R+λ mse L mse +λ 检测 L 检测 (3)
[0057] L 总 =R+λ ms-ssim L ms-ssim +λ 检测 L 检测 (4)
[0058] Where λ 检测 It is a positive weighting factor. During training, the model parameters of the machine vision network can be fixed, and only the parameters of the residual encoder 231 / decoder 238, entropy encoder 233 / decoder 236, and entropy model 235 are trained. In some embodiments, the parameters of the machine vision network can be... Figure 2 The rest of the network is trained together.
[0059] Figure 3 An embodiment of the architecture of a video coding machine (VCM), such as a hybrid video codec 200, is illustrated. Sensor output 300 travels along video encoding path 311 through VCM encoder 310 to VCM decoder 320, where it undergoes video decoding 321. Another path leads to feature extraction 312, feature transformation 313, feature encoding 314, and feature decoding 322. The output of VCM decoder 320 is primarily used for machine consumption, i.e., machine vision 306. In some cases, it may also be used for human vision 305. One or more machine tasks are then performed to understand the video content.
[0060] Figure 4 and Figure 5 The architecture of a learning-based image codec 230 in some embodiments is shown. The learning-based codec 230 can be an image codec that allows for frame-by-frame compression of the residual video signal described above without considering temporal redundancy between frames. For example, the learning-based image codec 230 can follow the following... Figure 4 and Figure 5 The autoencoder architecture is shown. In some instances, since quantization operations 401A and 401B are rounding operations (e.g., rounding floating-point numbers to their adjacent integers), the dequantization module can be removed. The corresponding dequantization module 237 can be an identity module and can be removed from the architecture. Therefore, it can be removed based on the operation of quantization module 232. Figure 2 The dequantizer module 237 in the middle.
[0061] Figure 4 and Figure 5The two architectures are similar autoencoder architectures. Figure 4 An example architecture is shown, including an analysis network 410, a synthesis network 420, quantization networks 401A and 401B, an arithmetic encoder 402A and 402B, a decoder 403A and 403B, and an entropy model 430. The differences lie in the details of the analysis network 410, the synthesis network 420, and the entropy model 430.
[0062] Furthermore, the network can be designed for compression of regular images. Therefore, the number of filters in each convolutional module can be larger. For example, in Figure 2 In this case, N = 128 and M = 192 or 320. Similarly, in Figure 3 In the middle, N = 192. Because in Figure 2 The learning-based image codec 230 can be used to compress the residual signal 202, thus significantly reducing the number of filters in the convolutional filter module to lower complexity with minimal performance degradation. Figure 5 In this context, the Generalized Split Normalization (GDN) module 511 and the Inverse Generalized Split Normalization (IGDN) module 521 can be replaced by rectified linear units (ReLU 531). To reduce complexity, these can be removed. Figure 4 The attention module in the Gaussian mixture entropy model can be used. Figure 2 The scale prior module shown is used instead.
[0063] In this disclosure, Figure 4 and Figure 5 The architecture in [the document] can serve as an example, and can be utilized by following [the principles of]... Figure 4 or Figure 5 Any autoencoder of the spirit or a simplified version thereof. For example, Figure 2 An example of this architecture is shown in the figure. In some embodiments, the learning-based codec 230 may be a video codec, thereby enabling the utilization of temporal redundancy between frames in the residual signal 202. Figure 7 An example of a learning-based video codec 230 is shown in the figure.
[0064] Figure 6A flowchart illustrating an embodiment of a process for training one or more networks is shown. The process may begin at operation S610, whereby input comprising at least one video or image is received at a hybrid codec. The hybrid codec may include a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision. For example, the hybrid codec may be a hybrid video codec 200 including a first codec and a second codec. The first codec may be a conventional codec 220, and the second codec may be a learning-based codec 230 receiving input 201, such as... Figure 2 As shown. The process proceeds to operation S620, where the input is compressed using a first codec. For example, input 201 can be compressed by the first codec 220. Compression may include downsampling the input using a downsampling module 210 and upsampling the compressed input using an upsampling module 240 that generates the residual signal 202. The process proceeds to operation S630, where the residual signal (e.g., residual signal 202) is quantized using a quantizer (e.g., quantizer 232) to obtain a quantized representation of the input. The process proceeds to operation S640, where the quantized representation of the input is entropy-encoded using one or more convolutional filter modules of an entropy model (e.g., entropy model 235). The process proceeds to operation S650, where one or more networks are trained using the entropy-encoded quantized representation of the input and a reconstructed video (e.g., reconstructed video 205).
[0065] To specify the hybrid video codec 200, several parameters need to be specified using high-level syntax, such as the sequence parameter set, picture parameter set, and picture header. Alternatively, this information can be passed through system-level metadata or using SEI messages.
[0066] Downsampling rate: A set of downsampling rates can be defined, for example {r0, r1, ..., r N-1} where N is the number of downsampling rates. Assume the height and width of the input image resolution are W and H, and the downsampling rate is r. n The height and width of the downsampled image are r and r, respectively. n W and r n H. Specifies the index n∈{0,1,…N-1} to indicate the sampling rate r used in the hybrid video codec. p The decoder needs to use an upsampling rate. The decoded low-resolution image / video is upsampled to obtain a high-resolution image / video. Index n can be binarized using fixed-length code or p-order exponent Golomb code and transmitted in the bitstream via bypass coding. In some embodiments, p = 0 or 1.
[0067] Upsampling Module (Upsampler): If only one type of upsampler is used in the hybrid video codec, it is not necessary to specify the upsampler in the bitstream. However, if multiple types of upsamplers can be used in the codec, each with different complexity and performance, information about the upsampler type needs to be specified in the bitstream. For example, if there are M types of upsamplers, the index m ∈ {0, 1, ..., M-1} is used to specify which type of upsampler to use. The index m can be binarized using fixed-length code or p-order exponent Golomb code and transmitted in the bitstream via bypass coding. In some embodiments, p = 0 or 1.
[0068] Codecs used for encoding downsampled images / video: If only one type of codec is used in the hybrid video codec used for encoding downsampled images / video, it is not necessary to specify the codec in the bitstream. However, if multiple types of codecs can be used in the system, such as VVC, HEVC, H.264, etc., information about the codec type needs to be specified in the bitstream to allow appropriate decoding in the decoder. For example, if there are Q types of upsamplers, the index q∈{0,1,…Q-1} is used to specify which type of codec to use. The index q can be binarized using fixed-length code or p-order exponent Golomb code and transmitted in the bitstream via bypass coding. In some embodiments, p = 0 or 1.
[0069] Codecs used for encoding residual images / videos: Similarly, if only one type of learning-based codec is used in a hybrid video codec for encoding residual images / videos, it is not necessary to specify the codec in the bitstream. However, if multiple types of codecs can be used in the system, for example... Figure 2-4 The codecs shown require specifying information about their type in the bitstream to allow for appropriate decoding in the decoder. For example, if there are L types of codecs, the index l ∈ {0, 1, ..., L-1} can be used to specify which type of codec to use. Index l can be binarized using fixed-length code or p-order exponent Golomb code and transmitted in the bitstream via bypass coding. In some embodiments, p = 0 or 1. Both the encoder and decoder should be aware of the different types of codecs.
[0070] In some embodiments, instead of using an index, a description of the network structure of the learning-based codec can be specified in the high-level syntax or metadata. For example, we can specify individual modules in the decoder network, such as a 3×3 convolutional module with N output filters, a 3×3 convolutional module with N output filters and 2× upsampling, etc. In addition to the network structure, corresponding decoder model parameters can be sent in the bitstream to allow the decoder to decode the bitstream and generate the reconstructed residual image / video. The decoder network must be symmetric to the encoder network. Partial reconstruction of the residual image / video is possible. Furthermore, instead of a floating-point implementation of the learning-based codec, a fixed-point implementation can be specified.
[0071] Parameter set selection for different machine vision tasks: In some embodiments, different parameter sets can be trained for different machine vision tasks for the residual video signal coding branch, while different tasks share the same network architecture. To perform residual coding on a specific video input, if the framework is to support more than one task, the target machine task should be specified. For example, if there are T types of machine vision tasks, the index t∈{0,1,…T-1} can be used to specify which type of machine task is being targeted. Therefore, the residual signal coding branch will switch to the corresponding parameter set trained for that task.
[0072] Quantization: Quantization typically refers to dividing continuously varying data into a finite number of levels and assigning a specific value to each level. The most basic form of quantization is uniform quantization. Uniform quantization is a method that uses quantization intervals of the same size within a certain range. For example, one method is to set the quantization interval size by dividing the minimum and maximum values of a specific input data by the number of quantization bits desired.
[0073] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations are possible based on the foregoing disclosure, or may be obtained from practice of the embodiments.
[0074] It is understood that the specific order or hierarchy of boxes in the process / flowcharts disclosed herein is illustrative of the exemplary methods. It is understood that the specific order or hierarchy of boxes in the process / flowcharts may be rearranged based on design preferences. Furthermore, some boxes may be combined or omitted. The appended method claims present the elements of the individual boxes in a sample order and are not intended to limit one to the specific order or hierarchy presented.
[0075] Some embodiments may relate to systems, methods, and / or computer-readable media at any possible level of integration technical detail. Furthermore, one or more of the above components may be implemented as instructions stored on a computer-readable medium and executable by at least one processor (and / or may include at least one processor). The computer-readable medium may include a computer-readable non-volatile storage medium (or medium) having computer-readable program instructions thereon for causing a processor to perform operations.
[0076] Computer-readable storage media can be tangible devices that can hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, but are not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital universal disc (DVD), memory sticks, floppy disks, mechanical encoding devices (e.g., punched cards or raised structures in recesses on which instructions are recorded), and any suitable combination of the foregoing. The computer-readable storage media used herein should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0077] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or downloaded to an external computer or external storage device. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the corresponding computing / processing device.
[0078] Computer-readable program code / instructions used to perform operations can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and procedural programming languages (such as the "C" programming language or similar programming languages). Computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network (including local area network (LAN) or wide area network (WAN)) or can be connected to an external computer (e.g., via the Internet provided by an Internet service provider). In some embodiments, electronic circuitry (including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs)) can be personalized by executing computer-readable program instructions using status information from the computer-readable program instructions to perform aspects or operations.
[0079] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to generate machine instructions that, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the operations specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the operations specified in one or more boxes of the flowchart and / or block diagram.
[0080] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operations to be performed on the computer, other programmable apparatus or other device, thereby producing a computer-implemented process, such that the instructions that execute on the computer, other programmable apparatus or other device implement the operations specified in one or more boxes of a flowchart and / or block diagram.
[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing one or more specified logical operations. The method, computer system, and computer-readable medium may include more, fewer, different, or differently arranged blocks than shown in the drawings. In some alternative implementations, the operations in the blocks may not occur in the order shown in the drawings. For example, two blocks shown consecutively may actually be executed simultaneously or substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functionality involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified operations or executes a combination of dedicated hardware and computer instructions.
[0082] Clearly, the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods limits these implementations. Therefore, the operation and behavior of the systems and / or methods are described herein without reference to specific software code—it should be understood that software and hardware can be designed to implement the system and / or method based on the description herein.
[0083] The above disclosure also includes the following listed embodiments:
[0084] (1) A method for encoding video for machine vision and human / machine hybrid vision, the method being executed by one or more processors, the method comprising: receiving an input at a hybrid codec, the input comprising at least one of video or image data, the hybrid codec comprising a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision; compressing the input using the first codec, wherein the compression comprises downsampling the input using a downsampling module and upsampling the compressed input using an upsampling module to generate a residual signal; quantizing the residual signal to obtain a quantized representation of the input; entropy encoding the quantized representation of the input using one or more convolutional filter modules; and training one or more networks using the entropy-encoded quantized representation.
[0085] (2) The method according to feature (1), wherein the conventional codec includes any one of VVC, HEVC, H264, JPEG or JPEG2000 codecs.
[0086] (3) The method according to feature (1) or (2), wherein the learning-based codec includes an image codec and compresses the residual signal frame by frame without considering temporal redundancy.
[0087] (4) The method according to any one of features (1)-(3), wherein the downsampling module is one of the classic image downsampling and learning-based image downsampling.
[0088] (5) The method according to any one of features (1)-(4), wherein the downsampling module uses a downsampling rate N, which is fixed and known in both the encoder and the decoder, or the downsampling rate N is user-defined.
[0089] (6) The method according to any one of features (1)-(5), wherein the upsampling module is one of the classic image upsampling and the learning-based image upsampling.
[0090] (7) The method according to any one of features (1)-(6), wherein the upsampled compressed input is subtracted from the input to generate a second residual signal, and the second residual signal is provided to the learning-based codec.
[0091] (8) The method according to any one of features (1)-(7), wherein the output of the second codec is added to the upsampled compressed input to form a reconstructed video suitable for machine vision tasks.
[0092] (9) The method according to any one of features (1)-(8), wherein the input of the hybrid codec is a truth value.
[0093] (10) The method according to any one of features (1)-(9), wherein the machine vision network is fixed and the parameters of the residual encoder, entropy encoder and entropy model of the second codec are trained.
[0094] (11) An apparatus for encoding video for machine vision and human / machine hybrid vision, comprising: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and operate according to instructions of the computer program code, the computer program code including: setting code configured to cause the at least one processor to receive input at a hybrid codec, the input including at least one of video or image data, the hybrid codec including a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision; compression code configured to cause the at least one A processor compresses the input using the first codec, wherein the compression code includes downsampling code configured to cause the at least one processor to downsample the input using a downsampling module, and the compression code includes upsampling code configured to cause the at least one processor to upsample the compressed input using an upsampling module to generate a residual signal; quantization code configured to cause the at least one processor to quantize the residual signal to obtain a quantized representation of the input; entropy encoding code configured to cause the at least one processor to entropy encode the quantized representation of the input using one or more convolutional filter modules; and training code configured to cause the at least one processor to train one or more networks using the entropy-encoded quantized representation.
[0095] (12) The apparatus according to feature (11), wherein the conventional codec includes any one of VVC, HEVC, H264, JPEG or JPEG2000 codecs.
[0096] (13) The apparatus according to feature (11) or (12), wherein the learning-based codec includes an image codec and compresses the residual signal frame by frame without taking into account temporal redundancy.
[0097] (14) The apparatus according to any one of features (11)-(13), wherein the downsampling module is one of a classical image downsampling device and a learning-based image downsampling device.
[0098] (15) The apparatus according to any one of features (11)-(14), wherein the downsampling module uses a downsampling rate N, which is fixed and known in both the encoder and the decoder, or the downsampling rate N is user-defined.
[0099] (16) The apparatus according to any one of features (11)-(15), wherein the upsampling module is one of a classic image upsampling device and a learning-based image upsampling device.
[0100] (17) The apparatus according to any one of features (11)-(16), wherein an upsampled compressed input is subtracted from the input to generate a second residual signal, and the second residual signal is provided to the learning-based codec.
[0101] (18) The apparatus according to feature (17), wherein the output of the second codec is added to the upsampled compressed input to form a reconstructed video suitable for machine vision tasks.
[0102] (19) The apparatus according to any one of features (11)-(18), wherein the input of the hybrid codec is a true value.
[0103] (20) A non-volatile computer-readable medium storing computer instructions that, when executed by at least one processor, cause the at least one processor to perform a method for encoding video for machine vision and human / machine hybrid vision, the method comprising: receiving an input at a hybrid codec, the input comprising at least one of video or image data, the hybrid codec comprising a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision; compressing the input using the first codec, wherein the compression comprises downsampling the input using a downsampling module and upsampling the compressed input using an upsampling module to generate a residual signal; quantizing the residual signal to obtain a quantized representation of the input; entropy encoding the quantized representation of the input using one or more convolutional filter modules; and training one or more networks using the entropy-encoded quantized representation.
Claims
1. A method for encoding video for machine vision and human / machine hybrid vision, characterized in that, The method includes: Input is received at a hybrid codec, the input including at least one of video or image data, the hybrid codec including a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision. The input is compressed using the first codec, wherein the compression includes downsampling the input using a downsampling module and upsampling the compressed input using an upsampling module to generate a residual signal; The residual signal is quantized to obtain a quantized representation of the input; The quantized representation of the input is entropy-encoded using one or more convolutional filter modules; and Use entropy-encoded quantized representations to train one or more networks; The upsampled compressed input is subtracted from the input to generate a second residual signal, and the second residual signal is provided to the second codec; the output of the second codec is added to the upsampled compressed input to form a reconstructed video suitable for machine vision tasks. The step of training one or more networks using entropy-encoded quantized representations includes: determining an index value that identifies the machine vision task for which the training is being performed.
2. The method according to claim 1, characterized in that, The conventional codecs include any one of VVC, HEVC, H264, JPEG, or JPEG2000 codecs.
3. The method according to claim 1, characterized in that, The learning-based codec includes an image codec and compresses the residual signal frame by frame without considering temporal redundancy.
4. The method according to claim 1, characterized in that, The downsampling module is one of the classic image downsampling and learning-based image downsampling modules.
5. The method according to claim 4, characterized in that, The downsampling module uses a downsampling rate N, which is fixed and known in both the encoder and decoder, or the downsampling rate N is user-defined.
6. The method according to claim 1, characterized in that, The upsampling module is one of the classic image upsampling and learning-based image upsampling modules.
7. The method according to claim 1, characterized in that, The input to the hybrid codec is a true value.
8. The method according to claim 1, characterized in that, The machine vision network is fixed, and the parameters of the residual encoder, entropy encoder, and entropy model of the second codec are trained.
9. An apparatus for encoding video for machine vision and human / machine hybrid vision, characterized in that, The device includes: At least one memory is configured to store computer program code; At least one processor is configured to access the computer program code and operate in accordance with the instructions of the computer program code to perform the method of any one of claims 1-8.
10. An apparatus for encoding video for machine vision and human / machine hybrid vision, characterized in that, The device includes: The setup module is configured to receive input at a hybrid codec, the input including at least one of video or image data, the hybrid codec including a first codec and a second codec, wherein the first codec is a conventional codec designed for human consumption, and the second codec is a learning-based codec designed for machine vision. A compression module is configured to compress the input using the first codec, wherein the compression module includes a downsampling unit configured to downsample the input using the downsampling module, and the compression module includes an upsampling unit configured to upsample the compressed input using the upsampling module to generate a residual signal; A quantization module is configured to quantize the residual signal to obtain a quantized representation of the input; An entropy coding module is configured to entropy code the quantized representation of the input using one or more convolutional filter modules; and The training module is configured to train one or more networks using entropy-encoded quantized representations; The upsampled compressed input is subtracted from the input to generate a second residual signal, and the second residual signal is provided to the second codec; the output of the second codec is added to the upsampled compressed input to form a reconstructed video suitable for machine vision tasks. The training module uses an entropy-encoded quantized representation to train one or more networks, including: determining an index value that identifies the machine vision task for which the training is being performed.
11. The apparatus according to claim 10, characterized in that, The conventional codecs include any one of VVC, HEVC, H264, JPEG, or JPEG2000 codecs.
12. The apparatus according to claim 10, characterized in that, The learning-based codec includes an image codec and compresses the residual signal frame by frame without considering temporal redundancy.
13. The apparatus according to claim 10, characterized in that, The downsampling module is one of the classic image downsampling and learning-based image downsampling modules.
14. The apparatus according to claim 13, characterized in that, The downsampling module uses a downsampling rate N, which is fixed and known in both the encoder and decoder, or the downsampling rate N is user-defined.
15. The apparatus according to claim 10, characterized in that, The upsampling module is one of the classic image upsampling and learning-based image upsampling modules.
16. A non-volatile computer-readable medium having stored thereon computer instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of claims 1-8.
Citation Information
Patent Citations
Compression of Images Having Overlapping Fields of View Using Machine-Learned Models
US20200304835A1
Cascaded Prediction-Transform Approach for Mixed Machine-Human Targeted Video Coding
US20210218997A1
Feature-Domain Residual for Video Coding for Machines
US20210314573A1