Image and video coding with adaptive quantization for machine-based applications
Through the adaptive quantization module adjusting the quantization matrix according to the machine model, the problem of inefficient encoding of machine consumption images and videos in the prior art is solved, and more efficient encoding and decoding effects are achieved.
Patent Information
- Application Number
- CN202380077174.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-07
- Filing Date
- 2023-09-06
- Publication Date
- 2025-06-13
AI Technical Summary
The existing image and video encoding technology cannot effectively meet the needs of machine consumption, resulting in inefficient encoding and inability to adapt to the sensitivity requirements of different machine tasks.
The adaptive quantization module (AQM) is used to generate frequency importance maps based on the machine model, and the quantization matrix is adjusted to optimize image and video encoding. Combined with classic quantization methods, it can adapt to the needs of different machine tasks.
Improves the efficiency of image and video encoding, ensuring higher accuracy and performance of encoded data in machine tasks.
Smart Images

Figure CN120153652A_ABST
Abstract
Description
Background Art
[0001] The growth trend in video transmission is that a significant portion of the images and videos recorded on-site are only consumed by machines and do not reach the human eye. These machines process images and videos with the goal of completing specific tasks such as object detection, object tracking, segmentation, event detection, etc. Recognizing that this trend is prevalent and will only accelerate in the future, international standardization bodies have been working to standardize image and video coding that is optimized mainly for machine consumption. For example, in addition to the already established standards such as compact descriptors for visual search and compact descriptors for video analysis, standards such as JPEG AI for machines and video coding have been proposed. Video and image data for machine consumption do not necessarily have the same requirements as those for human consumption. As the amount of data for machine consumption grows, solutions are needed to improve the efficiency of the encoded data for machine consumption compared to classical image and video coding techniques.
[0002] In classical image and video coding, a workflow including input preprocessing, partitioning, prediction, frequency transformation, quantization, and entropy coding is mainly used to build a hybrid system. Most steps result in so-called lossless compression, such that the output of the decoder can be the same as the input of the encoder. The only step that allows lossy compression is quantization. Here, the designer of the coding system can apply domain knowledge to remove redundant information from the input signal. In classical image and video coding systems, this is typically done by leveraging knowledge about the human visual system. For example, since the human visual system is more sensitive to low frequencies than high frequencies, quantization for human consumption is usually designed such that more information is retained for low spatial frequencies. However, when humans are replaced by machines as the end users, such quantization strategies may lead to suboptimal results. Depending on the task, the machine may be sensitive to any part of the spectrum. Designing an adaptive strategy in quantization can provide enhanced performance and efficiency. Summary of the Invention
[0003] Systems and methods for image and video coding are proposed, which include an encoder and a decoder. This system and method are preferably applied to use cases where a machine is processing the output of the decoder. This method preferably includes a method for adaptive quantization, which is customized for improved performance in the tasks that the machine performs on the output of the decoder. Compared with widely used systems that apply classical coding techniques developed for human end users, this system and method improve the efficiency of image and video coding.
[0004] In one embodiment, the system includes an encoder and a decoder. The encoder encodes an image and / or video using the proposed quantization method, and the decoder decodes the encoded image and / or video using the information provided by the encoder in the bitstream, or applies an additional method for post-processing quantization independent of the encoded bitstream information.
[0005] An Adaptive Quantization Module (AQM) for encoding and decoding image or video data of a machine using adaptive quantization will preferably obtain a machine model of a machine-based system for receiving the image and / or video data. Based on this machine model, the AQM will generate a frequency importance map. The frequency importance map is used to determine an adjustment matrix, which in turn is used to adjust a default quantization matrix. The AQM uses the adjusted quantization matrix to quantize the image or video data.
[0006] In some embodiments, the frequency importance map is implicitly determined from the machine model by iteratively testing the sensitivity of the model output to changes in each frequency band of interest. This can include using a gradient method based on the differential chain rule to test the frequency sensitivity of the machine model.
[0007] Alternatively, statistics of the sample data set on which the machine model is trained can be used to implicitly determine the frequency importance map.
[0008] The AQM preferably adjusts the default quantization matrix by computing the Hadamard product of the default quantization matrix and the adjustment matrix.
[0009] Preferably, a video encoding system for a machine includes an encoder having an AQM and a compatible decoder having a compatible AQM. In some embodiments, the AQM at the encoder site generates an adjusted quantization matrix and encodes the parameters of the matrix in the bitstream. In this case, the AQM at the decoder can extract the parameters from the bitstream to perform inverse quantization.
[0010] While options for adaptive quantization are available for machine-targeted encoding, the system allows the flexibility of using classical quantization when needed - as an example, for a hybrid use case where both machines and humans are consumers of decoded video, such as decoded video extracted from a bitstream or a substream within a bitstream. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a block diagram of an encoder and decoder system according to the present disclosure;
[0012] Figure 2 is a flowchart showing an adaptive quantization method according to the present disclosure;
[0013] Figure 3A and Figure 3B respectively show matrices of picture block pixels and transform block coefficients according to embodiments of the present disclosure;
[0014] Figure 4A 、 Figure 4B and Figure 4CThe default quantization matrix M, the calculated adjustment factor F, and the resulting adaptive quantization matrix M are respectively shown in an exemplary 4x4 matrix A ;
[0015] Figure 5 is a block diagram showing an exemplary decoder suitable for practicing the present system and method;
[0016] Figure 6 is a block diagram showing an exemplary encoder suitable for practicing the present system and method; and
[0017] Figure 7 is a simplified block diagram of an exemplary general - purpose computing device that can be programmed and configured to execute the present system and method. Detailed Description
[0018] Figure 1 is a simplified block diagram showing a system for encoding and decoding data according to the present disclosure. An input image / video is encoded by an encoder 100 including a pre - processor 105 that converts the input signal into an appropriate modality and format. Examples include color - space conversion from RGB to YCbCr, frame - rate down - sampling, resolution re - scaling, etc. The pre - processor 105 provides an output to a feature extractor 110 that can be used to extract machine - relevant parts of spatial and temporal information. Examples include edge detection, feature - map extraction, key - point extraction, etc. Video encoders 120, 125 are provided, which use hybrid image / video coding with the proposed adaptive quantization method to produce a bit - stream. Optionally, an optimizer 130 is provided and can negotiate redundancy minimization between the encoders 120, 125. For example, if video encoder 120 is encoding edges and other low - level elements of the image / video (as a base layer), then video encoder 125 can be configured to encode the remaining visual information (as an additional layer). In this usage case, on the decoder side, after decoding the base layer, the additional layer is decoded and added to the base layer to regenerate the complete visual information.
[0019] The multiplexer 135 is a multiplexer that combines two or more sub - streams into a unified bit - stream sent to the decoder. Thus, the multiplexer 135 receives the outputs from video encoder 120 and video encoder 125 to generate a synthetic bit - stream from encoder 100.
[0020] The encoder also includes an adaptive quantization module, the aQM module 115, which calculates an adaptive quantization matrix based on a machine model 160 and passes the calculated value to the video encoders 120, 125. The adaptive quantization method employed by the aQM module 15 is further discussed in detail below and depicted in the high - level flowchart of Figure 2 The adaptive quantization method employed by the aQM module 15 is depicted in the high - level flowchart of
[0021] The encoded video bitstream is provided to the decoder 102 via a communication channel. The demultiplexer 140 receives the bitstream and demultiplexes sub-streams from the bitstream and sends them to the appropriate video decoder. The video decoders 150, 155 decode the image / video using the information in the bitstream. Optionally, the video decoder may implement the proposed adaptive quantization method for post-processing independent of the encoder.
[0022] In the decoder 102, an aQM module 145 is provided to calculate an adaptive quantization matrix based on the machine model 160 and pass the calculated value to the video decoders 150, 155 to implement adaptive quantization. Alternatively, the aQM module 145 receives the parameters of the adaptive quantization matrix calculated by the encoder and signaled in the received bitstream.
[0023] In addition to the encoder 100 and the decoder 102, the machine model 160 can be stored or transmitted to the encoder 100 and the decoder 102. The machine model 160 preferably contains relevant information about the machine algorithm. This information can be used to calculate appropriate quantization adjustments.
[0024] Adaptive quantization method
[0025] In classical image / video coding, some variants of the discrete cosine transform, discrete sine transform, wavelet transform or similar transforms are used to transform the pixels representing the input signal or the residuals from the prediction process into frequency coefficients. The reason lies in the energy compressing and decorrelating characteristics of these transforms. In an image / picture pixel block, the energy is almost evenly distributed over the block. After the transform, the data has been horizontally and vertically decorrelated, and one dominant coefficient now contains a significant proportion of the energy.
[0026] For example, in Figure 3A and Figure 3B 8×8 picture pixel blocks and DCT-II transform coefficients are given respectively. Referring to Figure 3A and Figure 3B , the coefficients with significant values are mainly located near the upper left part of the matrix, which corresponds to low frequencies. Most of the values near the lower right are so small that they can be set to zero without significantly affecting the reconstruction quality. This is found to be a near-optimal design in the characteristics of the human visual system that is insensitive to artifacts in the high-frequency regions of the image / picture.
[0027] However, machines typically process visual information based on principles that are significantly different from those of the human visual system. Distortions in the high-frequency portion of the spectrum, which corresponds to the highly textured portions of an image / picture, can lead to a significant degradation in the accuracy of completing a given machine task. The sensitivity of a particular machine to different parts of the input spectrum is not obvious and usually has to be calculated based on machine model parameters.
[0028] Figure 2 Generally describes a method of adaptive quantization processing according to the present invention. This method is used to adjust the default quantization matrix M to the frequency response of a machine model, thereby providing quantization that is both effective in compression and very suitable for transmitting information to a machine-based system that receives the data. At a high level, the method includes the steps of obtaining a machine model (step 200) and using the model to derive a frequency importance map (step 205). The frequency importance map can be used to adjust the M matrix. The machine model and the frequency importance map can be used to derive an F matrix (step 215), and the F matrix is in turn used to calculate the adjusted quantization matrix M A (step 220). Then, the adjusted quantization matrix M is used on the encoder side A to quantize the data (step 225) for transmission through the channel and subsequent decoding at the machine site. The adjusted quantization matrix M A is also preferably signaled in the bitstream for inverse quantization at the decoder.
[0029] Calculation of the adaptive quantization matrix
[0030] The transform coefficient matrix T is quantized such that each value is divided by the corresponding value in the quantization matrix M:
[0031]
[0032] where, represents the Hadamard product of matrices T NxN and M NxN to obtain the quantized coefficient matrix Q NxN . Using the optimal value, the matrix M can achieve optimal rate-distortion performance. However, when the end user is not a human but a machine that performs a specific task for which it is trained, the definition of the distortion metric changes. With this adaptive quantization method, the values in the matrix M are calculated such that the resulting matrix Q achieves the highest performance in task completion with the lowest amount of data. Therefore, preferably M is calculated such that it reflects the importance mapping of the machine model, as Figure 2 shown.
[0033] Reference Figure 2, in step 200, information about the machine model is obtained. Information about the machine model 160 can be obtained explicitly or implicitly. The machine model information can include, for example, the type and size of the objects that the decoder is interested in. For example, if the machine task detects faces and license plates, it may be desirable to preserve details, so higher frequency coefficients (e.g., Figure 2 B and Figure 3A the coefficients in the lower right corner) should not be significantly quantized. On the other hand, if the machine task requires less detail, for example, if it detects the presence of a car, the higher frequency coefficients may be less important. Another example of machine model information is the type of task. For detection tasks, higher frequency coefficients may be important, while for tracking tasks, they are generally less important. Another consideration is the characteristics of the dataset used to train the detection / tracking network. For example, if the training dataset uses JPEG images, it is preferably to match the quantization with the quantization used in the default JPEG settings, etc.
[0034] Explicit information can be provided by the designer who builds and configures the machine model 160. This information has a direct and explicit mapping to quantization, which is trivial. Here, we are considering implicit derivations that can be performed from the training samples or from the model parameters.
[0035] The spectral domain information is represented by the transform coefficients. The magnitude of the coefficients corresponds to the importance of a given frequency. The information most relevant to the machine model is the information retained in the coefficients after quantization. The dynamic range of all the coefficients at a given frequency, given by the energy or the variance of all the samples, represents the importance of a given frequency. Therefore, machine-optimized quantization should retain the frequencies that are more important to the machine model and reduce or remove the less important frequencies.
[0036] In step 205, a frequency importance map can be derived for the machine model. To determine the frequency importance map for a given machine model, two implicit techniques can be used. First, a technique equivalent to backpropagation can be used to obtain the mapping from the model itself. A gradient method based on the differential chain rule can be used to test the sensitivity of the model output to changes in each frequency band. If the output loss function of the machine model is given as L and the input to the model is given as X, the differential chain gives the following relationship:
[0037]
[0038] where each X n is the output of the nth layer of the neural network. Since the input X is obtained by dequantizing the quantized frequency coefficients in Q, the further expansion of the chain is related to the derivative (or gradient in the case of a matrix):
[0039]
[0040] This derivative will determine the sensitivity of the model represented by the loss function L to changes in each quantized coefficient.
[0041] Using this information, the default quantization matrix M can be adjusted such that the elements M corresponding to the elements Q with the highest gradient are scaled down, and vice versa. If the backpropagation method is not available or is not easily computable, an alternative implicit technique is to use the statistics of the sample dataset on which the model is trained. The rationale for using the training set statistics is demonstrated by the fact that the sensitivity of the model is related to the variance of the training samples. In other words, and generally speaking, the model will recognize new inputs within the dynamic range of the training samples, while inputs outside the range of the learning samples will be prone to misclassification. i,j of i,j In step 215, an adjustment factor matrix F is determined. The calculation of the correlation between the training and new samples is preferably performed in the frequency domain. For each frequency coefficient position i,j, the variance quotient can be calculated and stored as an adjustment factor.
[0042] For example,
[0043] For example,
[0044]
[0045] T’ i,j is the coefficient in the transformed new sample, and T i,j is the coefficient in the transformed training sample at the same frequency. T is calculated as the average of all coefficients at a given frequency in the dataset. An additional coefficient k is introduced as a multiplicative factor that can be adjusted based on a specific usage scenario to control the rate, and by default, it has a value of 1.
[0046] The final value of the adaptive quantization matrix M A is obtained by calculating the Hadamard product of the default matrix M and the adjustment factor matrix F.
[0047]
[0048] Examples of the default quantization matrix M, the calculated adjustment factor F, and the resulting adaptive quantization matrix M A are given in Figure 4A , Figure 4B and Figure 4C respectively. It can be seen that the adaptive quantization values have a different distribution from the default quantization values, and the frequency importance direction is less uniform. For the sake of illustration, Figures 4A to 4C the example given in
[0049] Although the given example describes a machine model based on a neural network, similar methods can be applied to other models based on different statistical learning techniques. Each time the machine model is updated, the value of the adaptive quantization matrix can be calculated. For example, this can be done once when initializing the system based on the starting parameters of the machine model and then repeated each time the machine model is updated. The update can be passed to the aQM module, which recalculates the matrix values. Updates to the machine model 160 can be initiated based on retraining the existing model or replacing the model with a new one, or using the same machine framework reused / retrained for a new task, or replacing the entire framework. Adaptive quantization is a general process that does not depend on the specific configuration or dimensions of the machine model.
[0050] Regarding the operation of the aQM module in the decoder, it is important to note the role of its post - processing. The output of the decoder process is typically a pixel representation of the encoded signal. Since the formula for calculating the sensitivity of the model to the input signal can be applied in the pixel domain (X) or extended to the quantization domain (Q), the post - processing can be done directly on the decoder output in the pixel domain, or the decoder output can be transformed to the frequency domain, quantized, and then de - quantized back to the pixel domain. In both cases, when passed into the machine model, the resulting pixel representation is expected to yield better performance in terms of accuracy.
[0051] Bitstream
[0052] The bitstream typically includes standard elements such as a stream header, sub - stream headers, and payload information. The signaling of the adaptive quantization matrix can be signaled in the bitstream, for example, by using a picture header. Parameters can be specified at the lowest block level.
[0053] Each image / picture header can contain adaptive quantization parameters in the following format:
[0054] Block ID AQ Increment 10 [1.2,1.4,2.7,1.8,1.6,2.1,1.5,0.9,2.2,1.3,0.7,0.5,1.1,0.8,0.6,0.3] 11 [0.0,0.2,0.1,1.2,1.2,0.4,0.2,0.0,0.0,0.0,0.2,0.4,-0.2,-0.4,0.0,0.1] … …
[0055] The block ID represents the identification number of the predicted block (i.e., a block in image coding, a macro - block in video coding, or a coding unit) in the current image / picture. The AQ increment represents an adjustment to an element of matrix M. The elements are encoded as the value of the first block and the difference (increment) for each subsequent block. For example, to obtain the element value of the second block, the value of the first block from the first row is added to the value of the difference (increment) from the second row.
[0056] Figure 5 is a system block diagram showing an example of a decoder suitable for use as a video decoder 155 and / or a feature decoder 150. The decoder 500 can include an entropy decoder processor 504, an inverse quantization and inverse transform processor 508, a de - block filter 512, a frame buffer 516, a motion compensation processor 520, and / or an intra - prediction processor 524.
[0057] In operation, and still referring to Figure 5 , the bitstream 528 can be received by the decoder 500 and input to the entropy decoder processor 504, which can entropy decode portions of the bitstream into quantized coefficients. The quantized coefficients can be provided to the inverse quantization and inverse transform processor 508, which can perform inverse quantization and inverse transform to create a residual signal that can be added to the output of the motion compensation processor 520 or the intra prediction processor 524 depending on the processing mode. The outputs of the motion compensation processor 520 and the intra prediction processor 524 can include block predictions based on previously decoded blocks. The sum of the prediction and the residual can be processed by the deblocking filter 512 and stored in the frame buffer 516.
[0058] In an embodiment, and still referring to Figure 5 , the decoder 500 can include circuitry configured to implement any of the operations described above in any of the embodiments described above in any order and with any degree of repetition. For example, the decoder 500 can be configured to repeatedly execute a single step or sequence until a desired or commanded result is achieved; the output of a previous repetition can be used as the input for a subsequent repetition to iteratively and / or recursively execute the repetition of a step or sequence of steps, aggregate the repeated inputs and / or outputs to produce an aggregated result, decrement or reduce one or more variables (such as global variables), and / or divide a larger processing task into a set of smaller processing tasks addressed iteratively. The decoder 500 can execute any step or sequence of steps described in this disclosure in parallel, such as executing a step two or more times simultaneously and / or substantially simultaneously using two or more parallel threads, processor cores, etc.; the division of tasks between the parallel threads and / or processes can be performed according to any protocol suitable for dividing tasks between iterations. Those skilled in the art will know, upon reading the entirety of this disclosure, the various ways in which iterative, recursive, and / or parallel processing can be used to subdivide, share, or otherwise process steps, sequences of steps, processing tasks, and / or data.
[0059] Figure 6 is a system block diagram showing an exemplary embodiment of a video encoder 600 suitable for use as a video encoder and / or feature encoder. The exemplary video encoder 600 can receive an input video 604, which can be initially segmented or partitioned according to a processing scheme such as a macroblock segmentation scheme (e.g., quadtree plus binary tree) in the form of a tree structure. Examples of tree structure macroblock partitioning schemes can include dividing a picture frame into large block elements called coding tree units (CTUs). In some implementations, each CTU can be further segmented one or more times into multiple sub-blocks called coding units (CUs). The result of this segmentation can include groups of sub-blocks that can be referred to as prediction units (PUs). Transform units (TUs) can also be utilized.
[0060] Still referring to Figure 6 Figure 6 , example video encoder 600 may include an intra prediction processor 608, a motion estimation / compensation processor 612 (which may also be referred to as an inter prediction processor) capable of constructing a motion vector candidate list (including adding global motion vector candidates to the motion vector candidate list), a transform / quantization processor 616, an inverse quantization / inverse transform processor 620, a loop filter 624, a decoded picture buffer 628, and / or an entropy coding processor 632. Bitstream parameters may be input to the entropy coding processor 632 to be included in the output bitstream 636.
[0061] In operation, and still referring to Figure 6 Figure 6 , for each block of a frame of the input video, it may be determined whether to process the block via intra picture prediction or using motion estimation / compensation. Based on this determination, the block may be provided to the intra prediction processor 608 or the motion estimation / compensation processor 612. If the block is to be processed via intra prediction, the intra prediction processor 608 may perform processing to output a predictor. If the block is to be processed via motion estimation / compensation, the motion estimation / compensation processor 612 may perform processing including constructing a motion vector candidate list, including adding global motion vector candidates to the motion vector candidate list (if applicable).
[0062] Still further referring to Figure 6 Figure 6 , a residual may be formed by subtracting the prediction from the input video. The residual may be received by the transform / quantization processor 616, which may perform a transform process (e.g., a discrete cosine transform (DCT)) to produce coefficients that may be quantized. The quantized coefficients and any associated signaling information may be provided to the entropy coding processor 632 for entropy coding and included in the output bitstream 836. The entropy coding processor 632 may support the coding of signaling information related to coding the current block. Additionally, the quantized coefficients may be provided to the inverse quantization / inverse transform processor 620, which may regenerate pixels that may be combined with the prediction and processed by the loop filter 624, and the output of the loop filter 624 may be stored in the decoded picture buffer 628 for use by the motion estimation / compensation processor 612 capable of constructing a motion vector candidate list (including adding global motion vector candidates to the motion vector candidate list).
[0063] Continuing to refer to Figure 6 Figure 6 , although several variations have been described in detail above, other modifications or additions are possible. For example, in some embodiments, the current block may include any symmetric block (8x8, 16x16, 32x32, 64x64, 128x128, etc.) and any asymmetric block (8x4, 16x8, etc.).
[0064] In some implementations, and still referring to Figure 6 , a quadtree plus binary decision tree (QTBT) can be implemented. In the QTBT, at the coding tree unit level, the splitting parameters of the QTBT can be dynamically derived to adapt to local characteristics without incurring any overhead in terms of transmission. Subsequently, at the coding unit level, the joint classifier decision tree structure can eliminate unnecessary iterations and control the risk of misprediction. In some implementations, the LTR frame block update mode can be used as an additional option available at each leaf node of the QTBT.
[0065] In some embodiments, and still referring to Figure 6 , additional syntax elements can be signaled at different hierarchical levels of the bitstream. For example, a flag can be enabled for an entire sequence by including an enable flag encoded in the sequence parameter set (SPS). Additionally, CTU flags can be encoded at the coding tree unit (CTU) level.
[0066] Some embodiments may include a non-transitory computer program product (i.e., a physically embodied computer program product) storing instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations herein.
[0067] Still referring to Figure 6 , the encoder 600 can include circuitry configured to implement any of the operations described above in any of the embodiments in any order and with any degree of repetition.
[0068] For example, the encoder 600 can be configured to repeatedly execute a single step or sequence until a desired or commanded result is achieved; the repetition of steps or sequences of steps can be performed iteratively and / or recursively using the output of a previous repetition as the input to a subsequent repetition, aggregating the repeated inputs and / or outputs to produce an aggregated result, decrementing or reducing one or more variables (such as global variables), and / or dividing a larger processing task into a set of smaller processing tasks addressed iteratively. The encoder 600 can perform any of the steps or sequences of steps described in this disclosure in parallel, such as executing a step two or more times simultaneously and / or substantially simultaneously using two or more parallel threads, processor cores, etc.; the division of tasks between parallel threads and / or processes can be performed according to any protocol suitable for dividing tasks between iterations. Those skilled in the art will know, upon reading the entirety of this disclosure, the various ways in which iteration, recursion, and / or parallel processing can be used to subdivide, share, or otherwise process steps, sequences of steps, processing tasks, and / or data.
[0069] Continuing to refer to Figure 6, a non-transitory computer program product (i.e., a physically embodied computer program product) can store instructions that, when executed by one or more data processors of one or more computing systems, cause at least one data processor to perform the operations and / or steps described in this disclosure, including but not limited to any of the operations described above and / or any operations that a decoder and / or encoder may be configured to perform. Similarly, a computer system is also described that may include one or more data processors and a memory coupled to the one or more data processors. The memory may temporarily or permanently store instructions that cause at least one processor to perform one or more operations described herein. Additionally, a method may be implemented by one or more data processors within a single computing system or distributed between two or more computing systems. Such computing systems may be connected and may exchange data and / or commands or other instructions, etc., including connections via a network (e.g., the Internet, wireless wide area network, local area network, wide area network, wired network, etc.), direct connections between one or more of the multiple computing systems, etc.
[0070] It should be noted that any one or more aspects and embodiments described herein can be conveniently implemented using one or more machines (e.g., one or more computing devices used as user computing devices for electronic documents, one or more server devices (e.g., document servers), etc.) programmed according to the teachings of this specification, which is obvious to those of ordinary skill in the computer art. Based on the teachings of this disclosure, a skilled programmer can easily prepare appropriate software code, which is obvious to those of ordinary skill in the software art. The aspects and implementations using software and / or software modules discussed above may also include appropriate hardware for assisting in implementing the machine-executable instructions of the software and / or software modules.
[0071] Such software can be a computer program product employing a machine-readable storage medium. A machine-readable storage medium can be any medium that is capable of storing and / or encoding a sequence of instructions executable by a machine (e.g., a computing device) and causing the machine to perform any of the methods and / or embodiments described herein. Examples of machine-readable storage media include but are not limited to magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROMs, EEPROMs, and any combination thereof. As used herein, a machine-readable medium is intended to include a single medium as well as a collection of physically separate media, such as, for example, a collection of optical disks or one or more hard disk drives combined with computer memory. As used herein, a machine-readable storage medium does not include transient forms of signal transmission.
[0072] Such software may also include information (e.g., data) carried as a data signal on a data carrier (such as a carrier wave). For example, machine-executable information may be included as a data-bearing signal embodied in a data carrier, where the signal encodes a sequence of instructions or portions thereof to be executed by a machine (e.g., a computing device) and any associated information (e.g., data structures and data) that cause the machine to perform any of the methods and / or embodiments described herein.
[0073] Examples of computing devices include, but are not limited to, e-reader devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smart phones, etc.), network devices, network routers, network switches, bridges, any machine capable of executing a sequence of instructions specifying actions to be taken by that machine, and any combination thereof. In one example, a computing device may include and / or be included in an all-in-one machine.
[0074] Figure 7 A graphical representation of one embodiment of a computing device in an exemplary form of a computer system 700 is shown, within which a set of instructions can be executed to cause a control system to perform any one or more aspects and / or methods of the present disclosure. It is also contemplated that multiple computing devices may be utilized to implement a specially configured set of instructions for causing one or more devices to perform any one or more aspects and / or methods of the present disclosure. The computer system 700 includes a processor 704 and a memory 708 that communicate with each other via a bus 712 and with other components. The bus 712 may include any one of several types of bus structures using any of various bus architectures, including but not limited to a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof.
[0075] The processor 704 may include any suitable processor, such as, but not limited to, a processor that includes logic circuitry for performing arithmetic and logical operations, such as an arithmetic and logic unit (ALU), which may be regulated by a state machine and guided by operational inputs from memory and / or sensors; the processor 704 may be organized according to, as a non-limiting example, the von Neumann and / or Harvard architectures. The processor 704 may include, incorporate, and / or be incorporated in but not limited to a microcontroller, a microprocessor, a digital signal processor (DSP), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a graphics processing unit (GPU), a general-purpose GPU, a tensor processing unit (TPU), an analog or mixed-signal processor, a trusted platform module (TPM), a floating-point unit (FPU), and / or a system-on-chip (SoC).
[0076] The memory 708 may include various components (e.g., machine-readable media), including but not limited to random access memory components, read-only components, and any combination thereof. In one example, the basic input / output system 716 (BIOS), including basic routines that help transfer information between elements within the computer system 700 during startup, may be stored in the memory 708. The memory 708 may also include (e.g., stored on one or more machine-readable media) instructions (e.g., software) 720 that embody any one or more aspects and / or methods of the present disclosure. In another example, the memory 708 may also include any number of program modules, including but not limited to an operating system, one or more applications, other program modules, program data, and any combination thereof.
[0077] The computer system 700 may also include a storage device 724. Examples of storage devices (e.g., storage device 724) include but are not limited to hard disk drives, disk drives, optical disc drives in combination with optical media, solid state memory devices, and any combination thereof. The storage device 724 may be connected to the bus 712 through an appropriate interface (not shown). Example interfaces include but are not limited to SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FireWire), and any combination thereof. In one example, the storage device 724 (or one or more of its components) may be removably interfaced with the computer system 700 (e.g., via an external port connector (not shown)). In particular, the storage device 724 and the associated machine-readable medium 728 may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 700. In one example, the software 720 may reside entirely or partially within the machine-readable medium 728. In another example, the software 720 may reside entirely or partially within the processor 704.
[0078] The computer system 700 may also include an input device 732. In one example, a user of the computer system 700 may input commands and / or other information into the computer system 700 via the input device 732. Examples of the input device 732 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices, joysticks, gamepads, audio input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touchpads, optical scanners, video capture devices (e.g., still cameras, video cameras), touchscreens, and any combination thereof. The input device 732 may be connected to the bus 712 via any one of a variety of interfaces (not shown), which include (but are not limited to) serial interfaces, parallel interfaces, game ports, USB interfaces, FIREWIRE interfaces, direct interfaces to the bus 712, and any combination thereof. The input device 732 may include a touchscreen interface, which may be part of or separate from the display 736, and is further discussed below. The input device 732 may be used as a user selection device for selecting one or more graphical representations in the graphical interface as described above.
[0079] The user may also input commands and / or other information into the computer system 700 via the storage device 724 (e.g., removable disk drive, flash drive, etc.) and / or the network interface device 740. A network interface device, such as the network interface device 740, may be used to connect the computer system 700 to one or more of a variety of networks (such as the network 744) and one or more remote devices 748 connected thereto. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, enterprise networks), local area networks (e.g., networks associated with an office, building, campus, or other relatively small geographical space), telephone networks, data networks associated with telephone / voice providers (e.g., mobile communication provider data and / or voice networks), direct connections between two computing devices, and any combination thereof. Networks such as the network 744 may employ wired and / or wireless communication modes. Generally, any network topology may be used. Information (e.g., data, software 720, etc.) may be transmitted to and / or from the computer system 700 via the network interface device 740.
[0080] The computer system 700 may also include a video display adapter 752 for transmitting a displayable image to a display device, such as display device 736. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, and any combination thereof. The display adapter 752 and the display device 736 may be used in conjunction with the processor 704 to provide a graphical representation of aspects of the present disclosure. In addition to the display device, the computer system 700 may include one or more other peripheral output devices, including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices may be connected to the bus 712 via a peripheral interface 756. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, FIREWIRE connections, parallel connections, and any combination thereof.
[0081] The foregoing was a detailed description of illustrative embodiments of the invention. Various modifications and additions may be made without departing from the spirit and scope of the invention. The features of each of the various embodiments described above may be appropriately combined with the features of other described embodiments in order to provide a variety of feature combinations in associated new embodiments. Further, while multiple separate embodiments have been described above, what has been described herein is merely illustrative of the application of the principles of the invention. Additionally, although a particular method herein may be shown and / or described as being performed in a particular order, the order is highly variable within the ordinary skill in the art to achieve the methods, systems, and software in accordance with the present disclosure. Accordingly, the description is intended to be illustrative only and not to limit the scope of the invention in any other way.
[0082] Exemplary embodiments have been disclosed above and shown in the drawings. Those skilled in the art will understand that various changes, omissions, and additions may be made to what has been specifically disclosed herein without departing from the spirit and scope of the invention.
Claims
1. A video encoder for encoding video data for a machine using adaptive quantization, the encoder including a quantization processor having a default quantization matrix and performing steps including the following: Obtain a machine model of a machine-based system for receiving the video data; Generate a frequency importance map from the machine model; Determine an adjustment matrix based on the frequency importance map; Use the adjustment matrix to adjust the default quantization matrix; And Quantize the video data using the adjusted quantization matrix.
2. The encoder according to claim 1, Wherein, The frequency importance map is implicitly determined from the machine model by iteratively testing the sensitivity of the model output to changes in each band of interest.
3. The encoder according to claim 2, Wherein, Testing the frequency sensitivity of the machine model further includes: using a gradient method based on the differential chain rule.
4. The encoder according to claim 1, Wherein, The frequency importance map is implicitly determined using the statistics of the sample data set on which the machine model is trained.
5. The encoder according to claim 1, Wherein, Adjusting the default quantization matrix includes: calculating the Hadamard product of the default matrix and the adjustment matrix.
6. An adaptive quantization module for encoding or decoding video data, the adaptive quantization module having a processor programmed to perform an adaptive quantization method, the adaptive quantization method Includes: Obtain a machine model of a machine-based system for receiving the video data; Generate a frequency importance map from the machine model; Determine an adjustment matrix based on the frequency importance map; Use the adjustment matrix to adjust the default quantization matrix; And Quantize the video data using the adjusted quantization matrix.
7. The adaptive quantization module according to claim 6, Wherein, The frequency importance map is implicitly determined from the machine model by iteratively testing the sensitivity of the model output to changes in each band of interest.
8. The adaptive quantization module according to claim 7, Wherein, Testing the frequency sensitivity of the machine model further includes: using a gradient method based on the differential chain rule.
9. The adaptive quantization module according to claim 6, Wherein, The frequency importance map is implicitly determined using the statistics of the sample data set on which the machine model is trained.
10. The adaptive quantization module according to claim 6, Wherein, Adjusting the default quantization matrix includes: calculating the Hadamard product of the default matrix and the adjustment matrix.
11. A decoder for decoding a video bitstream for consumption by a machine, the decoder having an adaptive quantization module programmed to perform inverse quantization of a bitstream encoded with an adaptive quantization method, the inverse quantization Includes: Obtain a machine model of a machine-based system for receiving video data; Generate a frequency importance map from the machine model; Determine an adjustment matrix based on the frequency importance map; Use the adjustment matrix to adjust the default quantization matrix; And Inverse quantize the video data using the adjusted quantization matrix.
12. The decoder according to claim 11, wherein, the frequency importance map is implicitly determined from the machine model by iteratively testing the sensitivity of the model output to changes in each frequency band of interest.
13. The decoder according to claim 12, wherein, testing the frequency sensitivity of the machine model further includes: using a gradient method based on the differential chain rule.
14. The decoder according to claim 11, wherein, the frequency importance map is implicitly determined using the statistics of the sample data set on which the machine model is trained.
15. The decoder according to claim 11, wherein, adjusting the default quantization matrix includes: calculating the Hadamard product of the default matrix and the adjustment matrix.