Video Coding Machine (VCM) Encoders and Decoders for Combined Lossless and Lossy Coding
Patent Information
- Application Number
- JP2023574429
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-08
- Filing Date
- 2022-06-01
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional video encoding methods for applications like surveillance and intelligent transportation systems face challenges in efficiently transmitting and processing large volumes of video data from multiple cameras, requiring significant time for real-time analysis and decision-making.
A video coding machine (VCM) encoder and decoder system that separates video into subpictures, encoding visual signals for human consumption and feature data for machine analysis, using a combination of lossy and lossless encoding techniques to reduce data transmission while maintaining high-quality video and feature extraction.
The system significantly reduces data transmission while ensuring high-quality video and efficient feature extraction for machine analysis, facilitating faster real-time processing and decision-making in applications like surveillance and intelligent transportation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates generally to the field of video encoding and decoding, and more particularly to a video coding machine (VCM) encoder for combined lossless and lossy encoding. [Background technology]
[0002] A video codec may include electronic circuitry or software that compresses or decompresses digital video. The electronic circuitry or software may convert uncompressed video to a compressed format or compressed video to an uncompressed format. In the case of video compression, a device that compresses (and / or performs some function of) the video may typically be called an encoder, and a device that decompresses (and / or performs some function of) the video may be called a decoder.
[0003] The format of the compressed data can conform to standard video compression specifications. The compression can be lossy, in that the compressed video lacks some information that was present in the original video. As a result, the decompressed video can have lower quality than the original uncompressed video, since there is insufficient information to exactly reconstruct the original video.
[0004] There can be a complex relationship between video quality, the amount of data used to represent the video (e.g., as determined by the bit rate), the complexity of the encoding and decoding algorithms, susceptibility to data loss and errors, ease of editing, random access, end-to-end delay (e.g., latency), etc.
[0005] Motion compensation may include techniques for predicting a video frame or a portion thereof given a reference frame, such as a previous frame and / or a future frame, by considering the motion of a camera and / or objects in the video. Motion compensation may be employed in encoding and decoding video data for video compression, for example, encoding and decoding using the Moving Picture Experts Group (MPEG) Advanced Video Coding (AVC) standard (also called H.264). Motion compensation may describe a picture in terms of a transformation of a reference picture to a current picture. A reference picture may be a picture that is temporally earlier than the current picture, or a picture from the future than the current picture. Compression efficiency may be improved if an image can be accurately synthesized from previously transmitted and / or stored images. Summary of the Invention
[0006] A machine-oriented video coding (VCM) encoder is provided, the VCM encoder comprising a feature encoder configured to receive a source video, encode a sub-picture comprising a feature in a source input video, and provide an indication of the sub-picture. The VCM encoder also comprises a video encoder configured to receive a source video, receive an indication of the sub-picture from the feature encoder, and encode the sub-picture. A multiplexer is coupled to the feature encoder and the video encoder, and provides a VCM encoded bitstream having feature data and video data.
[0007] In some embodiments, the video encoder is a lossless encoder, a lossy encoder, or a combination thereof. The video encoder may encode the video according to any applicable encoding standard, such as VVC, AVC, etc.
[0008] The VCM decoder comprises a feature decoder that receives an encoded bitstream having encoded feature data and video data and provides decoded feature data for use by a machine. The VCM decoder also comprises a video decoder that receives the encoded bitstream and feature data from the feature decoder and provides decoded video, such as suitable for human viewing.
[0009] In some embodiments, the VCM decoder is configured to decode video encoded according to an applicable standard, such as VVC, AVC, etc. These and other aspects and features of the non-limiting embodiments of the present invention will become apparent to those of ordinary skill in the art upon consideration of the following description of specific non-limiting embodiments of the present invention in conjunction with the accompanying drawings.
[0010] For the purpose of illustrating the invention, there is shown in the drawings aspects of one or more embodiments of the invention, it being understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown in the drawings. [Brief description of the drawings]
[0011] [Figure 1] FIG. 2 is a block diagram illustrating an example embodiment of a VCC encoder. [Diagram 2] FIG. 2 is a block diagram illustrating an example embodiment of a VCM encoder. [Diagram 3] 1 is a screenshot of an example embodiment of an image having a subpicture containing features. [Figure 4] 1 is a block diagram illustrating an exemplary embodiment of a video decoder. [Diagram 5] 1 is a block diagram illustrating an example embodiment of a video encoder. [Figure 6] FIG. 1 is a block diagram of a computing system that can be used to implement any one or more of the methodologies disclosed herein and one or more portions thereof. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] The drawings are not necessarily to scale and may be shown with phantom lines, diagrammatic representations, and fragmentary views. In certain cases, details that are not necessary for an understanding of the embodiments or that obscure other details may be omitted.
[0013] In many applications, such as surveillance systems with multiple cameras, intelligent transportation, smart city applications, and intelligent industrial applications, traditional video encoding requires the compression of multiple videos from the cameras, transmission over a network to a machine, and consumption by humans. At the machine site, algorithms for feature extraction are then applied, typically using convolutional neural networks or deep learning techniques including object detection, event action recognition, pose estimation, and others. Figure 1 shows a standard VVC encoder applied to a machine.
[0014] A problem with the above approach is the large amount of video transmission from multiple cameras, which may require significant time for efficient and fast real-time analysis and decision making. The machine-directed video coding (VCM) approach embodiments described herein solve this problem by, but are not limited to, both encoding the video and extracting some features at the transmitter site, and then transmitting the resulting encoded bitstream to a VCM decoder. At the VCM decoder site, the video may be decoded for human vision and the features may be decoded for the machine. Referring now to FIG. 2, an exemplary embodiment of a machine-directed video coding (VCM) encoder is shown. The VCM encoder 200 may be implemented using any circuitry, including but not limited to digital and / or analog circuitry; the VCM encoder 200 may be configured using hardware, software, firmware, and / or any combination thereof. The VCM encoder 200 may be implemented as and / or as a component of a computing device, which may include but is not limited to any computing device described below. In one embodiment, the VCM encoder 200 may be configured to receive an input video 204 and generate an output bitstream 208. Receiving the input video 204 may be accomplished in any manner described below. The bitstream may include, but is not limited to, any of the bitstreams described below.
[0015] The VCM encoder 200 may include, but is not limited to, a pre-processor, a video encoder 212, a feature extractor 216, an optimizer, a feature encoder 220, and / or a multiplexer 224. The pre-processor may receive and analyze the input video 204 stream to separate the stream into video, audio, and metadata sub-streams. The pre-processor may include and / or communicate with a decoder, as described in more detail below. In other words, the pre-processor may have the ability to decode the input stream. This may enable decoding of the input video 204, in a non-limiting example, and may facilitate downstream pixel-domain analysis.
[0016] 2, the VCM encoder 200 may operate in a hybrid mode and / or a video mode, where when in the hybrid mode, the VCM encoder 200 may be configured to encode a visual signal for a human consumer and to encode a feature signal for a machine consumer, where the machine consumer may include, but is not limited to, any device and / or component, including, but not limited to, a computing device, as described in more detail below. The input signal may be passed through a pre-processor, for example, when in the hybrid mode.
[0017] Still referring to FIG. 2, the video encoder 212 may include, but is not limited to, any video encoder 212, as described in more detail below. When the VCM encoder 200 is in hybrid mode, the VCM encoder 200 may send the unaltered input video 204 to the video encoder 212, or may send a copy of the same input video 204 and / or an input video 204 that has been altered in some manner to the feature extractor 216. The alterations to the input video 204 may include any scaling, transformation, or other alterations that may occur to one of ordinary skill in the art in view of this disclosure as a whole. By way of example and not limitation, the input video 204 may be resized to a smaller resolution, a certain number of pictures in a sequence of pictures in the input video 204 may be discarded to reduce the frame rate of the input video 204, by way of example and not limitation, color information may be altered by converting RGB, the video may be converted to grayscale video, etc. Still referring to FIG. 2, the video encoder 212 and the feature extractor 216 may be connected and may exchange useful information in both directions. By way of example and not limitation, the video encoder 212 may forward the motion estimation information to the feature extractor 216 or the feature extractor 216 may forward the motion estimation information to the video encoder 212 .
[0018] The video encoder 212 may provide the quantization mapping and / or its data description to the feature extractor 216 based on regions of interest (ROIs) that the video encoder 212 and / or the feature extractor 216 may identify, or the feature extractor 120 may provide the quantization mapping and / or its data description to the video encoder 212 based on regions of interest (ROIs) that the video encoder 212 and / or the feature extractor 216 may identify. The video encoder 212 may provide data to the feature extractor 216 describing one or more partition decisions based on features present and / or identified in the input video 204, the input signal, and / or any frames and / or subframes thereof. The feature extractor 216 may provide data to the video encoder 212 describing one or more partition decisions based on features present and / or identified in the input video 204, the input signal, and / or any frames and / or subframes thereof. The video encoder 212 and the feature extractor 216 may share and / or transmit temporal information with each other for optimal group of pictures (GOP) determination. Each of these techniques and / or processes may be performed as described in further detail below, without limitation.
[0019] 2, the feature extractor 216 may operate in an offline or online mode. The feature extractor 216 may identify features and / or otherwise act on and / or manipulate features. A "feature," as used in this disclosure, is a particular structural and / or content attribute of data. Examples of features may include SIFT, audio features, color histograms, motion histograms, speech levels, loudness levels, etc. A feature may be a timestamp. Each feature may be associated with a single frame of a group of frames. The features may include high-level content features such as timestamps, labels of people and objects in the video, coordinates of objects and / or regions of interest, frame masks for region-based quantization, and / or any other features that may occur to one of skill in the art in view of this disclosure as a whole. As a further non-limiting example, the features may include features describing spatial and / or temporal characteristics of a frame or group of frames. Examples of features describing spatial and / or temporal characteristics may include motion, texture, color, brightness, edge count, blur, blockiness, etc. In the offline mode, all machine models described in more detail below may be stored in the encoder and / or in the encoder's memory and / or may be accessible to the encoder. Examples of such models may include, without limitation, full or partial convolutional neural networks, keypoint extractors, edge detectors, feature map constructors, etc. In the online mode, one or more models may be communicated to the feature extractor 216 by a remote machine in real time or at some point prior to extraction.
[0020] Still referring to FIG. 2, the feature encoder 220 is configured to encode the feature signal generated by the feature extractor 216, for example and not limitation. In one embodiment, after extracting the features, the feature extractor 216 may pass the extracted features to the feature encoder 220, which may generate a feature stream using, for example and not limitation, entropy coding and / or similar techniques, which may be passed to the multiplexer 224. The video encoder 212 and / or the feature encoder 220 may be connected via an optimizer 124, which may exchange useful information between them. For example and not limitation, information related to the codeword structure and / or the length of the entropy coding may be exchanged and reused via the optimizer for optimal compression.
[0021] In one embodiment, and continuing to refer to FIG. 2, the video encoder 212 may generate a video stream, which may be passed to a multiplexer 224. The multiplexer 224 may multiplex the video stream with the feature stream generated by the feature encoder 220, and alternatively or additionally, the video and feature bitstreams may be transmitted over separate channels, separate networks, to separate devices, and / or at separate times or time intervals (time multiplexing). Each of the video and feature streams may be implemented in any manner suitable for implementing any bitstream described in this disclosure. In one embodiment, the multiplexed video and feature streams may generate a hybrid bitstream, which may be transmitted as described in more detail below.
[0022] 2, when VCM encoder 200 is in video mode, VCM encoder 200 may use video encoder 212 for both video and feature encoding. Feature extractor 216 may send the features to video encoder 212, which may encode the features into a video stream, which may be decoded by a corresponding video decoder 232. Note that VCM encoder 200 may use a single video encoder 212 for both video encoding and feature encoding, in which case different sets of parameters may be used for the video and features, or alternatively, VCM encoder 200 may use two separate video encoders 212s, which may operate in parallel.
[0023] 2, system 100 may include and / or be in communication with a VCM decoder 228. VCM decoder 228 and / or elements thereof may be implemented using any circuitry and / or any type of configuration suitable for the configuration of VCM encoder 200 described above. VCM decoder 228 may include, without limitation, a demultiplexer. The demultiplexer may operate to demultiplex the bitstream if multiplexed as described above, for example, without limitation, the demultiplexer may separate a multiplexed bitstream including one or more video bitstreams and one or more feature bitstreams into separate video and feature bitstreams.
[0024] 2, the VCM decoder 228 may include a video decoder 232. The video decoder 232 may be implemented in any manner suitable for a decoder, without limitation, as described in further detail below. In one embodiment, without limitation, the video decoder 232 may generate an output video that may be viewed by a human or other living being and / or device with visual capabilities.
[0025] Still referring to FIG. 2, the VCM decoder 228 may include a feature decoder 236. In one embodiment, but not limited to, the feature decoder 236 may be configured to provide one or more decoded data to a machine. The machine may include any computing device, including but not limited to any microcontroller, processor, embedded system, system on chip, network node, etc., as described below. The machine may manipulate, store, train, receive input from, generate output from, and / or otherwise interact with a machine model, as described in more detail below. The machine may be included in the Internet of Things (IOT), which is defined as a network of objects having processing and communication components, some of which may not be traditional computing devices, such as desktop computers, laptop computers, and / or mobile devices. An object in the IoT may include, but is not limited to, any device incorporating a microprocessor and / or microcontroller and one or more components for interacting with a local area network (LAN) and / or a wide area network (WAN); the one or more components may include, by way of example and not limitation, a wireless transceiver communicating at 2.4-2.485 GHz, such as a Bluetooth transceiver following a protocol promulgated by Bluetooth SIG, Inc., Kirkland, Wash., and / or a network communication component operating according to the MODBUS protocol promulgated by Schneider Electric SE, Rueil-Malmaison, France, and / or the ZIGBEE specification of the IEEE 802.15.4 standard promulgated by the Institute of Electrical and Electronics Engineers (IEEE). Those skilled in the art will recognize, in light of this disclosure as a whole, alternative or additional communication protocols that may be employed consistently with this disclosure and devices supporting such protocols, each of which is contemplated to be within the scope of this disclosure.
[0026] 2, each of the VCM encoder 200 and / or VCM decoder 228 may be designed and / or configured to perform any of the methods, method steps, or sequences of method steps in any of the embodiments described in this disclosure in any order and to any degree of iteration. For example, each of the VCM encoder 200 and / or VCM decoder 228 may be configured to repeatedly perform a single step or sequence of steps until a desired or commanded result is achieved, and the iterations of the step or sequence of steps may be performed iteratively and / or recursively using the output of a previous iteration as input for a subsequent iteration, aggregating the inputs and / or outputs of the iteration to generate an aggregate result, using reduction or decrement of one or more variables, such as global variables, and / or using division of a larger processing task into a set of smaller processing tasks that are addressed repeatedly. Each of the VCM encoder 200 and / or VCM decoder 228 may perform any step or sequence of steps described in this disclosure in parallel, e.g., performing a step two or more times simultaneously and / or nearly simultaneously using two or more parallel threads, processor cores, etc.; the division of tasks among the parallel threads and / or processors may be performed according to any protocol suitable for dividing tasks among iterations. Those skilled in the art, in light of this disclosure as a whole, will recognize various ways in which steps, sequences of steps, processing tasks, and / or data may be subdivided, shared, and / or otherwise accommodate the use of iterations, recursion, and / or parallel processing.
[0027] In some embodiments, still referring to FIG. 2, the amount of data transmitted over the network may be encoded using a combination of lossless and lossy encoding, for example, but not limited to, in a bitstream format, which may be implemented suitable for example, but not limited to, combined lossless and lossy VVC encoding, as described below.
[0028] In one embodiment, and still referring to FIG. 2, when the VCM encoder 200 determines features to be extracted from the source video 204, the encoder may divide the source video 204 into sub-pictures, including but not limited to one or more sub-pictures that include the identified features. The VCM encoder 200 may inform the video encoder 212, which may include but is not limited to a VVC encoder, about the locations of the sub-pictures. The video encoder 212 may then implement a lossy encoding technique, such as but not limited to a simplified shape-adaptive DCT (SA-DCT) algorithm, to encode the identified sub-pictures. A "sub-picture" as described herein may include any portion of a frame and / or combination of such portions, and the portions may include any combination of rectangular forms into blocks, coding units, coding tree units, slices and / or tiles, and / or any shapes with polygonal and / or curved perimeters.
[0029] With further reference to FIG. 2, in one exemplary embodiment, given a rectangular array of pixels, the SA-DCT process performs the N j pixels to the top position and store them as column vector x j and grouping the column vector x j may then be transformed vertically by using a one-dimensional standard DCT, which may result in a corresponding vector with vertical transform coefficients per column. Then, an equivalent procedure may be repeated horizontally, in other words, the same row may be transformed i A column vector a belonging to j M i The elements are shifted to the leftmost positions to form row vector b i which may again be transformed with a one-dimensional standard DCT, but now transformed horizontally to give a row vector c i The one-dimensional standard DCT operation may be performed according to the following equation:
[0030]
number
[0031] Here, the DCT L and S L are respectively, LM i or N j represents the LxL matrix and shape adaptation pre-factor for x, y ...
[0032]
number
[0033] where the star values indicate that quantization has been performed. For a given transform L, the transform matrix DCT L may be given according to the following formula for row index p and column index k:
[0034]
number
[0035] Here, if p=0
[0036]
number
[0037] and 1 otherwise. In one embodiment, the SA-DCT approach may offer a reasonable trade-off between implementation complexity, coding efficiency, and full backward compatibility to existing DCT techniques. The SA-DCT may represent a low-complexity solution with conversion efficiency approaching more complex DCT solutions. Alternatively or additionally, any other DCT-based or other lossy coding protocol may be employed that may occur to one of skill in the art in view of this disclosure, including, but not limited to, other inter-coding, intra-coding, and / or DCT-based approaches.
[0038] 2, in some embodiments, the VCM decoder and / or the video decoder 232 may use a lossless encoding protocol to encode other subpictures and / or one or more video frames displayed in the video format. Alternatively or additionally, the feature encoder 220 may encode the subpictures including the features using a lossless encoding protocol, where the frames are encoded and decoded without loss of information or with negligible loss. A lossless encoding protocol may include, by way of non-limiting example, that the encoder and / or decoder may achieve lossless encoding, bypassing the transform encoding stage and directly encoding the residual. This technique, sometimes referred to in this disclosure as “transform skip residual coding,” may be accomplished by skipping the spatial to frequency domain transform of the residual, as described in further detail below, for example, by applying a transform from the family of discrete cosine transforms (DCTs), as performed in some forms of block-based hybrid video coding. Lossless encoding and decoding may be performed according to one or more alternative processes and / or protocols, including but not limited to those proposed in Core Experiment CE3-1 of JVET-Q00069 on normal and TS residual coding (RRC, TSRC) for lossless coding, and Core Experiment CE3-2 of JVET-Q0080 on modifications of RRC and TSRC for lossless and lossy modes of operation, enabling block differential pulse code modulation (BDPCM) and higher level techniques for lossless coding, and combinations of BDPCM with different RRC / TSRC techniques, etc.
[0039] With further reference to FIG. 2, an encoder described in this disclosure may be configured to encode one or more fields using TS residual coding, where the one or more fields may include, but are not limited to, any picture, sub-picture, coding unit, coding tree unit, tree unit, block, slice, tile, and / or any combination thereof. A decoder described in this disclosure may be configured to decode one or more fields according to and / or using TS residual coding. In transform skip mode, the residual of a field may be coded in non-overlapping sub-blocks or other units of sub-division of a given size, such as, but not limited to, a size of 4 pixels by 4 pixels. Instead of coding the last valid scan position, the quantization index of each scan position in the transformed field may be coded, and the last sub-block and / or sub-division position may be inferred based on the previous sub-division level. The TS residual coding may perform a diagonal scan in a forward manner rather than a backward manner. A forward scan order may be applied to scan sub-blocks within a transform block as well as positions within sub-blocks and / or sub-divisions, and in one embodiment, there may be no signaling of the final (x,y) position. As a non-limiting example, coded_sub_block_flag may be coded for all sub-blocks except the last sub-block when all previous flags are equal to 0. The significance flag context modeling may use a reduced template. The context model of the significance flag may depend on the upper and left neighboring values, and the context model of the abs_level_gt1 flag may also depend on the left and upper significant coefficient flag values.
[0040] As a non-limiting example, during the first scan pass in the TS residual coding process, a significance flag, a sign flag, an absolute level flag greater than 1, and a parity may be coded. For a given scan position, if the significant coefficient is equal to 1, the coefficient sign flag may be coded, followed by a flag specifying whether the absolute level is greater than 1. If abs_level_gtX_flag is equal to 1, par_level_flag may be additionally coded to specify the parity of the absolute level. During the second or subsequent scan pass, for each scan position whose absolute level is greater than 1, up to four abs_level_gtx_flag[i] for i=1...4 may be coded to indicate whether the absolute level at the given position is greater than 3, 5, 7, or 9, respectively. During the third or final "remainder" scan pass, the remainder, which may be stored as the absolute level abs_remainder, may be coded in bypass mode. The absolute level remainder may be binarized using a fixed Rice parameter value of 1.
[0041] Bins in the first and second or "greater than x" scan passes may be context coded until the maximum number of context coded bins in a field, such as, but not limited to, TU, is exhausted. The maximum number of context coded bins in a residual block may be limited to, in a non-limiting example, 1.75*block_width*block_height, or equivalently, 1.75 context coded bins per sample position on average. Bins in the last scan pass, such as the remainder scan pass as described above, may be bypass coded. A variable, such as, but not limited to, RemCcbs, may initially be set to the maximum number of context coded bins in a block or other field and may be reduced by 1 each time a context coded bin is coded. In a non-limiting example, while RemCcbs is 4 or greater, syntax elements in the first coding pass, which may include sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag, may be coded using context coded bins. In some embodiments, if RemCcbs becomes less than 4 during encoding of the first pass, the remaining coefficients not yet encoded in the first pass may be encoded in a remainder scan pass and / or a third pass.
[0042] After the first pass encoding is completed, if RemCcbs is equal to or greater than 4, syntax elements in the second encoding pass, which may include abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag, may be coded using context coded bins. If RemCcbs becomes smaller than 4 during encoding the second pass, the remaining coefficients not yet coded in the second pass may be coded in the residual and / or third scanning pass. In some embodiments, blocks coded using TS residual coding may not be coded using BDPCM coding. For blocks not coded in BDPCM mode, a level mapping mechanism may be applied to transform skip residual coding until the maximum number of context coded bins is reached. The level mapping may use the upper and left neighboring coefficient levels to predict the current coefficient level to reduce the signaling cost. For a given residual position, absCoeff may be denoted as the absolute coefficient level before mapping, and absCoeffMod may be denoted as the coefficient level after mapping. As a non-limiting example, if X0 denotes the absolute coefficient level of the left neighboring position and X1 denotes the absolute coefficient level of the upper neighboring position, the level mapping may be performed as follows:
[0043]
number
[0044] The absCoeffMod value may then be coded as described above. After all context-coded bins have been exhausted, level mapping may be disabled for all remaining scan positions in the current block and / or field and / or sub-division. If coded_subblock_flag is equal to 1, three scan passes as described above may be performed for each sub-block and / or other sub-division, which may indicate the presence of one or more non-zero quantized residuals in the sub-block.
[0045] In some embodiments, when transform skip mode is used for large blocks, the entire block may be used without zeroing any values. In addition, transform skip mode may eliminate transform shift. The statistical characteristics of the signal in TS residual coding may differ from the statistical characteristics of the transform coefficients. Residual coding for transform skip mode may specify a maximum luma and / or chroma block size, and as a non-limiting example, the setting may allow transform skip mode to be used for luma blocks of a size up to MaxTsSize×MaxTsSize, where the value of MaxTsSize may be signaled in the PPS and may have a global maximum possible value such as, but not limited to, 32. When a CU is coded in transform skip mode, its prediction residual may be quantized and coded using a transform skip residual coding process.
[0046] With continued reference to FIG. 2, the encoder described in this disclosure may be configured to encode one or more fields using BDPCM, where the one or more fields may include, but are not limited to, any picture, sub-picture, coding unit, coding tree unit, tree unit, block, slice, tile, and / or any combination thereof. The decoder described in this disclosure may be configured to decode one or more fields according to and / or using BDPCM. BDPCM can maintain full reconstruction at the pixel level. As a non-limiting example, the prediction process for each pixel with BDPCM may include four main steps, which may predict each pixel using its intrablock reference and then reconstruct each pixel to be used as an intrablock reference for subsequent pixels in the remainder of the block: (1) intrablock pixel prediction, (2) residual calculation, (3) residual quantization, and (4) pixel reconstruction.
[0047] Still referring to Figure 2, intrablock pixel prediction may use multiple reference pixels to predict each pixel, which may include, as a non-limiting example, a pixel α to the left of the predicted pixel p, a pixel β above p, and a pixel γ above and to the left of p. The prediction of p may be formulated, without limitation, as follows:
[0048]
number
[0049] Still referring to FIG. 2, once the prediction is calculated, the residual may be calculated. The residual at this stage is lossless and may not be accessible at the decoder side, so
[0050]
number
[0051] and may be calculated as the subtraction of the original pixel value o from the prediction p.
[0052]
number
[0053] 2, pixel-level independence may be achieved by skipping the residual transform and integrating spatial domain quantization, which may be performed by a linear quantizer Q to calculate the quantized residual value r as follows:
[0054]
number
[0055] To accommodate the correct rate-distortion ratio imposed by the quantizer parameter (QP), BDPCM may employ spatial domain normalization, such as, but not limited to, that used in the forward skip mode method, as described above. The quantized residual value r may be transmitted by the encoder.
[0056] Still referring to FIG. 2, another state of BDPCM may include pixel reconstruction using p and r from the previous step, which may be performed, for example, but not limited to, at or by the decoder as follows: c=p+r Once reconstructed, the current pixel may be used as an intrablock reference for other pixels in the same block.
[0057] If there is a relatively large residual when the original pixel value is away from its prediction, the prediction scheme in the BDPCM algorithm may be used. In screen content, this may occur when the current pixel belongs to the background layer while the intrablock reference belongs to the background layer, or when the current pixel belongs to the background layer while the intrablock reference belongs to the background layer. In this situation, which may be called a "layer transition" situation, the available information in the reference may not be enough for accurate prediction. At the sequence level, a BDPCM enable flag may be signaled in the SPS, and this flag may be signaled only if the transform skip mode is enabled in the SPS, for example, but not limited to, as described above. When BDPCM is enabled, if the CU size is less than or equal to MaxTsSize×MaxTsSize in terms of luma samples and the CU is intra-coded, a flag may be sent at the CU level, where MaxTsSize is the maximum block size for which the transform skip mode is allowed. This flag may indicate whether normal intra-coding or BDPCM is used. If BDPCM is used, a BDPCM prediction direction flag may be sent to indicate whether the prediction is horizontal or vertical. The block may then be predicted using a normal horizontal or vertical intra-prediction process with unfiltered reference samples.
[0058] 2, at the decoding site, a feature decoder 236 may assist a video decoder 232, such as but not limited to a VVC decoder, in decoding sub-pictures for human vision, and the decoded features, which in one embodiment may be decoded according to a lossless protocol, may be provided to the video decoder 232 for assembly of the entire video. In one embodiment, the techniques disclosed herein may significantly reduce the amount of data transmitted and still maintain high quality of the decoded video.
[0059] Referring now to FIG. 3, a non-limiting example of the techniques disclosed herein is presented. A VCM encoder 200 may perform face recognition in a video sequence. At the encoder side, a sub-picture 304 may be identified consisting of a person whose face is recognized. The face may be recognized using, but is not limited to, user input, an image classifier such as a neural net classifier, which may include, but is not limited to, a deep neural net classifier, a convolutional neural net classifier, a recurrent neural net classifier, etc., a classifier based on a Naive Bayes classifier, a K-nearest neighbor classifier, and / or a particle swarm optimization, an ant colony optimization, and / or a genetic algorithm classifier. The video with the recognized face may be encoded using, for example, but is not limited to, any combination of lossy and lossless encoding, where, as a non-limiting example, areas with high detail, high importance, etc., such as sub-pictures, may be encoded with lossless encoding and other areas may be encoded with lossy encoding.
[0060] The areas of high importance may include, but are not limited to, faces identified by facial recognition or the like. Alternatively or additionally, the identification of the first region may be performed by receiving semantic information regarding one or more blocks and / or portions of the frame and using the semantic information to identify blocks and / or portions of the frame for inclusion in the first region. The semantic information may include, but is not limited to, data characterizing face detection. The face detection and / or other semantic information may be performed by an automatic face recognition process and / or program and / or by receiving identification of face data, semantic information, etc. from a user. Alternatively or additionally, the semantic importance may be calculated using a significance score.
[0061] Still referring to FIG. 3, the encoder may identify the first region by determining an average measure of information for a number of blocks and using the average measure of information to identify the first region. The identification may include, for example, a comparison of the average measure of information to a threshold. The average measure of information may be determined by calculating a sum of a number of information measures for a number of blocks, which may be multiplied by a significance factor. The significance factor may be determined based on characteristics of the first area. Alternatively, the significance factor may be received from a user. The measure of information may include, for example, a level of detail of the area of the current frame. For example, a smooth area or a highly textured area may contain different amounts of information.
[0062] Further referring to FIG. 3, the average measure of information may be determined, as a non-limiting example, according to the sum of the information measures for each individual block in the area, which may be weighted and / or multiplied by a significance factor, as shown, for example, in the sum:
[0063]
number
[0064] where N is the serial number of the first area, S N is the significant coefficient, k is an index corresponding to one of the blocks that make up the first area, n is the number of blocks that make up the area, and B k is a measure of the information of one block among multiple blocks, A N is the first average measure of information. k may include, for example, a measure of spatial activity calculated using a discrete cosine transform of the block. For example, if the block is a 4×4 block of pixels, then the generalized discrete cosine transform matrix may include a generalized discrete cosine transform II matrix having the following form:
[0065]
number
[0066] where a is 1 / 2 and b is
[0067]
number
[0068] and c is
[0069]
number
[0070] It is. In some implementations, still referring to Figure 3, integer approximations of the transform matrices may be utilized that may be used for efficient hardware and software implementations. For example, if the blocks discussed above are 4x4 blocks of pixels, the generalized discrete cosine transform matrix may include a generalized discrete cosine transform II matrix having the following form:
[0071]
number
[0072] Block B i For , the frequency content of the block may be calculated using: FBi=TxBixT' where T' is the horizontal axis of the cosine transform matrix T, and B i where x is a block represented as a matrix of numbers corresponding to pixels in the block, such as a 4×4 matrix representing a 4×4 block as described above, and the operation x represents matrix multiplication. The measure of spatial activity may alternatively or additionally be performed using edge and / or corner detection, convolution with a kernel for pattern detection, and / or frequency analysis, such as, but not limited to, FFT processing, as described in more detail below.
[0073] Continuing to refer to FIG. 3, if the encoder is further configured to determine a second area within the video frame, as described in more detail below, the encoder may be configured to determine a second average measure of information for the second area, where determining the second average measure of information may be accomplished as described above for determining the first average measure of information.
[0074] Further referring to Figure 3, the significance coefficient S N may be supplied by an external expert and / or calculated based on the characteristics of the area. A "characteristic" of an area, as used herein, is a measurable attribute of an area that is identified based on its content, and the characteristic may be expressed numerically using the output of one or more calculations performed on the first area. The one or more calculations may include any analysis of any signal represented by the first area. One non-limiting example includes the use of a higher S for areas with a smooth background in quality modeling applications. N and assign a lower S to less smooth areas. Nas a non-limiting example, smoothness may be determined using Canny edge detection to determine the number of edges, where a lower number indicates a higher smoothness. Further examples of automatic smoothness detection may include the use of a Fast Fourier Transform (FFT) over signals in spatial variables over an area, where the signals may be analyzed over any two-dimensional coordinate system and over channels representing red-green-blue color values, etc., where a relative dominance of low frequency components in the frequency domain calculated using the FFT may indicate higher smoothness, while a relative dominance of high frequency components may indicate more frequent and rapid transitions in color and / or shade values over the background area, which may lead to a lower smoothness score, and semantically significant objects may be identified by user input. Alternatively or additionally, semantic significance may be detected according to edge configurations and / or texture patterns. The background may be identified by receiving and / or detecting a portion of an area representing significant or "foreground" objects, such as, but not limited to, faces or other items, including semantically significant objects. Another example is that areas containing semantically important objects such as human faces have higher S N may be assigned.
[0075] With further reference to FIG. 3, identifying the first region may include determining a spatial activity measure for each block of the plurality of blocks and using the spatial activity measure to identify the first region. As used in this disclosure, a "spatial activity measure" is a quantity that indicates how frequently and with what amplitude the texture changes within a block, set of blocks, and / or area of a frame. In other words, a flat area such as sky may have a low spatial activity measure, while a complex area such as grass will receive a high spatial activity measure. Determining the respective spatial activity measure may include determining using a transform matrix, such as, but not limited to, a discrete cosine transform matrix. Determining the respective spatial activity measure for each block may include determining using a generalized discrete cosine transform matrix, which may include, but is not limited to, any discrete cosine transform matrix as described above. For example, determining the respective spatial activity measure for each block may include using a generalized discrete cosine transform matrix, a generalized discrete cosine transform II matrix, and / or an integer approximation of a discrete cosine transform matrix.
[0076] In one embodiment, and still referring to FIG. 3, the video encoder 212 may be notified of the sub-picture containing the identified face and / or person, including the video clip size, for example, from frame 700 to frame 756. The video encoder 212 may then apply a lossy encoder to the sub-picture and / or clip, using a simplified SA-DCT. The feature encoder 220 may encode the feature and / or sub-picture containing the feature using lossless encoding, and the feature and / or sub-picture may be decoded by the feature decoder 236 using lossless decoding corresponding to the lossless encoding protocol and combined with the decoded video in the video decoder 232.
[0077] FIG. 4 is a system block diagram illustrating an example decoder 400 capable of adaptive cropping. The decoder 400 may include an entropy decoder processor 404, an inverse quantization and inverse transform processor 408, a deblocking filter 412, a frame buffer 416, a motion compensation processor 420, and / or an intra prediction processor 424. In operation, still referring to FIG. 4, a bitstream 428 may be received by the decoder 400 and input to the entropy decoder processor 404, which may decode portions of the bitstream into quantized coefficients. The quantized coefficients may be provided to the inverse quantization and inverse transform processor 408, which may perform inverse quantization and inverse transform to create a residual signal, which may be added to the output of the motion compensation processor 420 or the intra prediction processor 424 depending on the processing mode. The output of the motion compensation processor 420 and the intra prediction processor 424 may include block predictions based on previously decoded blocks. The sum of the prediction and the residual may be processed by a deblocking filter 412 and stored in a frame buffer 416 .
[0078] In one embodiment, still referring to FIG. 4, the decoder 400 may include circuitry configured to perform any of the operations described above in any embodiment in any order and to any degree of iteration. For example, the decoder 400 may be configured to repeatedly perform a single step or sequence of steps until a desired or commanded result is achieved, where the iterations of the step or sequence of steps may be performed iteratively and / or recursively using the output of a previous iteration as input for a subsequent iteration, aggregating the inputs and / or outputs of the iterations to generate an aggregate result, using reduction or decrement of one or more variables, such as global variables, and / or dividing a larger processing task into a set of smaller processing tasks that are addressed iteratively. The decoder may perform any step or sequence of steps described in this disclosure in parallel, e.g., performing a step two or more times simultaneously and / or nearly simultaneously using two or more parallel threads, processor cores, etc.; the division of tasks among the parallel threads and / or processors may be performed according to any protocol suitable for dividing tasks among the iterations. Those skilled in the art, upon consideration of this disclosure as a whole, will recognize various ways in which steps, sequences of steps, processing tasks, and / or data may be subdivided, shared, and / or otherwise addressed using iteration, recursion, and / or parallel processing.
[0079] 5 is a system block diagram illustrating an example video encoder 500 capable of adaptive cropping. The example video encoder 500 may receive an input video 504 that may be first segmented or divided according to a processing scheme such as a tree-structured macroblock partitioning scheme (e.g., quadtree+bintree). An example of a tree-structured macroblock partitioning scheme may include partitioning a picture frame into large block elements called coding tree units (CTUs). In some implementations, each CTU may be further partitioned into several sub-blocks called coding units (CUs). The end result of this partitioning may include a set of sub-blocks that may be called prediction units (PUs). Transform units (TUs) may be utilized.
[0080] 5, an example video encoder 500 may include an intra prediction processor 508, a motion estimation / compensation processor 512, sometimes referred to as an inter prediction processor, which may construct motion vector candidates including adding global motion vector candidates to a motion vector candidate list, a transform / quantization processor 516, an inverse quantization / inverse transform processor 520, an in-loop filter 524, a decoded picture buffer 528, and / or an entropy coding processor 532. Bitstream parameters may be input to the entropy coding processor 532 and included in the output bitstream 536.
[0081] In operation, and still referring to Figure 5, for each block of a frame of the input video 504, a determination may be made as to whether the block should be processed via intra-picture prediction or using motion estimation / compensation. The block may be provided to an intra-prediction processor 508 or a motion estimation / compensation processor 512. If the block is to be processed via intra-prediction, the intra-prediction processor 508 may perform processing and output a predictor. If the block is to be processed via motion estimation / compensation, the motion estimation / compensation processor 512 may perform processing including building a motion vector candidate list, including adding a global motion vector candidate to the motion vector candidate list, if applicable.
[0082] 5, a residual may be formed by subtracting the predictor from the input video 54. The residual may be received by a transform / quantization processor 516, which may perform a transform operation (e.g., a discrete cosine transform (DCT)) to generate coefficients, which may be quantized. The quantized coefficients and any associated signaling information may be provided to an entropy coding processor 532, where they may be entropy coded and included in the output bitstream 536. The entropy coding processor 532 may support coding of signaling information related to the coding of the current block. In addition, the quantized coefficients may be provided to an inverse quantization / inverse transform processor 520, which may reconstruct the pixels, which may be combined with the predictor and processed by an in-loop filter 524, the output of which may be stored in a decoded picture buffer 528 for use by the motion estimation / compensation processor 512, which may build a motion vector candidate list, including adding global motion vector candidates to the motion vector candidate list.
[0083] 5, a few variations have been described in detail above, but other modifications or additions are possible. For example, in some embodiments, the current block may include any symmetric block (8x8, 16x16, 32x32, 64x64, 128x128, etc.) and any asymmetric block (8x4, 16x8, etc.).
[0084] In some implementations, still referring to FIG. 5, a quadtree plus binary decision tree (QTBT) may be implemented. In QTBT, at the coding tree unit level, the partition parameters of QTBT may be dynamically derived to fit local characteristics without transmitting any overhead. Subsequently, at the coding unit level, a joint classifier decision tree structure may eliminate unnecessary repetitions and control prediction error risk. In some implementations, an LTR frame block update mode may be available as an additional option available at every leaf node of QTBT.
[0085] In some embodiments, still referring to Figure 5, additional syntax elements may be conveyed at different hierarchical levels of the bitstream. For example, a flag may be enabled for the entire sequence by including an enablement flag coded in the sequence-specific parameter set (SPS). Additionally, a CTU flag may be coded at the coding tree unit (CTU) level.
[0086] Some embodiments may include a non-transitory computer program product (i.e., a physically embodied computer program product) having stored thereon instructions that, when executed by one or more data processors of one or more computing systems, cause the one or more data processors to perform the operations herein. Still referring to FIG. 5, the encoder 500 may include circuitry configured to perform any of the operations described above in any embodiment in any order and to any degree of iteration. For example, the encoder 500 may be configured to repeatedly perform a single step or sequence of steps until a desired or commanded result is achieved, and the iterations of the step or sequence of steps may be performed iteratively and / or recursively using the output of a previous iteration as input for a subsequent iteration, aggregating the inputs and / or outputs of the iteration to generate an aggregated result, using reduction or decrement of one or more variables, such as global variables, and / or using division of a larger processing task into a set of smaller processing tasks that are addressed repeatedly. Encoder 500 may perform any step or sequence of steps described in this disclosure in parallel, e.g., performing a step two or more times simultaneously and / or nearly simultaneously using two or more parallel threads, processor cores, etc.; division of tasks among parallel threads and / or processors may be performed according to any protocol suitable for dividing tasks among iterations. Those skilled in the art, in light of this disclosure as a whole, will recognize various ways in which steps, sequences of steps, processing tasks, and / or data may be subdivided, shared, and / or otherwise accommodate the use of iterations, recursion, and / or parallel processing.
[0087] Continuing with reference to FIG. 5, a non-transitory computer program product (i.e., a computer program product that is physically embodied) may store instructions that, when executed by one or more data processors of one or more computing systems, cause the one or more data processors to perform operations and / or steps thereof described in this disclosure, including, without limitation, any operations described above and / or any operations that the decoder 900 and / or the encoder 500 may be configured to perform. Similarly, computer systems are described that may include one or more data processors and a memory coupled to the one or more data processors. The memory may store, either temporarily or permanently, instructions that cause the one or more processors to perform one or more of the operations described herein. In addition, the methods can be implemented by one or more data processors within a single computing system or distributed across two or more computing systems. Such computing systems may be connected and exchange data and / or commands or other instructions, etc., via one or more connections, including connections via a network (e.g., the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, etc.), direct connections between one or more of the computing systems, etc.
[0088] It should be noted that, as would be apparent to one skilled in the computer art, any one or more of the aspects and embodiments described herein may be conveniently implemented using one or more machines (e.g., one or more computing devices utilized as user computing devices for electronic documents, one or more server devices such as document servers, etc.) programmed in accordance with the teachings herein. Appropriate software coding may be readily prepared by skilled programmers based on the teachings of the present disclosure, as would be apparent to one skilled in the software art. The above aspects and embodiments employing software and / or software modules may include hardware suitable to assist in the execution of the machine-executable instructions of the software and / or software modules.
[0089] Such software may be a computer program product employing a machine-readable storage medium. A machine-readable storage medium may be any medium capable of storing and / or encoding sequences of instructions for execution by a machine (e.g., a computing device) to cause the machine to perform any one of the methodologies and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROMs, EEPROMs, and / or any combination thereof. Machine-readable media, as used herein, is intended to include both single media and collections of physically separate media, such as, for example, a compact disc in combination with a computer memory or a collection of one or more hard disk drives. As used herein, machine-readable storage media does not include a transitory form of signal transmission.
[0090] Such software may also include information (e.g., data) carried as a data signal on a data carrier, such as a carrier wave. For example, the machine-executable information may be contained in a data carrier embodied in a data carrier, the signal encoding a sequence of instructions, or portions thereof, executed by a machine (e.g., a computing device) and any associated information (e.g., data structures and data) that cause the machine to perform any one of the methodologies and / or embodiments described herein.
[0091] Examples of computing devices include, but are not limited to, e-book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smart phones, etc.), web appliances, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions that define actions to be taken by the machine, and any combination thereof. In one example, a computing device may include and / or be included in a kiosk.
[0092] 6 shows a diagrammatic representation of one embodiment of a computing device in the exemplary form of a computer system 600 within which a set of instructions may be executed that causes a control system to perform any one or more of the aspects and / or methodologies of the present disclosure. It is also contemplated that multiple computing devices may be utilized to execute a specifically configured set of instructions that causes one or more of the devices to perform any one or more of the aspects and / or methodologies of the present disclosure. The computer system 600 includes a processor 604 and a memory 608 that communicate with each other and other components via a bus 612. The bus 612 may include any of several types of bus structures, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof using any of a variety of bus architectures.
[0093] The processor 604 may include any suitable processor, such as, without limitation, a processor incorporating logic circuits that perform arithmetic and logical operations, such as an arithmetic logic unit (ALU), that may be coordinated with a state machine and directed by operational inputs from memory and / or sensors, and the processor 604 may be organized according to the Von Neumann and / or Harvard architectures, as non-limiting examples. The processor 604 may include, incorporate, and / or be incorporated into, without limitation, a microcontroller, a microprocessor, a digital signal processor (DSP), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a graphical processing unit (GPU), a general purpose GPU, a tensor processing unit (TPU), an analog or mixed signal processor, a trusted platform module (TPM), a floating point unit (FPU), and / or a system on a chip (SoC).
[0094] Memory 608 may include a variety of components (e.g., machine-readable media), including, but not limited to, random access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 616 (BIOS), containing the basic routines that help to transfer information between elements within computer system 600, such as during start-up, may be stored in memory 608. Memory 608 may also include (e.g., stored on one or more machine-readable media) instructions (e.g., software) 620 that embody any one or more of the aspects and / or methodologies of the present disclosure. In another example, memory 608 may further include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0095] Computer system 600 may also include a storage device 624. Examples of storage devices (e.g., storage device 624) include, but are not limited to, hard disk drives, magnetic disk drives, optical disk drives in combination with optical media, solid-state memory devices, and any combination thereof. Storage device 624 may be connected to bus 612 by an appropriate interface (not shown). Examples of interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE®), and any combination thereof. In one example, storage device 624 (or one or more components thereof) may removably interface with computer system 600 (e.g., via an external port connector (not shown)). In particular, storage device 624 and associated machine-readable media 628 may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 600. In one example, the software 620 may reside, completely or partially, within the machine-readable medium 628. In another example, the software 620 may reside, completely or partially, within the processor 604.
[0096] Computer system 600 may also include input devices 632. In one example, a user of computer system 600 may input commands and / or other information to computer system 600 via input devices 632. Examples of input devices 632 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices, joysticks, gamepads, audio input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touchpads, optical scanners, video capture devices (e.g., still cameras, video cameras), touch screens, and any combination thereof. Input devices 632 may interface with bus 612 via any of a variety of interfaces (not shown), including, but not limited to, a serial interface, a parallel interface, a game port, a USB interface, a FIREWIRE interface, a direct interface to bus 612, and any combination thereof. Input devices 632 may include a touch screen interface, which may be part of or separate from display 636, which is described further below. The input device 632 may be utilized as a user selection device for selecting one or more graphical representations within a graphical interface such as those described above.
[0097] A user may also input commands and / or other information to computer system 600 via storage device 624 (e.g., removable disk drive, flash drive, etc.) and / or network interface device 640. A network interface device such as network interface device 640 may be utilized to connect computer system 600 to one or more of a variety of networks, such as network 644, and one or more remote devices 648 connected thereto. Examples of network interface devices include, but are not limited to, a network interface card (e.g., a mobile network interface card, a LAN card), a modem, and any combination thereof. Examples of networks include, but are not limited to, a wide area network (e.g., the Internet, a corporate network), a local area network (e.g., a network associated with an office, a building, a campus, or other relatively small geographic space), a telephone network, a data network associated with a telephone / voice provider (e.g., a data and / or voice network of a mobile communications provider), a direct connection between two computing devices, and any combination thereof. A network such as network 644 may employ wired and / or wireless modes of communication. In general, any network topology may be used. Information (eg, data, software 620 , etc.) may be communicated to and / or from computer system 600 and computer system 1200 via network interface device 640 .
[0098] Computer system 600 may further include a video display adapter 652 for communicating images displayable on a display device, such as display device 636. Examples of display devices include, but are not limited to, a liquid crystal display (LCD), a cathode ray tube (CRT), a plasma display, a light emitting diode (LED) display, and any combination thereof.
[0099] A display adapter 652 and a display device 636 may be utilized in combination with the processor 604 to provide graphical representations of aspects of the disclosure. In addition to a display device, the computer system 600 may include one or more other peripheral output devices, including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices may be connected to the bus 612 via a peripheral interface 656. Examples of peripheral interfaces include, but are not limited to, a serial port, a USB connection, a FIREWIRE connection, a parallel connection, and any combination thereof.
[0100] The above was a detailed description of exemplary embodiments of the present invention. Various modifications and additions can be made without departing from the spirit and scope of the present invention. The features of each of the various embodiments described above can be combined with the features of other embodiments described as appropriate to provide many feature combinations in related new embodiments. Moreover, although the above describes several separate embodiments, what has been described herein is merely illustrative of the application of the principles of the present invention. Furthermore, although certain methods herein may be shown and / or described as being performed in a particular order, the order can be varied considerably within ordinary skill to achieve the methods, systems, and software according to the present disclosure. Thus, this description is intended to be taken as an example only, and is not intended to limit the scope of the present invention.
[0101] Exemplary embodiments have been disclosed above and illustrated in the accompanying drawings. It will be understood by those skilled in the art that various modifications, omissions, and additions may be made to what is specifically disclosed herein without departing from the spirit and scope of the present invention.
Claims
1. A video coding for machines (VCM) encoder, comprising: A feature encoder configured to receive a source video, encode a sub-picture including features in the input video, and provide an index of the sub-picture; A video encoder configured to receive a source video, receive the index of the sub-picture from the feature encoder, and encode the sub-picture using a non-reversible coding protocol; A multiplexer coupled to the feature encoder and the video encoder and configured to provide an encoded bitstream. The VCM encoder comprises the multiplexer.
2. The VCM encoder according to claim 1, further comprising a feature extractor configured to identify the sub-picture.
3. The VCM encoder according to claim 1, wherein the feature encoder is further configured to encode the sub-picture using a reversible coding protocol.
4. The VCM encoder according to claim 3, wherein the reversible coding protocol is a conversion skip residual coding protocol.
5. The VCM encoder according to claim 3, wherein the encoder enables block differential pulse code modulation.
6. The VCM encoder according to claim 1, wherein the feature encoder is further configured to encode the sub-picture using a non-reversible coding protocol.
7. The VCM encoder according to claim 1, wherein the non-reversible coding protocol includes a discrete cosine transform coding protocol.
8. The VCM encoder according to claim 7, wherein the discrete cosine transform coding protocol includes a shape-adaptive discrete cosine transform coding protocol.
9. The VCM encoder according to claim 1, further configured to signal the sub-picture to a decoder.
10. The VCM encoder according to claim 8, wherein signaling the sub-picture further includes signaling a series of frames including the sub-picture.
11. The VCM encoder according to claim 8, wherein signaling the sub-picture further includes signaling the types of features included in the sub-picture.
12. A VCM decoder, comprising: A feature decoder that receives an encoded bitstream having symbolized feature data and video data and provides the feature data decoded for machine consumption. A video decoder that receives the encoded bitstream and the feature data from the feature decoder and provides a decoded video suitable for a human viewer. The VCM decoder comprises the feature decoder and the video decoder. **Claim 13** The VCM decoder according to claim 12, wherein the video decoder is configured to decode an encoded bitstream encoded according to the VVC standard. **Claim 14** The VCM decoder according to claim 12, wherein the video decoder is configured to decode the encoded bitstream encoded using a conversion skip residual encoding protocol. **Claim 15** The VCM decoder according to claim 12, wherein the video decoder is further configured to decode a bitstream encoded using block differential pulse code modulation.