Machine visual perception-oriented convertible video feature coding method
By extracting middle-level features, motion compensation, detail reconstruction and feature space transformation, the video feature encoding method for machine vision perception solves the problems of redundancy and waste of computing resources in the prior art, and realizes efficient video encoding and highly adaptable scalability.
Patent Information
- Application Number
- CN202510282291.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
AI Technical Summary
Existing video encoding methods for machine vision perception fail to make full use of prior information, resulting in redundancy in compressed structures and waste of computing resources, while requiring frequent retraining and deployment of codecs and downstream task networks, limiting their scalability in practical applications.
A transformable video feature encoding method for machine vision perception is proposed. The middle-level features are extracted from the original frame sequence through a feature extractor, and the video feature codec is used for motion compensation and detail reconstruction. Then, the reconstructed features are transformed through the feature space transformation module to meet the needs of different downstream tasks.
This method effectively eliminates time and space redundancy, realizes targeted bit rate allocation, reduces the consumption of computing resources, and improves the adaptability and scalability of the coding scheme.
Smart Images

Figure CN120151544A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video coding technology, and in particular, to a transformable video feature coding method for machine vision perception. Background Art
[0002] In recent years, neural video compression frameworks for the human visual system have made great progress and provided excellent video compression performance. However, the comprehensive exploration of video coding for machine vision (VCM) is still in its infancy.
[0003] Video codecs for the human visual system, such as H.265 / HEVC and H.264 / AVC, are often used to compress videos for downstream analysis. However, these methods encounter two limitations in video coding schemes for machine vision perception. First, these compression frameworks focus on minimizing the relevant distortions (PSNR, MS-SSIM) in the pixel domain and the human visual system, rather than meeting the specific needs of machine vision applications, which is not the best way to solve machine vision perception tasks. Second, machine vision tasks usually only require a subset of the image content. For example, transmitting unnecessary background information in the task of image classification will result in an increase in bitrate overhead, and more customized methods are needed in machine-centric scenarios.
[0004] Some studies have deeply explored the Analyze-Then-Compress (ATC) paradigm to solve the above problems. This paradigm starts with extracting features from images, followed by feature compression for specific downstream tasks. For video feature compression, there are two strategies. One strategy requires optimizing the video feature codec through a specific downstream loss. The other strategy focuses on freezing the codec while fine-tuning the entire downstream task network.
[0005] In summary, the current machine-oriented video feature coding methods have the following disadvantages: 1. They do not make full use of prior knowledge, resulting in redundant compression structures and time consumption; 2. They need to retrain and redeploy the upstream codec or the downstream machine vision network to adapt to new downstream tasks, thus consuming more computing resources and limiting their scalability in practical applications. Summary of the Invention
[0006] The purpose of this application is to overcome the deficiencies of the prior art and provide a transformable video feature coding method for machine vision perception.
[0007] To achieve the above purpose, this application has taken the following technical solutions.
[0008] In a first aspect, the present application provides a transformable video feature encoding method for machine vision perception, including:
[0009] Extract mid-level features from the original frame sequence through a feature extractor;
[0010] Perform motion compensation on the extracted mid-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and mid-level features to obtain reconstructed features;
[0011] Perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features into at least one downstream task to obtain prediction results.
[0012] In some embodiments, the video feature codec includes a feature buffer, a motion estimation module, a motion codec, a motion compensation module, and a conditional codec, and the feature buffer stores reference features;
[0013] Performing motion compensation on the mid-level features through a video feature codec, and performing detail reconstruction on the obtained compensated features and mid-level features to obtain reconstructed features, includes:
[0014] Estimate the motion information between the reference features and the mid-level features through the motion estimation module, encode and decode the motion information through the motion codec, and perform motion compensation on the encoded and decoded motion information based on the reference features through the motion compensation module to obtain compensated features;
[0015] Based on the encoding perception condition and the decoding perception condition, and reconstruct the detail information between the compensated features and the mid-level features through the conditional codec to obtain reconstructed features.
[0016] In some embodiments, the video feature codec further includes a channel reduction module and an entropy model, and the conditional codec includes a conditional encoder and a conditional decoder;
[0017] Before estimating the motion information between the reference features and the mid-level features through the motion estimation module, the method further includes:
[0018] Perform channel reduction on the reference features and the mid-level features respectively through the channel reduction module to obtain the reduced reference features and the reduced mid-level features;
[0019] Estimating the motion information between the reference features and the mid-level features through the motion estimation module, includes:
[0020] Estimate the motion information between the reduced reference features and the reduced mid-level features through the motion estimation module;
[0021] Based on the encoding perception condition and the decoding perception condition, and by means of a conditional codec, the detailed information between the reconstructed compensation feature and the middle-level feature is reconstructed to obtain a reconstructed feature, including:
[0022] Based on the encoding perception condition, and by means of a conditional encoder, the compensation feature and the reduced middle-level feature are encoded to obtain an encoded feature;
[0023] The compensation feature is information-compressed through an entropy model to obtain a compressed compensation feature;
[0024] Based on the decoding perception condition, and by means of a conditional decoder, the compensation feature, the encoded feature, and the compressed compensation feature are decoded to obtain a reconstructed feature.
[0025] In some embodiments, the video feature codec further includes a feature pyramid network;
[0026] The encoding perception condition is obtained by extracting from the middle-level feature through the feature pyramid network;
[0027] The decoding perception condition is obtained by extracting from the compensation feature through the feature pyramid network.
[0028] In some embodiments, the feature space transformation module includes a first branch, a second branch, and a third branch. The first branch includes a downsampling module and an upsampling module connected in sequence. The second branch includes three bottleneck residual blocks connected in sequence. The third branch includes an upsampling module and a downsampling module connected in sequence;
[0029] Feature transformation is performed on the reconstructed feature through the feature space transformation module, including:
[0030] The information of the reconstructed feature is retained through the first branch to obtain a first transformed feature;
[0031] The shape of the reconstructed feature is migrated through the second branch to obtain a second transformed feature;
[0032] The global information of the reconstructed feature is extracted through the third branch to obtain a third transformed feature;
[0033] The first transformed feature, the second transformed feature, and the third transformed feature are concatenated to obtain a final transformed feature.
[0034] In some embodiments, the downstream tasks include object detection, semantic segmentation, and instance segmentation.
[0035] In a second aspect, the present application further provides a transformable video feature encoding device for machine vision perception, including:
[0036] A feature extraction module, configured to extract middle-level features from an original frame sequence through a feature extractor;
[0037] A compensation and reconstruction module, configured to perform motion compensation on the extracted middle-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and the middle-level features to obtain reconstructed features;
[0038] A transformation and prediction module, configured to perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features into at least one downstream task to obtain prediction results.
[0039] In a third aspect, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method is implemented.
[0040] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented.
[0041] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented.
[0042] Advantages of the present application: The transformable video feature encoding method for machine vision perception provided by the present application extracts middle-level features from the original frame sequence through a feature extractor. The middle-level features retain the original spatial structure and filter out information irrelevant to the downstream tasks, making them easier to compress. Motion compensation and reconstruction are implemented on the video feature codec, which ensures that there are only minor differences between the original frame sequence and the predicted frame sequence, realizes targeted bitrate allocation, maximally eliminates temporal redundancy through motion compensation, and effectively removes spatial redundancy. The reconstructed features are transferred to other feature spaces through the feature space transformation module to adapt to different downstream tasks, thus solving the problem that it requires extremely high computing resources to retrain and redeploy the upstream feature codec or the entire downstream task network in practical applications. In summary, the present application provides a highly adaptable and scalable video coding solution for machine vision perception.
[0043] Additional aspects and advantages of the present application will be given in part in the following description, which will become apparent from the following description, or can be learned through the practice of the present application. Description of the Drawings
[0044] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 One of the flow diagrams of a transformable video feature encoding method for machine vision perception provided by an embodiment of the present application;
[0046] Figure 2 Another flow diagram of the transformable video feature encoding method for machine vision perception provided by an embodiment of the present application;
[0047] Figure 3 The structural diagram of a video feature codec provided by an embodiment of the present application;
[0048] Figure 4 The structural diagram of a feature space transformation module provided by an embodiment of the present application;
[0049] Figure 5 (a), Figure 5 (b), and Figure 5 (c) are respectively the schematic diagrams of performance comparison for object detection tasks, semantic segmentation tasks, and instance segmentation tasks;
[0050] Figure 6 The schematic diagram of instance segmentation prediction result comparison provided by an embodiment of the present application;
[0051] Figure 7 The visualization schematic diagram of compensation features, motion representations, and motion patterns provided by an embodiment of the present application;
[0052] Figure 8 The schematic diagram of perception-guided conditional encoding provided by an embodiment of the present application. Detailed implementation manners
[0053] The following details the implementation manners of the present application. The examples of the implementation manners are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The implementation manners described below with reference to the drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.
[0054] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more of the associated listed items.
[0055] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as such here.
[0056] For the convenience of understanding the embodiments of this application, the following will further explain with several specific embodiments in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of this application.
[0057] Embodiment 1
[0058] As Figure 1 and Figure 2 shown, a transformable video feature encoding method for machine vision perception includes the following steps:
[0059] S101, extract middle-level features from the original frame sequence through a feature extractor.
[0060] In this step, obtain the original frame sequence from the video data to be processed, and extract the middle-level features from the original frame sequence through a feature extractor. Among them, the feature extractor can be a convolutional neural network such as ResNet, VGGNet, DenseNet, etc. Schematically, the feature extractor is the res2 layer of the ResNet-50 backbone network in Faster R-CNN.
[0061] It should be understood that the middle-level features here contain more general information about the image compared to the subsequent high-level features and provide the potential for multi-task analysis. In addition, the middle-level features retain the spatial structure in the original frame, making it possible to more effectively eliminate redundancy through a neural network. Compared with shallow features, the intermediate features are initially extracted to filter out information irrelevant to the machine vision task, making them easier to compress.
[0062] S102, perform motion compensation on the extracted middle-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and the middle-level features to obtain reconstructed features.
[0063] S103, perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features into at least one downstream task to obtain a prediction result.
[0064] Among them, the downstream task is a task related to machine vision, which can be object detection, semantic segmentation, instance segmentation, object tracking, image reconstruction, style transfer, etc., and is not limited thereto.
[0065] In this step, the feature space transformation module transfers the reconstructed features to other feature spaces to adapt to different downstream tasks.
[0066] It should be emphasized that the number of feature space transformation modules is the same as the number of downstream tasks. Different feature space transformation modules can perform different feature transformations on the reconstructed features according to the different requirements of different downstream tasks for channels and spatial shapes, so that the output transformed features can be aligned with the downstream task features.
[0067] The transformable video feature encoding method for machine vision perception provided by the embodiments of the present application extracts middle-level features from the original frame sequence through a feature extractor. The middle-level features retain the original spatial structure and filter out information irrelevant to the downstream task, making them easier to compress. Motion compensation and reconstruction are implemented through a video feature codec, which ensures that there are only minor differences between the original frame sequence and the predicted frame sequence, realizes targeted bitrate allocation, maximally eliminates temporal redundancy through motion compensation, and effectively removes spatial redundancy. The reconstructed features are transferred to other feature spaces through a feature space transformation module to adapt to different downstream tasks, thus solving the problem that re-training and re-deploying the upstream feature codec or the entire downstream task network requires extremely high computing resources in practical applications. In summary, the present application provides a highly adaptable and scalable video coding solution for machine vision perception.
[0068] In some embodiments of the present application, such as Figure 3As shown in the figure, the video feature codec includes a feature buffer, a motion estimation module, a motion codec, a motion compensation module, and a conditional codec. The feature buffer stores reference features, and the motion codec includes a motion encoder and a motion decoder.
[0069] The middle-level features are motion-compensated by the video feature codec, and the compensated features and the middle-level features are subjected to detail reconstruction to obtain reconstructed features, including:
[0070] The motion information between the reference features and the middle-level features is estimated by the motion estimation module, the motion information is encoded and decoded by the motion codec, and the encoded and decoded motion information is motion-compensated based on the reference features by the motion compensation module to obtain compensated features.
[0071] In this step, the motion information between the reference features and the middle-level features is used and compressed by the motion encoder to form a compact motion representation, and then the compact motion representation is decoded by the motion decoder to obtain the encoded and decoded motion information. In motion compensation, an inter-frame feature domain prediction method of combining motion patterns is adopted to generate various potential motion schemes from the reference features, and they are selectively combined based on the encoded and decoded motion information to obtain compensated features.
[0072] Based on the encoding perception condition and the decoding perception condition, and through the conditional codec, the detailed information between the compensated features and the middle-level features is reconstructed to obtain the reconstructed features.
[0073] In this step, the encoding perception condition and the decoding perception condition are prior knowledge. The detailed information between the compensated features and the middle-level features is compressed using the encoding perception condition and the decoding perception condition as the prior, so as to obtain the reconstructed features, which can better remove spatial redundancy.
[0074] It should be understood that the task of video coding for machine vision perception is to ensure a certain compression ratio while not having too much impact on the performance of its downstream tasks.
[0075] In some embodiments of the present application, the video feature codec further includes a channel reduction module and an entropy model, and the conditional codec includes a conditional encoder and a conditional decoder.
[0076] Before estimating the motion information between the reference features and the middle-level features by the motion estimation module, the method further includes:
[0077] The reference features and the middle-level features are respectively subjected to channel reduction by the channel reduction module to obtain the reduced reference features and the reduced middle-level features.
[0078] That is, for the reference feature F ref and the middle-level feature F tPerform channel reduction to reduce the dimensions of two features for compression, and finally obtain the reduced reference feature f ref and the reduced middle-level feature f t . Schematically, the 256-dimensional reference feature F ref and the middle-level feature F t can be reduced to 64 dimensions, or can be reduced to 128 dimensions, or reduce 128 dimensions to 64 dimensions, etc., which is not limited herein.
[0079] Estimate the motion information between the reference feature and the middle-level feature through the motion estimation module, including:
[0080] Estimate the motion information between the reduced reference feature and the reduced middle-level feature through the motion estimation module, encode and decode the motion information through the motion codec, and perform motion compensation on the encoded and decoded motion information based on the reduced reference feature through the motion compensation module to obtain the compensated feature.
[0081] Specifically, perform motion estimation, motion encoding and decoding, and motion compensation on the reduced reference feature f ref and the reduced middle-level feature f t according to the following formula:
[0082]
[0083] In the formula, ME(·) represents motion estimation, MC(·) represents motion compensation, E m (·) and D m (·) represent the motion encoder and the motion decoder respectively, represents the quantization operation, represents the compensated feature.
[0084] Among them, the motion estimation module consists of a convolutional layer and a residual block, and uses the motion estimation module to estimate the motion information m t between the reference feature and the middle-level feature. Input the motion information m t into the motion codec to generate the encoded and decoded motion information
[0085] In the motion compensation module, adopt the inter-frame feature domain prediction method of motion mode combination, and use the reduced reference feature f ref to generate various potential motion schemes and the encoded and decoded motion information through multiple depthwise separable convolutions and generate multi-scale features through the 1×1 convolutional kernel, and selectively combine them to obtain the compensated feature
[0086] Based on the encoding perception condition and the decoding perception condition, and reconstruct the detailed information between the compensated feature and the middle-level feature through a conditional codec to obtain a reconstructed feature, including:
[0087] Based on the encoding perception condition Cenc , and encode the compensated feature and the reduced middle-level feature through a conditional encoder to obtain an encoded feature.
[0088] Compress the information of the compensated feature through an entropy model to obtain a compressed compensated feature.
[0089] Based on the decoding perception condition Cdec , and decode the compensated feature, the encoded feature and the compressed compensated feature through a conditional decoder to obtain a reconstructed feature.
[0090] Specifically, input the compensated feature and the reduced middle-level feature f t into the perception-guided conditional encoder, take the encoding perception condition C enc as prior knowledge, and encode the compensated feature and the reduced middle-level feature f t to obtain an encoded feature. In addition, the information of the compensated feature is also compressed through an entropy model to obtain a compressed compensated feature. The conditional decoder takes the decoding perception condition C dec as prior knowledge, and decodes the encoded feature, the compressed compensated feature and the compensated feature to obtain a reconstructed feature Among them, the two perception conditions are strategically inserted at the positions aligned with the 1 / 4, 1 / 8 and 1 / 16 scales, and used as prior conditions for encoding and decoding to enhance the overall performance of the codec.
[0091] Schematically, the reconstructed feature is obtained according to the following formula
[0092]
[0093] where E c (·) and D c (·) respectively represent the perception-guided conditional encoder and the perception-guided conditional decoder.
[0094] In addition, the video feature codec further includes a channel restoration module, and the channel dimension of the reconstructed feature is restored to the channel dimension before being processed by the channel reduction module through the channel restoration module, so as to obtain the final reconstructed feature and input it into the feature buffer as a new reference feature for subsequent use.
[0095] In some embodiments of the present application, such asFigure 3 As shown, the motion estimation module includes a convolutional layer and multiple residual blocks. After concatenating the reference feature and the middle-level feature, they are input into the motion estimation module and processed by the convolutional layer and multiple residual blocks to obtain motion information.
[0096] The motion compensation module includes depthwise separable convolution, depthwise separable convolution blocks, 1×1 convolution, and residual blocks.
[0097] The motion encoder includes a convolutional layer, a generalized normalization layer, and residual blocks. The motion decoder includes residual blocks, an inverse generalized normalization layer, and transposed convolution.
[0098] In some embodiments of the present application, as Figure 3 shown, the video feature codec further includes a feature pyramid network.
[0099] The encoding perception condition is extracted from the middle-level feature by the feature pyramid network.
[0100] The decoding perception condition is extracted from the compensated feature by the feature pyramid network.
[0101] Specifically, the encoding perception condition C enc is the multi-scale high-level feature inferred by the feature pyramid network based on Faster R-CNN from the middle-level feature F t . Due to the invisibility of the middle-level feature F t during decoding, the decoding perception condition C dec is the multi-scale high-level feature inferred by the feature pyramid network based on the compensated feature after channel recovery . It is the feature after the compensated feature undergoes channel dimension recovery through the above-mentioned channel recovery module. The two perception conditions can be strategically inserted at the positions aligned with 1 / 4, 1 / 8, and 1 / 16 scales. As Figure 3 shown, the encoding perception condition C enc = {p2, p3, p4} and the decoding perception condition C dec = {p2, p3, p6} are inserted into different positions respectively.
[0102] In some embodiments of the present application, as Figure 4As shown, the feature space transformation module includes a first branch, a second branch, and a third branch. The first branch includes a downsampling module and an upsampling module connected in sequence. The second branch includes three bottleneck residual blocks connected in sequence. The third branch includes an upsampling module and a downsampling module connected in sequence. The upsampling module includes multiple transposed convolutions and bottleneck structures, which can gradually restore the resolution of the feature map. The downsampling module includes a convolutional layer, a normalization layer, a RELU function, a convolutional layer, and two bottleneck residual blocks connected in sequence, which can reduce the resolution to obtain a feature map containing rich global information.
[0103] Performing feature transformation on the reconstructed features through the feature space transformation module includes:
[0104] Performing information retention on the reconstructed features through the first branch to obtain the first transformed features.
[0105] Performing shape migration on the reconstructed features through the second branch to obtain the second transformed features.
[0106] Performing global information extraction on the reconstructed features through the third branch to obtain the third transformed features.
[0107] Concatenating the first transformed features, the second transformed features, and the third transformed features to obtain the final transformed features.
[0108] Specifically, roughly reconstructing the current frame through the first branch can retain content information in the pixel domain, thereby enhancing the information retention ability during the feature transfer process. The second branch can promote the feature migration of the original shape. The third branch focuses on the extraction of global information. By concatenating the transformed features generated by the three branches and using the last convolutional layer to align the channels and spatial shapes of the concatenated transformed features with the specific downstream task features, the final transformed features are obtained. Then the final transformed features are input into their respective downstream tasks to obtain the prediction results corresponding to each downstream task.
[0109] The final transformed features are obtained through the following formula:
[0110]
[0111] In the formula, FST i (·) represents the i-th feature space transformation module, represents the final transformed features applicable to the i-th downstream task.
[0112] The final transformed feature vector Input into the downstream task network. Specifically, in this instance, three downstream task networks are set up: object detection, semantic segmentation, and instance segmentation, and the final prediction results y are obtained respectively. i . The specific principle can be expressed by the formula:
[0113]
[0114] Among them, TASK i (·) represents the i-th downstream task network, and y i represents the prediction result of the i-th downstream task (1 represents object detection, 2 represents semantic segmentation, and 3 represents instance segmentation).
[0115] In some embodiments of the present application, the transformable video feature encoding method provided by the present application for machine vision perception determines the effect based on the following settings:
[0116] Use the following three downstream task networks: use the CrossVIS framework for video instance segmentation; use Deeplab-v3 for semantic segmentation; use Faster R-CNN for object detection. The parameters of all downstream task networks are frozen throughout the experiment.
[0117] Experiments were conducted on the YoutubeVIS-2019 (YTVIS-2019) and Video Scene Parsing in the Wild (VSPW) datasets. Among them, the YTVIS-2019 dataset is a large video dataset, which includes 2883 videos and has frame-level annotations of 40 categories for video instance segmentation. The VSPW dataset is a large video dataset, which contains 3536 480P resolution videos of 231 scenes, and it has frame-level annotations of 124 categories for video semantic segmentation.
[0118] All experiments were completed on a single NVIDIA RTX 3090 24GB GPU, and the batch size on the GPU was 4. The network was trained in two stages. In the first stage, the video feature codec was trained, and it was trained on the YTVIS-2019 dataset and iteratively trained for 5 stages. Among them, the learning rate in the first 3 stages was The learning rate in the 4th fine-tuning stage was The learning rate in the last fine-tuning stage was In the second stage, the feature space transformation module was trained, and the total number of training iterations was 100K, and the learning rate was set to In addition, the input features during training were cropped to 128*128.
[0119] To evaluate the method provided in this application, it is also compared with traditional hybrid codecs VTM-23.1, HM-18.0, and x265 (FFmpeg-4.2.7), and contrasted with the open-source neural video compression (NVC) framework, such as DCVC-DC, DCVC-HEM, DCVC-TCM, DCVC, and FVC. For the compared NVC methods, all available pre-trained models are evaluated by different metrics (PSNR, MS-SSIM, YUV), and only the models with the best performance are shown. In addition, the video codec SMC++ oriented to machine vision perception is also used as a comparison method. VTM-23.1 serves as an anchor for calculating BD-Rate (a lower BD-Rate means more bitrate savings).
[0120] Figure 5 (a), Figure 5 (b), and Figure 5 (c) are respectively schematic diagrams of performance comparison for object detection tasks, semantic segmentation, and instance segmentation. In the upper part of the data graphs in Figure 5 (a), Figure 5 (b), and Figure 5 (c), the horizontal axis represents the number of bits required to encode per pixel (Bit per pixel, Bpp), and the vertical axis represents the average precision (Average Precision, AP) of the downstream task performance metric; in the lower part of the data graphs, the horizontal axis represents the execution time (Execution Time), and the vertical axis represents the mean intersection over union (mean Intersection over Union, mIoU). Compared with other video coding methods, the method provided in this application (i.e., Trans VFC) achieves the best performance and the best rate performance on the object detection and instance segmentation test datasets, exceeding the latest open-source neural video compression methods and traditional codecs HM and VTM. In terms of semantic segmentation, it is superior to the best deep learning-based method SMC++ in terms of rate-task performance. In addition, compared with the open-source neural video compression methods, it achieves the best balance of speed performance.
[0121] The comparison schematic diagram of instance segmentation is as shown in Figure 6 . As can be seen from Figure 6 , the present invention obtains better subjective segmentation results at different bitrates. Although the quality of the reconstructed frames is high, it is difficult for the downstream task network CrossVIS to maintain the segmentation consistency of the main objects (such as skateboards and people with umbrellas), and often mis-segments them into multiple instances. In contrast, the present invention better maintains the consistency of instances and maximally retains the original segmentation results.
[0122] As shown inFigure 7 As shown, the inter-frame feature domain prediction method for motion pattern combinations mainly focuses on generating potential motion patterns and combining them through motion representations. Motion representations carry the motion information between two frames, including local edge motion ( Figure 7 schematic diagrams shown in channels 0 and 8 in Figure 7 and large-scale motion (
[0123] such as Figure 8 the fast movement of the vehicle in the schematic diagrams shown in channels 4 and 6 in
[0124] Embodiment 2
[0125] Based on Embodiment 1, this Embodiment 2 provides a transformable video feature encoding device for machine vision perception. This transformable video feature encoding device for machine vision perception corresponds to the above-mentioned transformable video feature encoding method for machine vision perception, and specifically includes:
[0126] A feature extraction module, configured to extract middle-level features from the original frame sequence through a feature extractor;
[0127] A compensation and reconstruction module, configured to perform motion compensation on the extracted middle-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and middle-level features to obtain reconstructed features;
[0128] A transformation and prediction module, configured to perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features into at least one downstream task to obtain a prediction result.
[0129] For specific details, refer to the description in the part of the transformable video feature encoding method for machine vision perception, which will not be elaborated here.
[0130] Example 3
[0131] Example 3 of the present application provides an electronic device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute a transformable video feature encoding method for machine vision perception. The method includes the following process steps:
[0132] Extract middle-level features from the original frame sequence through a feature extractor;
[0133] Perform motion compensation on the extracted middle-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and the middle-level features to obtain reconstructed features;
[0134] Perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features into at least one downstream task to obtain a prediction result.
[0135] Example 4
[0136] Example 4 of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements a transformable video feature encoding method for machine vision perception. The method includes the following process steps:
[0137] Extract middle-level features from the original frame sequence through a feature extractor;
[0138] Perform motion compensation on the extracted middle-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and the middle-level features to obtain reconstructed features;
[0139] Perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features into at least one downstream task to obtain a prediction result.
[0140] Example 5
[0141] Example 5 of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a transformable video feature encoding method for machine vision perception. The method includes the following process steps:
[0142] Extract middle-level features from the original frame sequence through a feature extractor;
[0143] Perform motion compensation on the extracted middle-level features through a video feature codec, and perform detail reconstruction on the obtained compensated features and the middle-level features to obtain reconstructed features;
[0144] The reconstructed features are subjected to feature transformation through the feature space transformation module, and the transformed features are input into at least one downstream task to obtain a prediction result.
[0145] Those of ordinary skill in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily essential for implementing the present application.
[0146] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for method or system embodiments, since they are basically similar to method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The method and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0147] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A transformable video feature encoding method for machine vision perception, characterized in that: include: Extract mid-level features from the original frame sequence through a feature extractor; Performing motion compensation on the extracted middle-level features through a video feature codec, and reconstructing details of the obtained compensated features and the middle-level features to obtain reconstructed features; The reconstructed features are transformed by a feature space transformation module, and the transformed features are input into at least one downstream task to obtain a prediction result.
2. The method according to claim 1, characterized in that The video feature codec includes a feature buffer, a motion estimation module, a motion codec, a motion compensation module and a conditional codec, and the feature buffer stores reference features; The method of performing motion compensation on the middle-level features by using a video feature codec, and reconstructing the obtained compensated features and the middle-level features in detail to obtain reconstructed features, includes: Estimate motion information between the reference feature and the middle-level feature by a motion estimation module, encode and decode the motion information by the motion codec, and perform motion compensation on the encoded and decoded motion information by the motion compensation module based on the reference feature to obtain a compensated feature; Based on the encoding perception condition and the decoding perception condition, detail information between the compensation feature and the middle-level feature is reconstructed through the conditional codec to obtain a reconstructed feature.
3. The method according to claim 2, characterized in that The video feature codec also includes a channel reduction module and an entropy model, and the conditional codec includes a conditional encoder and a conditional decoder; Before estimating the motion information between the reference feature and the mid-level feature by the motion estimation module, the method further includes: Perform channel reduction on the reference feature and the middle-level feature respectively by the channel reduction module to obtain a reduced reference feature and a reduced middle-level feature; The estimating the motion information between the reference feature and the middle-level feature by a motion estimation module includes: Estimating motion information between the reduced reference features and the reduced middle-level features by a motion estimation module; The step of reconstructing detail information between the compensation feature and the middle-level feature based on the encoding perception condition and the decoding perception condition by the conditional codec to obtain the reconstructed feature includes: Based on the encoding perception condition, the compensating feature and the reduced middle-level feature are encoded by the conditional encoder to obtain an encoded feature; Compressing the information of the compensation feature by using the entropy model to obtain a compressed compensation feature; Based on the decoding perceptual condition, the compensation feature, the encoded feature and the compressed compensation feature are decoded by the conditional decoder to obtain the reconstructed feature.
4. The method according to claim 2, characterized in that: The video feature codec also includes a feature pyramid network; The encoding perception condition is obtained by extracting from the middle-level features through the feature pyramid network; The decoding perception condition is obtained by extracting from the compensation feature through the feature pyramid network.
5. The method according to claim 1, characterized in that The feature space transformation module includes a first branch, a second branch and a third branch, the first branch includes a downsampling module and an upsampling module connected in sequence, the second branch includes three bottleneck residual blocks connected in sequence, and the third branch includes an upsampling module and a downsampling module connected in sequence; The step of performing feature transformation on the reconstructed features by using a feature space transformation module includes: Retaining information of the reconstructed feature through the first branch to obtain a first transformed feature; Performing shape migration on the reconstructed feature through the second branch to obtain a second transformed feature; Performing global information extraction on the reconstructed feature through the third branch to obtain a third transformed feature; The first transformed feature, the second transformed feature and the third transformed feature are concatenated to obtain a final transformed feature.
6. The method according to any one of claims 1 to 5, characterized in that: The downstream tasks include object detection, semantic segmentation, and instance segmentation.
7. A transformable video feature encoding device for machine vision perception, characterized in that: include: A feature extraction module, used to extract mid-level features from the original frame sequence through a feature extractor; A compensation and reconstruction module, used for performing motion compensation on the extracted middle-level features through a video feature codec, and performing detailed reconstruction on the obtained compensated features and the middle-level features to obtain reconstructed features; The transformation and prediction module is used to perform feature transformation on the reconstructed features through a feature space transformation module, and input the transformed features to at least one downstream task to obtain a prediction result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.