Video encoding and decoding acceleration method and system based on learnable task perception mechanism
By introducing image group optimization and multi-scale feature processing based on learningable task perception mechanism in video encoding and decoding technology, combining dynamic quantization and intelligent entropy coding, the problems of waste of computing resources and low encoding and decoding efficiency in the prior art are solved, and efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202411518429.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-29
AI Technical Summary
When existing video encoding and decoding technologies deal with video streams containing a large amount of redundant information or static scenes, computing resources are severely wasted, and the way of processing keyframes and non-keyframes is lacking in flexibility, resulting in low encoding and decoding efficiency.
The video encoding and codec acceleration method based on the learnable task perception mechanism is adopted, and the image group is constructed and optimized, combined with high/low resolution codecs for multi-scale feature extraction and restoration, and dynamic quantization and intelligent entropy encoding technology are used to optimize the encoding and codec process.
While maintaining visual quality, the complexity of encoding and decoding operations is reduced, the encoding and decoding efficiency is improved, and the bit rate can be reduced by 30-40% on average, which improves the transmission efficiency and quality of the video stream.
Smart Images

Figure CN119031147B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video compression, and in particular to a video encoding and decoding acceleration method and system based on a learnable task perception mechanism. Background Art
[0002] In today's era of rapid development of multimedia technology, video content has become the main body of Internet traffic, posing unprecedented challenges to the efficiency and real-time performance of video encoding and decoding technology. In many application scenarios, such as power inspection, real-time video communication and unmanned driving, with the in-depth development of artificial intelligence and deep learning, video processing technology is undergoing an unprecedented transformation, especially in the field of video encoding and decoding.
[0003] Although traditional video coding and decoding methods have made remarkable achievements in compression efficiency, they often ignore the intelligent perception and processing of the dynamic characteristics of video content, resulting in a large waste of computing resources, especially when processing video streams containing a large amount of redundant information or static scenes; at the same time, traditional technologies are relatively fixed in the processing of key frames and non-key frames, lacking flexibility, resulting in low coding and decoding efficiency.
[0004] Therefore, further research and innovation are needed to solve the above problems existing in the prior art. Summary of the invention
[0005] The purpose of the invention is to provide a video encoding and decoding acceleration method and system based on a learnable task-aware mechanism to solve the above-mentioned problems existing in the prior art.
[0006] The technical solution is a video encoding and decoding acceleration method based on a learnable task-aware mechanism, comprising the following steps:
[0007] S1. Obtaining original video sequence data from a video source, and preprocessing it to obtain preprocessed video frame data; based on the preprocessed video frame data, obtaining a preset number of continuous video frame data to form an initial image group; based on the initial image group, extracting video frame features; based on the video frame features, calculating an optimal image group length; based on the optimal image group length and the preprocessed video frame data, constructing final image group structure data;
[0008] S2. Based on the video frames in the image group structure data, a pre-trained target detection model is used to obtain the target detection result of each frame; based on the target detection result, the target importance score of each frame is calculated; based on the adjacent frames in the image group structure data, an optical flow estimation algorithm is used to obtain the inter-frame motion information; based on the target importance score and the inter-frame motion information, a feature vector sequence is constructed; the feature vector sequence is input into the pre-trained image group selection network to obtain the importance prediction value of each frame; according to the preset importance threshold, the importance prediction value is binarized to obtain the validity mark of each frame; the validity mark is combined with the image group structure data to generate the optimized image group structure data;
[0009] S3, inputting the optimized image group structure data into the pre-trained high-resolution encoder and low-resolution encoder in parallel to obtain a high-resolution feature map and a low-resolution feature map; applying the self-attention mechanism to them respectively to obtain enhanced high-resolution feature maps and low-resolution feature maps; performing feature fusion on the enhanced high-resolution feature maps and low-resolution feature maps to obtain a multi-scale feature representation; obtaining the current quantization parameter and calculating the dynamic quantization step size; quantizing the multi-scale feature representation using the dynamic quantization step size to obtain discretized feature data; applying entropy coding to the discretized feature data to obtain a preliminarily compressed bit stream; packaging the preliminarily compressed bit stream and the dynamic quantization step size to form a quantized data packet;
[0010] S4, parsing the quantized data packet to obtain quantized feature values; using a pre-trained dynamic space selection network to process the quantized feature values to generate a spatial importance map; constructing a binary mask based on the spatial importance map and a preset threshold; processing the quantized feature values based on the binary mask to obtain retained feature data; converting the retained feature data into a one-dimensional sequence, entropy encoding the one-dimensional sequence using context-adaptive binary arithmetic coding to generate a compressed bit stream; generating a coded data packet based on the binary mask and the compressed bit stream;
[0011] S5. Parse the encoded data packet to obtain a compressed bit stream and encoding metadata; use an arithmetic entropy decoder to decode the compressed bit stream to obtain quantized feature data and a spatial selection mask; reconstruct the quantized feature data into a two-dimensional feature map based on the spatial selection mask; use the quantization parameter in the encoding metadata to perform a dequantization operation on the two-dimensional feature map to obtain a dequantized feature map; input the dequantized feature map into a cascaded low-resolution decoder and a high-resolution decoder to obtain the output of the high and low decoders; use a feature fusion module to merge the outputs of the high and low decoders to obtain a final reconstructed video frame; based on the reconstructed video frame, use a deblocking effect filter for processing to obtain an optimized reconstructed video frame; based on the optimized reconstructed video frame, generate a final video sequence.
[0012] Video codec acceleration system based on learnable task-aware mechanism, including:
[0013] at least one processor; and,
[0014] a memory communicatively connected to at least one of the processors; wherein,
[0015] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the video encoding and decoding acceleration method based on the learnable task-aware mechanism.
[0016] Beneficial effect: the present invention reduces the waste of resources and calculations by constructing an image group; multi-scale feature extraction and restoration of the current frame are performed through a high / low resolution codec, which reduces the complexity of the encoding and decoding operation and improves the encoding and decoding efficiency while maintaining the visual quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of the present invention.
[0018] Figure 2 This is a flow chart of step S1 of the present invention.
[0019] Figure 3 This is a flow chart of step S2 of the present invention.
[0020] Figure 4 This is a flow chart of step S3 of the present invention.
[0021] Figure 5 This is a flow chart of step S4 of the present invention.
[0022] Figure 6 This is a flow chart of step S5 of the present invention.
[0023] Figure 7 This is a network structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0024] like Figure 1 As shown, the present application proposes a video encoding and decoding acceleration method based on a learnable task-aware mechanism, comprising the following steps:
[0025] S1. Obtaining original video sequence data from a video source, and preprocessing it to obtain preprocessed video frame data; based on the preprocessed video frame data, obtaining a preset number of continuous video frame data to form an initial image group; based on the initial image group, extracting video frame features; based on the video frame features, calculating an optimal image group length; based on the optimal image group length and the preprocessed video frame data, constructing final image group structure data;
[0026] S2. Based on the video frames in the image group structure data, a pre-trained target detection model is used to obtain the target detection result of each frame; based on the target detection result, the target importance score of each frame is calculated; based on the adjacent frames in the image group structure data, an optical flow estimation algorithm is used to obtain the inter-frame motion information; based on the target importance score and the inter-frame motion information, a feature vector sequence is constructed; the feature vector sequence is input into the pre-trained image group selection network to obtain the importance prediction value of each frame; according to the preset importance threshold, the importance prediction value is binarized to obtain the validity mark of each frame; the validity mark is combined with the image group structure data to generate the optimized image group structure data;
[0027] S3, inputting the optimized image group structure data into the pre-trained high-resolution encoder and low-resolution encoder in parallel to obtain a high-resolution feature map and a low-resolution feature map; applying the self-attention mechanism to them respectively to obtain enhanced high-resolution feature maps and low-resolution feature maps; performing feature fusion on the enhanced high-resolution feature maps and low-resolution feature maps to obtain a multi-scale feature representation; obtaining the current quantization parameter and calculating the dynamic quantization step size; quantizing the multi-scale feature representation using the dynamic quantization step size to obtain discretized feature data; applying entropy coding to the discretized feature data to obtain a preliminarily compressed bit stream; packaging the preliminarily compressed bit stream and the dynamic quantization step size to form a quantized data packet;
[0028] S4, parsing the quantized data packet to obtain quantized feature values; using a pre-trained dynamic space selection network to process the quantized feature values to generate a spatial importance map; constructing a binary mask based on the spatial importance map and a preset threshold; processing the quantized feature values based on the binary mask to obtain retained feature data; converting the retained feature data into a one-dimensional sequence, entropy encoding the one-dimensional sequence using context-adaptive binary arithmetic coding to generate a compressed bit stream; generating a coded data packet based on the binary mask and the compressed bit stream;
[0029] S5. Parse the encoded data packet to obtain a compressed bit stream and encoding metadata; use an arithmetic entropy decoder to decode the compressed bit stream to obtain quantized feature data and a spatial selection mask; reconstruct the quantized feature data into a two-dimensional feature map based on the spatial selection mask; use the quantization parameter in the encoding metadata to perform a dequantization operation on the two-dimensional feature map to obtain a dequantized feature map; input the dequantized feature map into a cascaded low-resolution decoder and a high-resolution decoder to obtain the output of the high and low decoders; use a feature fusion module to merge the outputs of the high and low decoders to obtain a final reconstructed video frame; based on the reconstructed video frame, use a deblocking effect filter for processing to obtain an optimized reconstructed video frame; based on the optimized reconstructed video frame, generate a final video sequence.
[0030] like Figure 2As shown, according to one aspect of the present application, step S1 is further:
[0031] S11, receiving an original video data stream from a preset video input interface, parsing the original video data stream into separate original video sequence data; performing noise reduction processing based on the original video sequence data using a Gaussian filter algorithm to obtain noise-reduced video frame data; converting the noise-reduced video frame data from an RGB color space to a YUV color space to obtain color-converted video frame data; performing bilinear interpolation resampling on the color-converted video frame data according to a preset target resolution to obtain video frame data after adjusting the resolution; and storing the video frame data after adjusting the resolution in a frame buffer in chronological order;
[0032] S12, reading a preset number of continuous video frame data from the frame buffer, performing inter-frame difference calculation on the continuous video frame data to obtain an inter-frame difference value sequence; using a sliding window method to analyze the inter-frame difference value sequence, calculating a local peak position, and obtaining a scene switching candidate point;
[0033] S13, based on the previous and next frames of each scene switching candidate point, using an edge detection algorithm, calculate the similarity of edge distribution to obtain a scene switching probability value; according to a preset scene switching threshold, screen the scene switching probability value to determine the final scene switching point; with the scene switching point as the boundary, divide the continuous video frame data into a predetermined number of subsequences; calculate the average motion vector amplitude for each subsequence to obtain a motion complexity index; based on the motion complexity index and a preset image group length range, determine the optimal image group length for each subsequence;
[0034] S14. Based on the optimal GOP length, each subsequence is divided into a predetermined number of GOPs; for each GOP, the first frame is designated as a key frame, and the remaining frames are allocated as forward prediction frames or bidirectional prediction frames according to a preset encoding strategy, to form final GOP structure data.
[0035] In one embodiment of the present application, the original video data stream is read from a preset video input interface. A multi-threaded reading mechanism is used to simultaneously read multiple video clips, each of which is 4MB in size. The read video data stream is input into a video decoder, and the video data is decoded using an H.264 or H.265 decoding algorithm. Individual video frames are extracted from the decoded data to obtain an original video frame sequence. The extracted original video frame sequence is stored in a frame buffer.
[0036] Read the original video frame sequence from the frame buffer. Apply the adaptive Gaussian filter algorithm to each frame of video data for spatial domain noise reduction. The filter σ value is dynamically adjusted according to the local noise level. The calculation formula is σ(x, y) = σ base +k*noise level(x, y), where σ base is the base σ value (such as 1.5), k is the adjustment coefficient, and noise level is the estimated local noise level. For high-motion regions, a time-domain noise reduction algorithm based on Kalman filtering is applied. The results of spatial-domain and time-domain noise reduction are weighted and fused to obtain a noise-reduced video frame sequence.
[0037] Receive the noise-reduced video frame sequence. Use an optimized RGB-to-YUV conversion algorithm to perform color space conversion on each frame. The conversion formulas are: Y = 0.299R + 0.587G + 0.114B, U = -0.147R - 0.289G + 0.436B, V = 0.615R - 0.515G - 0.100B, where Y represents luminance, R represents the red component, G represents the green component, B represents the blue component, U represents the blue difference component, and V represents the red difference component; use the Single Instruction Multiple Data (SIMD) instruction set (such as SSE or AVX) to accelerate the conversion process. Output the video frame sequence in the YUV color space.
[0038] Read the video frame sequence in the YUV color space. Implement an adaptive resolution adjustment strategy, and use different downsampling rates for different regions according to the preset target resolution and video content characteristics. Use the Lanczos resampling algorithm for resolution adjustment. The Lanczos kernel function is defined as: if -a < x < a, Lanczos(x) = sinc(x) * sinc(x / a); otherwise 0, where a is the parameter of the Lanczos kernel, usually taking 2 or 3, and sinc( ) represents the sinc function. Output the video frame sequence with adjusted resolution.
[0039] Receive the video frame sequence with adjusted resolution. Apply the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm to each frame for adaptive contrast enhancement. Then use the Unsharp Masking technique for image sharpening. The sharpening formula is: I sharp = I + λ(I - I blur ), where I is the original image, I blur is the image after Gaussian blur, and λ is the sharpening intensity. Finally, output the enhanced video frame sequence as the final result of preprocessing.
[0040] Read preprocessed continuous video frames from the video frame buffer. Perform pixel-level difference calculation on adjacent frames to obtain an inter-frame difference map. Calculate the mean and standard deviation of the inter-frame difference map as statistical features of inter-frame differences. Use the sliding window method to analyze the inter-frame difference sequence and calculate the local peak position. Output the inter-frame difference statistical features and local peak position information. Receive the inter-frame difference statistical features and local peak position information. Apply an edge detection algorithm, such as the Canny edge detector, to the frames before and after each local peak position. Calculate the similarity of the edge distribution of adjacent frames to obtain the scene switching probability value. Set an adaptive scene switching threshold and dynamically adjust it according to the global statistical characteristics of the video content. Compare the scene switching probability value with the threshold to determine the final scene switching point. Output a list of scene switching points.
[0041] Read the preprocessed video frame sequence and scene switching point list. Use block matching algorithm or optical flow estimation method to calculate the motion vector between adjacent frames. Perform statistical analysis on the motion vectors within each scene and calculate the average motion amplitude and directional consistency. Combine the motion amplitude and directional consistency to evaluate the motion complexity of each scene. Output the motion complexity index of each scene. Receive the scene switching point list and the motion complexity index. Based on the scene switching points, divide the video sequence into multiple subsequences. For each subsequence, use an adaptive algorithm to determine the optimal GoP length based on its motion complexity index and the preset GoP length range. Consider the balance between coding efficiency and random access performance, and fine-tune the GoP length. Output the GoP length of each subsequence.
[0042] Read the subsequence GoP length information and the preprocessed video frame sequence. According to the determined GoP length, further split each subsequence into multiple GoPs. For each GoP, specify the first frame as a key frame (I frame). According to the preset encoding strategy and motion complexity, assign the remaining frames as forward prediction frames (P frames) or bidirectional prediction frames (B frames). Generate complete GoP structure data, including the length, frame type sequence and key reference relationship of each GoP. Associate the GoP structure data with the corresponding video frame data and output the optimized GoP structure information.
[0043] In another embodiment of the present application, a 30-minute 1080p traffic surveillance video with a frame rate of 30fps is prepared; a pre-trained target detection model (such as YOLOv5) is used for vehicle detection; a pre-trained optical flow estimation model (such as PWC-Net); a pre-trained high / low resolution codec network; a pre-trained dynamic space selection network. Input video sequence and perform GoP selection: divide the 30-minute video into 10-second segments, each with 300 frames; perform the following operations on each segment: apply Gaussian filtering for noise reduction, σ=1.5; convert RGB color space to YUV; adjust the resolution to 720p. Calculate the scene complexity score C for each segment: C = w1*E +w2*M+ w3*T, where: E represents edge density (using Canny edge detection); M represents average motion amplitude (using optical flow estimation); T represents texture complexity (using gray-level co-occurrence matrix) w1, w2, w3 are weights, initially set to 0.3, 0.4, 0.3. Dynamically adjust the GoP length L according to the scene complexity score C: L = max(Lmin, min(Lmax, round(α / C))), where: Lmin = 15, Lmax = 60, α = 1000 (adjustable parameter).
[0044] This embodiment improves the efficiency and flexibility of video encoding by introducing an adaptive GoP structure adjustment mechanism based on scene complexity. The length and structure of GoP are dynamically adjusted by real-time analysis of the complexity of video content, including factors such as edge density, texture complexity and motion intensity. This adaptive adjustment strategy enables the encoder to optimize processing for different video scene features. For scenes with high complexity, such as fast motion or rich details, the system will automatically shorten the GoP length and increase the frequency of key frames to ensure encoding quality and reduce error propagation. On the contrary, for scenes with low complexity, such as static backgrounds or slowly changing pictures, the system will extend the GoP length and reduce redundant key frames, thereby improving compression efficiency. This dynamic adjustment not only optimizes bit rate allocation, but also improves overall compression efficiency. In practical applications, the bit rate consumption can be reduced by an average of 15-20% while maintaining the same video quality. At the same time, by reducing unnecessary key frames, the encoding delay is reduced, which is particularly suitable for real-time video transmission scenarios. In addition, the adaptive GoP structure also improves the robustness of the video stream because it can quickly adjust the encoding strategy according to changes in network conditions and reduce the problem of image freeze and quality degradation in an unstable network environment. In the video content analysis task, more key visual information can be retained, improving the accuracy of subsequent tasks such as target detection and tracking. This embodiment not only improves the coding efficiency, but also enhances the system's adaptability to different types of video content and network environments, providing support for efficient and flexible video coding.
[0045] like Figure 3 As shown, according to one aspect of the present application, step S2 is further:
[0046] S21. Based on the video frames in the image group structure data, a pre-trained feature extraction network is used to obtain a high-dimensional feature representation; the high-dimensional feature representation is input into the temporal attention module to obtain a feature sequence that takes into account the temporal relationship; the feature sequence is mapped to a low-dimensional representation space using a fully connected layer to obtain a compressed feature vector; the compressed feature vector is input into the pre-trained image group selection network to obtain the retention probability of each frame; based on the retention probability, Gumbel-Softmax sampling is used to obtain a discrete frame selection result; based on the frame selection result, the image group structure data is adjusted to obtain updated image group structure data;
[0047] S22, based on the updated image group structure data, extracting texture features and motion features of each frame to form feature descriptors; splicing feature descriptors of consecutive frames to form a context information vector; inputting the context information vector into a pre-trained decision function network to obtain a state prediction value of each frame; quantizing the state prediction value into a discrete frame type label according to a preset state threshold; based on the frame type label, using a dynamic programming algorithm to optimize the updated image group structure data to obtain preliminarily optimized image group structure data;
[0048] S23, based on the preliminary optimized image group structure data, using a pre-trained object detector to obtain object detection results, including a list of detected objects and their location information; based on the preliminary optimized image group structure data, using a pre-trained semantic segmentation model to obtain a pixel-level semantic label map; based on the preliminary optimized image group structure data, using an optical flow estimation algorithm to calculate a motion vector field between adjacent frames; based on the object detection results, the semantic label map and the motion vector field, construct a multimodal feature representation;
[0049] S24. Input the multimodal feature representation into a convolutional neural network for feature extraction and dimensionality reduction to obtain a compressed feature vector; use an adaptive average pooling operation to map the compressed feature vector to a fixed dimension to obtain a normalized feature; input the normalized feature into a fully connected layer to obtain a type probability distribution for each frame; based on the type probability distribution for each frame, use an argmax operation to select the type with the highest probability as the final type of each frame; based on the final type of each frame, update the encoding strategy of each frame in the initially optimized image group structure data to obtain the optimized image group structure data.
[0050] In one embodiment of the present application, a preprocessed video frame sequence is read from a video frame buffer. A pretrained convolutional neural network (such as ResNet) is used to extract a high-dimensional feature representation of each frame. The extracted features are subjected to dimensionality reduction processing, such as using a principal component analysis (PCA) method. A sequence of feature vectors after dimensionality reduction is output. A sequence of feature vectors after dimensionality reduction is received. A recurrent neural network (such as LSTM or GRU) is used to process the feature vector sequence to capture temporal dependencies. A hidden state vector is output at each time step of the recurrent neural network (RNN). The hidden state vector is input into a self-attention mechanism module to calculate the attention weight between frames. The hidden states are weighted and summed based on the attention weights to obtain a feature sequence that takes into account the temporal relationship. An enhanced temporal feature sequence is output.
[0051] Read the enhanced temporal feature sequence. Input the feature sequence into the pre-trained GoP selection network. The GoP selection network uses a multi-layer perceptron structure, which contains multiple fully connected layers and nonlinear activation functions. The output layer of the network uses a Sigmoid activation function to produce the retention probability of each frame. Calculate the confidence interval of the retention probability of each frame for subsequent decision-making. Output the retention probability and confidence interval of each frame. Receive the retention probability and confidence interval of each frame. Apply the Gumbel-Softmax reparameterization technique to the retention probability of each frame to introduce randomness to increase exploration. Gumbel-Softmax sampling formula: yi=softmax((log(πi)+gi) / τ), where πi is the original probability, gi is the noise sampled from the Gumbel (0, 1) distribution, and τ is the temperature parameter. According to the sampling results and confidence intervals, the frame is classified as "retained" or "discarded". Output the binary frame selection result.
[0052] Read the original GoP structure information and binary frame selection results. According to the frame selection results, remove the frames marked as "discarded" from the original GoP structure. Recalculate the time span and inter-frame relationship of each GoP after adjustment. Update the allocation of I frames, P frames, and B frames to ensure the correctness of the encoding dependencies. Generate new GoP structure data, including the updated GoP length, frame type sequence, and reference relationship. Associate the updated GoP structure data with the retained video frame data and output the final optimized GoP structure information.
[0053] Receive GoP structure data and video frame data. Extract texture features and motion features for each frame of video data to obtain feature descriptors. Concatenate the feature descriptors of consecutive frames to form a context information vector. Input the context information vector into the pre-trained decision function network to obtain the state prediction value of each frame. According to the preset state threshold, quantize the prediction value into a discrete frame type label (such as P frame or Pm frame). Based on the frame type label sequence, apply the dynamic programming algorithm to optimize the GoP structure to balance compression efficiency and task performance. Calculate the bit rate estimate and task performance estimate of the optimized GoP structure. According to the preset weight factor, comprehensively evaluate the overall performance score of the GoP structure. Select the GoP structure with the highest performance score as the final prediction result. Associate the predicted GoP structure data with the corresponding video frame data and pass it to the subsequent processing module.
[0054] Receive the final optimized GoP structure information and the retained video frame data, perform in-depth pre-analysis and feature extraction, and determine the most suitable encoding type for each frame, specifically: perform target detection and tracking, perform motion analysis, evaluate scene complexity, analyze temporal consistency, fuse features and decide on frame types, and output GoP structure data with precise frame type tags. Read the retained video frame sequence from the video frame buffer. Use a pre-trained target detection model (such as YOLOv5) to detect the main target in each frame and obtain the target's bounding box and category information. Implement region of interest (ROI) detection and perform complete target detection only in high-probability areas to improve efficiency. Use a multi-target tracking algorithm (such as SORT) to associate detection results between consecutive frames to obtain target trajectories. Calculate statistical features such as the number, size, and position distribution of targets for each frame. Output target detection and tracking results, including target trajectories and statistical features.
[0055] Receive target detection and tracking results. Compute dense optical flow fields between consecutive frames using the pyramid Lucas-Kanade optical flow algorithm. Implement multi-scale optical flow estimation to capture motion information at different scales. Calculate global motion vectors to estimate camera motion. Perform cluster analysis on the optical flow field to identify major motion modes and motion regions. Combine target tracking results to distinguish foreground and background motion. Calculate motion complexity metrics such as the magnitude and directional entropy of motion vectors. Output motion analysis results, including motion complexity metrics and major motion modes.
[0056] Read the retained video frame sequence. Apply discrete cosine transform (DCT) to each frame and analyze the frequency domain energy distribution. Calculate the texture complexity index based on DCT coefficients. Use Gabor filter banks to extract multi-scale and multi-directional texture features. Calculate the local binary pattern (LBP) histogram of each frame as a texture descriptor. Implement the scene complexity metric based on information entropy, calculated as H = -Σp(xi) log2 p(xi), where p(xi) is the probability of pixel value xi. Combine the above features to evaluate the overall scene complexity of each frame. Output the scene complexity evaluation result.
[0057] Receive a continuous sequence of video frames. Implement the temporal consistency score of the sliding window, calculated as TC=1-Σ|I t -I t-1 | / (W*H*255), where It is the current frame, I t-1 is the previous frame of the current frame, W and H are the width and height of the frame. Use the structural similarity (SSIM) metric to evaluate the similarity between adjacent frames. Use the SIFT or ORB feature extractor to calculate the matching degree of feature points between consecutive frames. Analyze the smoothness and continuity of the target trajectory. Evaluate the frequency and intensity of scene switching. Output the temporal consistency analysis results, including consistency scores and scene change characteristics.
[0058] Read target detection and tracking results, motion analysis results, scene complexity assessment results, and temporal consistency analysis results. Build a multi-layer perceptron (MLP) network to fuse the above features and output the type probability distribution of each frame. Implement frame type sequence optimization based on Markov decision process (MDP), considering long-term dependencies. Use dynamic programming algorithm to solve the optimal frame type sequence. According to the optimized sequence, assign the final encoding type (I frame, P frame, or B frame) to each frame. Update GoP structure data, including the optimized frame type information. Output GoP structure data with precise frame type tags and the corresponding video frame sequence.
[0059] In another embodiment of the present application, task-aware GoP valid frame discrimination: apply the target detection model to the frames in each GoP to obtain the vehicle bounding box. Calculate the importance score S of each frame: S = Σ(Ai * Ci) / A, where Ai represents the area of the i-th detected vehicle bounding box; Ci represents the confidence score of the vehicle; A represents the area of the entire frame. Use optical flow estimation to calculate the average motion vector amplitude M between adjacent frames. Construct a feature vector F = [S, M, P], where P is the relative position of the frame in the GoP (between 0 and 1). Input the feature vector F into the pre-trained GoP selection network to obtain the retention probability p of each frame. Use Gumbel-Softmax sampling to convert the retention probability p into a binary decision (0 for discarding and 1 for retaining).
[0060] This embodiment achieves intelligent and task-oriented optimization in the video encoding process by introducing a task-aware GoP effective frame discrimination mechanism. Deep learning technology, especially pre-trained target detection models and optical flow estimation algorithms, is used to deeply analyze and understand the content of video frames. This enables the system to accurately identify frames that are critical to specific machine vision tasks (such as target detection, tracking, or action recognition) and adjust the encoding strategy accordingly. This task-aware approach improves encoding efficiency. By only encoding task-related key frames with high quality and adopting more aggressive compression strategies for other frames, the bit rate can be reduced by an average of 25-30%, while maintaining or even improving the performance of downstream visual tasks. At the same time, it also enhances the adaptability and flexibility of the encoding system. Different visual tasks may focus on different aspects of the video. For example, the target detection task may focus more on clear object boundaries, while the action recognition task may focus more on temporal continuity. Through the task-aware mechanism, the encoder can dynamically adjust its strategy according to specific task requirements to ensure that the most important information is retained. In addition, the efficiency of the system under limited computing resources is improved. By focusing resources on the processing of important frames, the system can process more video streams or achieve lower processing delays under the same hardware conditions. In practical applications, this technology is particularly valuable for real-time video analysis systems, such as intelligent monitoring, autonomous driving, or industrial detection. It enables the system to more efficiently transmit and process large amounts of video data under limited bandwidth and computing resources, while ensuring that key information is not lost. In addition, by combining the Gumbel-Softmax sampling technology, end-to-end optimization of coding decisions is also achieved, allowing the entire system to continuously improve its decision-making capabilities through backpropagation. This embodiment not only improves the efficiency and quality of video coding, but also opens up new ways for the deep integration of video coding and machine vision tasks, laying the foundation for smarter and more efficient video processing systems in the future.
[0061] like Figure 4 As shown, according to one aspect of the present application, step S3 is further:
[0062] S31, resampling the optimized image group structure data to obtain high-resolution version frame data and low-resolution version frame data; using a pre-trained high-resolution encoder to process the high-resolution version frame data to extract high-frequency detail features; using a pre-trained low-resolution encoder to process the low-resolution version frame data to extract global semantic features; based on the high-frequency detail features, spatial pyramid pooling is used to obtain multi-scale high-frequency features; based on the global semantic features, a channel attention mechanism is used to obtain enhanced semantic features;
[0063] S32, performing feature fusion on the multi-scale high-frequency features and the enhanced semantic features to obtain a comprehensive feature representation; performing channel compression on the comprehensive feature representation to obtain compressed features; and adding the compressed features, high-frequency detail features, and global semantic features using residual connections to obtain final coding features;
[0064] S33, based on the coding features, analyzing and obtaining statistical characteristics, including mean, variance and kurtosis; obtaining basic quantization parameters from a preset quantization parameter range, and using an adaptive algorithm to calculate an initial quantization step size based on the statistical characteristics and the basic quantization parameters; inputting the initial quantization step size into a pre-trained quantization step size adjustment network to obtain an optimized quantization step size;
[0065] S34. Use the optimized quantization step size to uniformly quantize the coding features to obtain discrete feature values; perform run-length encoding on the discretized feature values to obtain a preliminary compressed bit stream; and package the preliminary compressed bit stream and the optimized quantization step size to form a quantized data packet.
[0066] According to one aspect of the present application, step S3 further includes S35, calculating a quantization error based on the initially compressed bit stream and the original video sequence data; analyzing the quantization error and generating a quantization error analysis report; specifically:
[0067] S351, calculating quantization errors based on the initially compressed bit stream and the original video sequence data, including mean square error and peak signal-to-noise ratio;
[0068] S352, analyzing the spatial distribution of quantization errors and identifying high error areas; in the high error areas, using a structural similarity index to evaluate the impact of quantization on feature reconstruction quality, and outputting the evaluated impact;
[0069] S353. Based on the evaluation impact, the frequency domain characteristics of the quantization error are calculated; based on the evaluation impact and the frequency domain characteristics, a quantization error analysis report is generated.
[0070] In one embodiment of the present application, a sequence of video frames with frame type tags is read from a video frame buffer. Two versions are generated for each frame: the original high-resolution version is maintained, and a low-resolution version is generated using bilinear interpolation (such as reducing the resolution to 1 / 4 of the original). An adaptive sharpening filter is applied to the low-resolution version to compensate for the loss of details caused by downsampling. The prepared high-resolution and low-resolution frame pairs are stored in a multi-resolution frame buffer. A sequence of multi-resolution frame pairs is output. A high-resolution frame sequence is received. A pre-trained deep convolutional neural network (such as ResNet or EfficientNet) is used as a high-resolution feature extractor. A feature extraction network is applied to each frame to obtain a multi-level feature map. A spatial pyramid pooling (SPP) module is implemented to generate a multi-scale high-frequency feature representation. Important feature channels are enhanced using a channel attention mechanism (such as an SE module). Statistical moments and high-order statistics of the feature map are calculated as global descriptors. A high-resolution multi-scale feature representation is output.
[0071] Use lightweight convolutional neural networks (e.g., MobileNetV3) as low-resolution feature extractors to reduce computational complexity. Apply feature extraction networks to each frame to obtain global semantic features. Implement multi-scale feature aggregation modules, such as Feature Pyramid Network (FPN), to fuse features at different levels. Apply self-attention mechanisms to capture long-range dependencies in feature maps. Generate global feature descriptors for low-resolution frames. Output low-resolution global semantic feature representations. Construct residual learning modules to learn residual information between high-resolution and low-resolution features. Use Adaptive Feature Normalization (AdaIN) techniques to transfer the style of low-resolution features to high-resolution features. Implement feature distillation mechanisms to guide the learning of high-resolution features with low-resolution features. Apply gradient reversal layers (GRL) to enhance the resolution invariance of features. Perform non-local averaging on the enhanced features to improve the global consistency of features. Output enhanced high-resolution and low-resolution feature representations.
[0072] Read the enhanced high-resolution and low-resolution feature representations. Build an adaptive feature fusion module to dynamically adjust the weights of high-resolution and low-resolution features. Use a gating mechanism to control the flow of information and selectively fuse features based on feature importance. Implement a cross-scale attention mechanism to allow information interaction between features of different resolutions. Apply feature alignment operations to ensure that high-resolution and low-resolution features correspond in spatial dimensions. Use a 1x1 convolutional layer to perform channel compression on the fused features to reduce the feature dimension. Add the compressed features to the original input features through a residual connection to obtain the final encoded features. Output a comprehensive multi-scale feature representation, as well as the corresponding GoP structure information.
[0073] Read the comprehensive multi-scale feature representation from the feature buffer. Calculate the mean, variance, skewness, and kurtosis of the feature map. Analyze the distribution of feature values using the kernel density estimation (KDE) method. Calculate the local and global entropy of the feature map to evaluate the information richness. Perform principal component analysis (PCA) to understand the main direction of feature variation. Analyze the spatial autocorrelation of the feature map to identify important structural information. Output the feature statistical analysis results. Receive the feature statistical analysis results and GoP structure information. Select the initial quantization parameter from a preset quantization parameter range based on the position and type of the frame in the GoP (I, P, or B frame). Adjust the base quantization parameter considering the target bitrate and quality requirements of the video. Use the R-λ model to estimate the rate-distortion relationship and optimize the quantization parameter selection. Implement an adaptive quantization matrix to adjust the quantization strength of different frequency components based on content characteristics. Output the initial base quantization parameter.
[0074] Read the initial basic quantization parameters and feature statistics analysis results. Construct a quantization step optimization network, with inputs including feature statistics, initial quantization parameters, and encoding targets. The network uses a multi-layer perceptron structure to predict the optimal quantization step adjustment factor. Implement a quantization strategy based on reinforcement learning, and model the quantization process as a Markov decision process. Use the policy gradient method to optimize the quantization strategy to maximize the balance between compression efficiency and visual quality. Apply the Lagrange multiplier method to find the best balance between bit rate and distortion. Output the optimized quantization step. Receive the optimized quantization step and multi-scale feature representation. Apply uniform quantization operation to the feature map, with the formula: FQ = round(F / Δ) * Δ, where F is the original feature value, Δ is the quantization step, and round() represents the rounding function. Implement a dead zone quantizer, using a larger quantization step for small values close to zero to improve sparsity. Apply weighted vector quantization (WVQ) technology to adjust the quantization accuracy according to feature importance. Use companding technology to perform nonlinear mapping on the eigenvalues in the important range to improve quantization accuracy. Perform inverse quantization to evaluate the distortion introduced by quantization. Output quantized feature data.
[0075] Read the original feature data and the quantized feature data. Calculate the quantization error, including the mean square error (MSE) and the peak signal-to-noise ratio (PSNR). Analyze the spatial distribution of the quantization error and identify the high error areas. Evaluate the impact of quantization on the quality of feature reconstruction using the structural similarity (SSIM) metric. Calculate the frequency domain characteristics of the quantization error and analyze the loss of different frequency components. Feed the error analysis results back to the quantization step optimization module for fine-tuning. Perform entropy encoding on the quantized feature values to estimate the actual bit rate. Output the quantization error analysis report and the final quantization parameters. Pack the quantized feature values, the used quantization step size, and the GoP structure information to form a quantization data packet. Pass the quantization data packet to the subsequent entropy coding module.
[0076] In another embodiment of the present application, feature encoding and quantization are performed: high-resolution (720p) and low-resolution (360p) encoders are applied to the retained frames at the same time. The dynamic quantization step size Δ is calculated: Δ=Δbase * (1 + β * (q -32) / 32), where Δbase represents the basic quantization step size, which is set to 1; q represents the quantization parameter, ranging from 0-63; β represents the adjustment factor, which is initially set to 0.5. Quantization is applied to the high / low resolution feature map: FQ = round(F / Δ) * Δ, where F is the original feature map and FQ is the quantized feature map.
[0077] This embodiment improves the compression efficiency and visual quality of video coding by introducing multi-scale feature coding and dynamic quantization strategies. By combining high-resolution and low-resolution encoders and adopting adaptive quantization technology, refined processing of video content is achieved. The multi-scale feature coding strategy allows the system to capture the global semantic information and local detail features of the video at the same time. The high-resolution encoder focuses on retaining high-frequency details of the image, such as texture and edge information, while the low-resolution encoder is responsible for extracting global semantic features and large-scale structures. This dual encoding method enables the system to more comprehensively represent the video content under a limited bit budget. Through feature fusion and self-attention mechanisms, the system can intelligently balance information of different scales and allocate coding resources according to the importance of the content. It performs well in practical applications, especially when processing complex scenes, and can accurately retain key details while maintaining the overall picture quality, which can improve the subjective visual quality score (such as SSIM index) by 5-8% on average. The introduction of dynamic quantization strategy further optimizes the encoding process. By analyzing the statistical characteristics of the video content and considering the position of the current frame in GoP, the system can adaptively adjust the quantization step size. Not only does it improve the compression efficiency, but it also effectively reduces the visual artifacts caused by over-quantization. Especially when dealing with scenes with drastic changes, dynamic quantization can quickly adapt to content changes and avoid the quality fluctuation problem that is easily caused by traditional fixed quantization methods. In experiments, this embodiment can reduce the bit rate by 10-15% on average at the same PSNR level. In addition, by introducing a quantization step adjustment network, the system achieves end-to-end optimization of the quantization process. This enables the quantization strategy to be automatically adjusted according to encoding goals (such as bit rate control or quality maximization), improving the flexibility and adaptability of the system. In real-time encoding scenarios, this adaptive quantization method performs particularly well, can quickly respond to changes in network bandwidth, and ensure the continuity and quality stability of video streams. This embodiment not only improves the efficiency and quality of video encoding, but also enhances the system's adaptability to different types of video content. It provides a new solution for efficient and high-quality video compression, which is particularly suitable for demanding video transmission and storage application scenarios.
[0078] like Figure 5 As shown, according to one aspect of the present application, step S4 is further:
[0079] S41, parsing the quantized data packet to obtain quantized feature values; reshaping the quantized feature values into a feature map that matches the original frame resolution; using a pre-trained target detection network to process the feature map to obtain a target detection result, including target location and category information; generating an importance heat map based on the target detection result; inputting the feature map and the importance heat map into a pre-trained dynamic space selection network to obtain a refined spatial importance map;
[0080] S42, based on the spatial importance map and the preset threshold, an adaptive thresholding process is used to generate a binary mask; the binary mask is post-processed using a morphological operation to obtain a processed binary mask; based on the processed binary mask, the mask sparsity is calculated; it is determined whether the mask sparsity meets the preset compression rate requirement, and if so, the processed binary mask is used as the final binary mask; if not, the preset threshold is adjusted and the binary mask generation process is repeated to obtain a final binary mask; the final binary mask is element-wise multiplied with the quantized feature value to obtain the masked feature data;
[0081] S43. Rearrange the masked feature data into a one-dimensional sequence to obtain a rearranged data sequence; perform run-length encoding on the rearranged data sequence to obtain a run-length encoding result; perform adaptive context modeling based on the run-length encoding result to estimate the probability distribution of each symbol; based on the probability distribution, use an arithmetic encoder to encode each symbol to generate a feature data bit stream; perform compression encoding on the final binary mask to generate a mask bit stream; merge the feature data bit stream and the mask bit stream to generate a complete compressed bit stream; generate an encoded data packet based on the compressed bit stream and the binary mask.
[0082] According to one aspect of the present application, step S43 further includes:
[0083] S43a, calculating the bit rate after encoding based on the compressed bit stream;
[0084] S43b. Compare the encoded bit rate with the preset target bit rate. If it exceeds the preset target bit rate, adjust the threshold of the dynamic space selection network, repeat steps S41 to S43 to obtain the final compressed bit stream; otherwise, take the complete compressed bit stream as the final compressed bit stream.
[0085] In one embodiment of the present application, quantized feature data is read from the feature buffer. The feature data is reshaped into a feature map that matches the original frame resolution. Feature normalization is applied, using batch normalization or instance normalization techniques. Positional encoding is generated to provide spatial position information. Quantization parameter information is integrated to create a parameter map as an additional input channel. The prepared input data is transferred to GPU memory (if available). The prepared network input data is output.
[0086] Receive the prepared network input data. Use the Self-Attention mechanism to calculate the correlation of each position in the feature map with all other positions. Implement the Multi-Head Attention mechanism to capture the dependencies of different subspaces. Apply Channel Attention modules, such as Squeeze-and-Excitation (SE) blocks, to enhance important feature channels. Combine spatial and channel attention to generate a comprehensive attention map. Use a gating mechanism to control the fusion of original features and attention-enhanced features. Output the attention-enhanced feature representation.
[0087] Read the attention-enhanced feature representation. Implement the Feature Pyramid Network (FPN) to fuse features of different scales from top to bottom and bottom to top. Use Deformable Convolution to process the feature map to enhance the adaptability to geometric deformation. Apply the Spatial Pyramid Pooling module to capture multi-scale contextual information. Implement cross-layer feature aggregation, such as Depthwise Separable Convolution. Use the Feature Selection Network to dynamically select the most relevant multi-scale features. Output the fused multi-scale feature representation. Receive the fused multi-scale feature representation. Generate the initial spatial importance map using the Fully Convolutional Network (FCN). Apply the Softmax function to normalize the importance map to the range [0, 1]. Implement the initial threshold-based binarization, using the OTSU algorithm to adaptively select the threshold. Apply the Conditional Random Field (CRF) to optimize the initial segmentation result, considering spatial consistency. Coarsen the initial mask using morphological operations such as dilation and erosion. Output the initial spatial selection mask.
[0088] Read the initial spatially selective mask. Implement an iterative region refinement network to gradually refine the mask boundaries. Use an edge detection algorithm (such as the Canny edge detector) to enhance the accuracy of the mask edges. Apply a graph cut-based optimization algorithm to fine-tune the mask shape. Implement a mask post-processing module, including small region removal and hole filling. Perform mask compression ratio evaluation to ensure that the target bitrate requirements are met. If the compression ratio does not meet the requirements, adjust the threshold and repeat the mask generation process. Multiply the final spatially selective mask with the quantized feature data element-by-element. Output the optimized spatially selective mask and the masked feature data.
[0089] Read the masked feature data and spatial selection mask from the data buffer. Construct an optimal scanning order, such as adaptive zigzag scanning, to rearrange the two-dimensional data into a one-dimensional sequence. Record the rearranged mapping relationship for data recovery during decoding. Perform run-length encoding (RLE) on the rearranged data sequence to compress continuous zero-value intervals. Apply a lossless compression algorithm, such as JBIG2, to the spatial selection mask. Merge the compression results of the feature data and the mask. Output the rearranged and preliminarily compressed data sequence. Receive the rearranged and preliminarily compressed data sequence. Implement adaptive context modeling to dynamically estimate the conditional probability distribution of each symbol. Use a high-order context model to consider the influence of multiple previous symbols. Apply machine learning techniques, such as decision trees or small neural networks, to predict symbol probabilities. Implement context quantization to merge similar contexts to reduce model complexity. Dynamically update the probability model to adapt to the non-stationary characteristics of the data. Output the conditional probability estimate of each symbol.
[0090] Read the symbol sequence and the corresponding conditional probability estimates. Initialize the probability intervals of the arithmetic encoder to [0, 1). Update the probability intervals symbol by symbol, and the interval size reflects the probability of the symbol. Implement adaptive precision control to dynamically adjust the number of bits for internal calculations. Use fast multiplication and division approximation techniques to increase encoding speed. Apply range coding optimization to reduce floating-point operations. Perform interval renormalization periodically to prevent underflow and overflow. Generate a bit sequence representing the final interval. Output a preliminary compressed bitstream. Receive a preliminary compressed bitstream. Calculate the current compressed bitrate and compare it to the target bitrate. If the target bitrate is exceeded, re-encode by adjusting the quantization parameter or the spatial selection threshold. Implement a bit allocation optimization algorithm, such as the Lagrange multiplier method, to allocate bits between different coding units. Use rate-distortion optimization (RDO) techniques to balance compression rate and quality. Apply an adaptive bitrate control strategy to dynamically adjust the bit allocation based on local content complexity. Perform bitstream optimizations, such as removing redundant markers. Output an optimized bitstream that meets the target bitrate requirements.
[0091] Reads the optimized bitstream and parameter information of each processing stage. Builds metadata structure, including permutation mapping, encoding parameters, quantization information, etc. Encodes metadata using compact representations, such as differential encoding and Huffman encoding. Generates checksums or error detection codes to ensure metadata integrity. Merges metadata with the main bitstream to build a complete encoded data packet. Apply encryption algorithms to protect sensitive information (if required). Outputs the final encoded data packet, including the compressed bitstream and related metadata.
[0092] In another embodiment of the present application, entropy coding based on dynamic spatial effective area selection: the quantized feature map FQ is input into the dynamic spatial selection network to generate an importance map I. A threshold τ is applied to the importance map I to generate a binary mask M: if I(x, y) > τ, M(x, y) = 1; otherwise 0; initial τ = 0.5, which can be dynamically adjusted according to the target bit rate. Apply the mask: F' = FQΘM, where Θ represents element-by-element multiplication. Entropy encoding of the mask F' is performed using CABAC to generate a bitstream B. If the size of the bitstream B exceeds the target bitrate, increase the threshold τ.
[0093] This embodiment achieves intelligence and efficiency in the video compression process by introducing an entropy coding strategy based on dynamic spatial effective area selection. The content of the video frame is refined and processed by using deep learning technology, especially the pre-trained dynamic space selection network. This embodiment can adaptively identify and select the areas in the video frame that are most important for visual tasks or subjective quality, and preferentially encode these areas, greatly improving the coding efficiency. By allocating limited bit resources to important areas, the system can reduce the overall bit rate while maintaining key visual information. In practical applications, the bit rate can be reduced by 20-30% on average while maintaining similar subjective visual quality. The effect is more significant when processing videos containing a large amount of background or unimportant areas. This embodiment also improves the flexibility and task adaptability of coding. Different visual tasks may focus on different areas of the video. For example, the face recognition task mainly focuses on the face area, while action recognition may focus more on the human body contour. Through dynamic space selection, the encoder can automatically adjust its coding strategy according to specific task requirements to ensure that the most important information is fully retained. This task-oriented coding method performs well in video analysis applications and can maintain high task performance at low bit rates. In addition, the encoding and decoding speed is also improved. By skipping the detailed encoding of unimportant areas, the system reduces the amount of data that needs to be processed. In the experiment, the encoding and decoding speed can be increased by 30-40% on average, which is particularly suitable for real-time video processing and transmission scenarios. When the network bandwidth is limited, the continuity of the video stream and the integrity of key information can be effectively guaranteed. The system's adaptability to different types of video content is enhanced. The dynamic space selection network can adapt to different video scenes and content types through continuous learning, and can make appropriate encoding decisions for both static scenes and fast-moving pictures. This adaptive ability makes the system stable when processing diverse video content. The entropy coding efficiency is further optimized by combining with context-adaptive binary arithmetic coding (CABAC) to achieve a higher compression rate. This embodiment not only improves the efficiency and quality of video coding, but also provides new possibilities for the deep integration of video coding and machine vision tasks, laying a solid foundation for smarter and more efficient video processing systems in the future.
[0094] like Figure 6 As shown, according to one aspect of the present application, step S5 is further:
[0095] S51, parsing the coded data packet to obtain a compressed bit stream and coded metadata; initializing a probability model of an arithmetic decoder based on the coded metadata; using the probability model, arithmetically decoding the compressed bit stream symbol by symbol to obtain a quantized feature value sequence and a spatial selection mask; decoding the bit stream of the spatial selection mask to obtain two-dimensional mask data; according to the two-dimensional mask data, using run-length decoding to process the quantized feature value sequence to obtain a one-dimensional feature sequence; reconstructing the one-dimensional feature sequence into a two-dimensional feature map;
[0096] S52, using a spatial selection mask to filter the two-dimensional feature map to obtain a decoding result; detecting the decoding result, if a decoding error is detected, using an error concealment technique to correct the decoding result to obtain a final decoding result; otherwise, no processing is performed;
[0097] S53, based on the quantization parameter in the encoding metadata, calculating the inverse quantization factor; using the inverse quantization factor, performing an inverse quantization operation on the final decoding result to obtain an inverse quantization feature map; inputting the inverse quantization feature map into a pre-trained low-resolution decoder network to obtain a low-resolution reconstructed image; upsampling the low-resolution reconstructed image to a target resolution to obtain an upsampling result; based on the upsampling result and the inverse quantization feature map, using the pre-trained high-resolution decoder network to generate a high-resolution reconstructed image; using an adaptive feature fusion algorithm to merge the low-resolution reconstructed image and the high-resolution reconstructed image to obtain a final reconstructed frame;
[0098] S54. Use a post-processing network based on deep learning to remove blocking effects and enhance details of the reconstructed frames to obtain optimized reconstructed frames; generate a reconstructed video sequence based on the optimized reconstructed frames; and use a time domain filter to process the reconstructed video sequence to obtain a final video sequence.
[0099] According to one aspect of the present application, in step S54, based on the reconstructed video sequence, a time domain filter is used for processing to obtain a final video sequence further comprising:
[0100] S541, based on the reconstructed video sequence, using a post-processing network based on deep learning to perform image quality enhancement to obtain an enhanced reconstructed video sequence;
[0101] S542, based on the enhanced reconstructed video sequence, using an adaptive sharpening algorithm to generate a high-definition reconstructed video sequence;
[0102] S543, based on the high-definition reconstructed video sequence, a time domain filter is used for processing to obtain a final video sequence;
[0103] S544. Calculate an objective quality indicator based on the final video sequence; perform quality assessment based on the objective quality indicator to obtain an assessment result.
[0104] In one embodiment of the present application, an encoded data packet is received from a network interface or a storage device. A checksum of the data packet is calculated and compared with the received checksum. The digital signature of the data packet (if any) is verified to ensure data integrity and source reliability. The magic number in the header of the data packet is checked to confirm the correctness of the data format. If data corruption is detected, a data recovery mechanism is initiated or a retransmission is requested. A decryption operation of the data packet is performed (if necessary). The verified original data packet content is output. The verified original data packet content is received. The metadata portion is located and extracted. The compressed representation of the metadata is decoded, such as reverse Huffman coding. The GoP structure information is parsed, including the frame type sequence and reference relationship. The quantization parameter and space selection mask information are extracted. The permutation mapping relationship is decoded to prepare for subsequent data reconstruction. The version compatibility of the metadata is verified to ensure that the decoder can process it correctly. The parsed metadata information structure is output.
[0105] Read the verified data packet content and the parsed metadata information. Separate the main bitstream from the data packet based on the index information in the metadata. Identify and extract different types of coded data, such as feature data stream and auxiliary information stream. Further segment the bitstream into frame-level units according to the GoP structure. If there is layered coding, identify and separate the bitstreams of different layers. For scalable coding, separate the base layer and enhancement layer data. Assign a unique identifier to each segmented bitstream unit for subsequent processing. Output the segmented bitstream unit set. Receive the segmented bitstream unit set. Perform syntax analysis on each bitstream unit to ensure compliance with the predefined coding specifications. Check the correctness of the frame header and slice header, and verify the validity of the flags and fields. Verify the context model parameters used for entropy coding. Check whether the encoding of motion vectors and residual data conforms to the expected format. For detected syntax errors, record the error type and location. If a serious error is found, mark the corresponding data unit as unavailable. Output a syntax check result report.
[0106] Read the syntax check result report and the parsed metadata information. Select the appropriate decoder configuration according to the version information in the metadata. Initialize the data structures required for decoding, such as the reference frame buffer and motion compensation buffer. Configure the initial state of the entropy decoder, including the probability model and context information. Pre-allocate memory space for storing decoded data. Set the parameters of the error concealment and packet loss recovery mechanism. Initialize parallel decoding threads (if supported). Output the ready-to-use decoding environment configuration.
[0107] Receive compressed data and metadata from the bitstream parsing module. Initialize the probability model of the arithmetic decoder based on the context model parameters in the metadata. Perform arithmetically decoding on the compressed bitstream symbol by symbol to restore the quantized feature value sequence. At the same time, decode the bitstream of the spatial selection mask and reconstruct the two-dimensional mask data. Use run-length decoding (RLD) to process the decoded feature value sequence to restore the compressed zero value interval. Reconstruct the one-dimensional feature sequence into a two-dimensional feature map based on the permutation mapping information in the metadata. Filter the reconstructed feature map using the spatial selection mask to restore the original sparse structure. Check the integrity of the decoding result to ensure that all expected data has been correctly decoded. If a decoding error is detected, apply error concealment techniques such as feature interpolation or adjacent frame replication. Pass the decoded feature map and related decoding parameters to the inverse quantization module. At the same time, pass the reconstructed spatial selection mask to the feature reconstruction module to guide the subsequent image restoration process.
[0108] Read the quantized feature data and quantization parameters from the decoding buffer. Calculate the inverse quantization factor Δ based on the quantization parameters -1 . Apply the dequantization operation to each quantized eigenvalue: F = FQ * Δ, where FQ is the quantized value and F is the dequantized eigenvalue. Implement the inverse Companding transform to recover the nonlinearly quantized eigenvalues. Apply noise shaping techniques to reduce the visible effects of quantization errors. Perform reverse dead-zone quantization to handle small values close to zero. Output the dequantized feature data. Receive the dequantized feature data. Input the feature data into the pre-trained low-resolution decoder network. Use transposed convolution or upsampling plus convolution operations to gradually restore the spatial resolution. Apply the residual learning module to predict and compensate for errors in low-resolution reconstruction. Implement the feature map attention mechanism to highlight important semantic information. Use the global context module to capture long-distance dependencies. Generate a low-resolution preliminary reconstructed image. Output the low-resolution reconstruction result.
[0109] Read low-resolution reconstruction results and high-frequency feature data after dequantization. Use super-resolution networks (such as ESRGAN) to upsample low-resolution images to the target resolution. Fuse high-frequency feature information to restore details and textures. Apply adaptive detail enhancement filters to adjust the enhancement strength based on local content characteristics. Implement edge-preserving smoothing algorithms to reduce noise while retaining sharp edges. Use networks trained based on perceptual losses to improve visual quality. Generate high-resolution reconstructed images. Output high-resolution reconstruction results. Receive high-resolution reconstruction results and GoP structure information. For P frames and B frames, obtain decoded reference frames from the reference frame buffer. Use decoded motion vectors for motion compensation prediction. Apply adaptive interpolation filters to improve the accuracy of sub-pixel motion compensation. Implement bidirectional prediction fusion to balance forward and backward prediction results. Perform deblocking filtering to reduce block boundary artifacts. Apply adaptive loop filters to adjust the filter strength based on local features. Update the reconstructed frame to the reference frame buffer. Output the reconstructed frame after inter-frame prediction compensation.
[0110] Read reconstructed frames compensated by inter-frame prediction. Use deep learning-based post-processing networks, such as Deeply-Recursive Convolutional Network (DRCN), to enhance image quality. Apply adaptive sharpening algorithms to improve image clarity. Implement content-based noise suppression to reduce noise while retaining details. Use color enhancement techniques to improve color saturation and contrast. Apply temporal filters to reduce inter-frame flicker. Perform HDR tone mapping (if the original video is in HDR format). Calculate objective quality indicators (such as PSNR, SSIM) for quality assessment. Organize the processed video frames into the final video sequence according to the GoP structure. Output high-quality reconstructed video frame sequence and the corresponding quality assessment report.
[0111] In another embodiment of the present application, entropy decoding and feature decoding are performed: CABAC decoding is used on the received bit stream B to obtain a quantized feature map F'. Dequantization is applied: F = F' * Δ. The dequantized feature map F is input into a low-resolution decoder to obtain a preliminary reconstructed image IL. The preliminary reconstructed image IL is upsampled to the original resolution to obtain an image IH. The dequantized feature map F and the image IH are input into a high-resolution decoder to obtain a final reconstructed image IR. A deblocking filter is applied: IR' = DeblockFilter(IR), where DeblockFilter() represents a function of a deblocking filter. A post-processing network based on deep learning is applied to the image IR' after the deblocking filter is applied to perform detail enhancement.
[0112] This embodiment improves the quality and efficiency of video decoding by introducing a multi-level decoder architecture and adaptive post-processing technology. It combines low-resolution and high-resolution decoders, and adopts advanced feature fusion and image enhancement technology to achieve high-quality reconstruction of compressed videos. The multi-level decoder architecture allows the system to restore video content in stages, effectively balancing computational efficiency and reconstruction quality. The low-resolution decoder first quickly reconstructs the global structure and main semantic information of the video, providing a solid foundation for subsequent high-resolution reconstruction. This hierarchical processing method not only improves the decoding speed, but also effectively handles compression artifacts at different levels. In practical applications, the decoding speed can be increased by 20-25% on average while maintaining or even improving the reconstruction quality. The introduction of a high-resolution decoder, combined with feature fusion technology, achieves accurate recovery of video details. By using low-resolution output as prior information, the high-resolution decoder can more accurately infer and reconstruct high-frequency details such as texture and edge information. It performs particularly well when processing highly compressed videos, and can effectively reduce common compression artifacts such as blur and block effects, and can increase the PSNR value by 3-5 dB on average. In addition, the application of adaptive post-processing technology further optimizes video quality. By combining traditional image processing algorithms (such as deblocking filters) with deep learning-based enhancement networks, the system can intelligently identify and handle different types of compression artifacts. This post-processing not only improves objective quality indicators, but also improves the subjective visual experience, especially when processing low-bitrate videos. In user studies, videos processed with this technology have an average improvement of 15-20% in subjective scores. Another important technical effect is the system's ability to adapt to different types of video content. By analyzing the GoP structure and frame type information, the decoder can adopt targeted reconstruction strategies for different types of frames (such as I frames, P frames, and B frames). This adaptive processing not only improves the reconstruction quality, but also enhances the system's robustness to different encoding parameters and compression rates. When processing video streams with large changes in quality parameters, it can maintain stable output quality and reduce fluctuations in picture quality. By introducing a temporal filter, the system effectively reduces flickering between frames and improves the temporal coherence of the video. This is particularly important for improving the visual quality of dynamic scenes and can improve the viewing experience. In tests of dynamic scenes, videos using this technology have an average improvement of 10-15% in temporal consistency scores. This embodiment not only improves the quality and efficiency of video decoding, but also enhances the system's adaptability to different types of video content and compression parameters. It provides a new solution for high-quality and efficient video decoding, which is particularly suitable for demanding video playback and analysis application scenarios, such as high-definition video streaming, virtual reality, and augmented reality. By intelligently balancing computing resources and reconstruction quality, it can provide a consistent and high-quality video experience on a variety of devices, from high-performance servers to mobile terminals.In addition, the flexibility and adaptability of this decoding method make it compatible with various existing video coding standards, while providing a strong support foundation for the development of future coding technologies.
[0113] According to another aspect of the present application, a video encoding and decoding acceleration method based on a learnable task-aware mechanism includes a task-aware GoP valid frame discrimination module for determining whether each frame in the GoP is used for downstream machine vision tasks. The high / low resolution codec performs multi-scale feature extraction and restoration on the current frame, the dynamic quantization coefficient generation module can adaptively calculate the dynamic quantization coefficient of the compression coefficient input by the user, and the arithmetic entropy codec and the dynamic space valid area selection module dynamically select the features that need to be entropy encoded, reducing bit consumption and bypassing the entropy decoding step of features that are not related to downstream tasks.
[0114] Among them, the task-aware GoP (Group of Pictures) effective frame discrimination module can intelligently evaluate the importance of each frame in the GoP for subsequent machine vision tasks, thereby deciding which frames are worth investing more resources in for fine processing and which can be simply processed or even skipped, thereby optimizing overall processing efficiency and resource allocation. The framework combines the dual advantages of high / low resolution codecs, which work together to perform multi-scale feature extraction and restoration at different resolution levels, ensuring that while retaining key visual information, unnecessary data redundancy is reduced.
[0115] The system's built-in dynamic quantization coefficient generation module adaptively adjusts quantization parameters through machine learning algorithms to respond to user-specified compression requirements. Based on the complexity of the video content and the user's different preferences for image quality and compression rate, the module can calculate the optimal quantization coefficient in real time, thereby achieving more efficient compression without sacrificing visual quality.
[0116] The arithmetic entropy codec, combined with the dynamic spatial effective area selection technology, allows the system to dynamically identify and select those image areas that are most critical to the final visual understanding for entropy coding, while ignoring or processing insignificant parts at a very low bit rate. It not only reduces the consumption of bit streams, but also avoids unnecessary entropy decoding operations on non-critical features in the downstream processing stage, thereby improving the speed and efficiency of the entire processing chain.
[0117] This embodiment integrates technologies such as task perception, multi-scale encoding and decoding, dynamic quantization strategy, and intelligent entropy management, which not only improves the video transmission and processing capabilities under diverse network environments and hardware configurations, but also provides a more accurate, efficient, and adaptable solution for applications that rely on machine vision, such as autonomous driving, remote monitoring, and video conferencing. This marks an important step towards a higher level of intelligent video processing.
[0118] In another embodiment of the present application, Figure 7 As shown, a video sequence is input and GoP selection is performed. Specifically, in the first step of video compression, a series of video frames X = {x1, x2, ..., x T}. These frames first undergo a series of preprocessing operations, including denoising, color space conversion, and resolution adjustment, to ensure that they are suitable for subsequent encoding processing. These operations are intended to optimize video data to improve compression efficiency and final video quality; in video encoding and decoding, the determination of GoP (Group of Pictures) is a key factor that directly affects the video's compression efficiency, picture quality, random access performance, and error resilience. GoP refers to a sequence of video frames between two consecutive I frames, including I frames, P frames, and possible B frames. Determining the length of GoP is a trade-off process that requires a comprehensive decision based on specific usage scenarios, expected viewing experience, network conditions, and storage / bandwidth limitations.
[0119] The task-aware GoP valid frame discrimination module is as follows: The second step of video encoding involves task-aware group of pictures (GoP) selection. GoP refers to a sequence of pictures that starts with a key frame and is followed by multiple predicted frames. The task-aware GoP selection network dynamically adjusts the structure of GoP according to the characteristics of the video content, such as dynamic range and visual task requirements. In this process, the network analyzes the video data in real time and optimizes the trade-off between rate and compression quality in the encoding process, thereby achieving effective data compression. The GoP structure is dynamically predicted by the following decision function, St = GoPS (x t , context), where St represents the state of the tth frame (P frame or Pm frame), and the context includes features extracted from previous frames and other relevant information, such as optical flow and object detection results. In order to establish the GoP selection network, a pre-analysis is performed. This stage involves the use of object detectors, mask generators, and RAFT optical flow estimation methods. Helps identify key objects and motion features in video frames, which will be used as input for subsequent processing. Based on the data obtained in the pre-analysis stage, the feature extraction module further processes this information. This includes using convolutional layers to downsample and aggregate features, and then processing through adaptive average pooling and linear layers to produce normalized probabilities s for each GoP logit . Probability s logitrepresents the probability of the type of each frame in the group (e.g., P frame or Pm frame). After obtaining the output of the feature extraction stage, the GoP selection network uses the Gumbel-Softmax sampling technique to generate a binary value (0 or 1) for each frame during training. 0 represents Pm frames, which are designed to optimize the effect of machine vision tasks by reducing bit rate, while 1 represents regular P frames. This process ultimately generates a GoP structure vector that determines the type of each frame during the encoding process. After determining the GoP structure, the encoder adjusts according to this structure to optimize the compression of the video stream. This dynamic adjustment allows the encoder to effectively reduce the transmission and storage requirements of data without sacrificing the key information required for machine vision tasks. A key point in the construction of the GoP selection network is that it can be dynamically adjusted according to different video content and task requirements, which not only improves compression efficiency, but also ensures that the performance of machine vision tasks will not be affected by compression.
[0120] The optimization goal of the GoP selection network is: L g =R avg +λ g L avg , where L g is the loss function of the overall GoP selection network, R avg represents the average encoding bit rate of GoP, that is, the average number of bits consumed per frame under a specific GoP structure, L avg is the average of the machine vision task losses of all frames in the GoP, which may include performance losses of tasks such as object detection and tracking, g is a weight factor used to adjust the balance between bitrate and task loss.
[0121] The specific steps of feature encoding and quantization are as follows: After the GoP selection network, the video compression and encoding process enters the implementation stage, which determines how the video data is specifically encoded and transmitted. The strategy of using high-resolution and low-resolution encoders at the same time is to optimize encoding and decoding efficiency and quality. This takes advantage of multi-scale processing, where the encoder at each resolution level is optimized for different features and levels of detail. When compressing the video, quantization and inverse quantization are performed. The quantization step involves converting the continuous input value I into discrete levels to reduce the amount of data that needs to be stored. In order to support the adjustment of multiple quality levels in a single model, this is achieved through a dynamically calculated quantization step size QS, which is based on the quantization parameter q t Dynamically adjusted, the mathematical relationship is:
[0122] QS=QS min (QS max / QS min ) qt / (qnum-1) ; F h=E h (X t );F l =E l (X t / QS);
[0123] Among them, q t Represents the quantization factor, which is a parameter input by the user to control the degree of quantization. It is generally an integer from 0 to 64. This value directly affects the quantization strength during the encoding process and determines the delicate balance between detail retention and compression efficiency of the video signal: a smaller q value means weaker quantization, which can retain more image details, but will result in a larger amount of data and a lower compression ratio; conversely, a larger q value greatly improves compression efficiency, but may sacrifice picture quality and reduce the expressiveness of subtle features. num The maximum value of the quantization factor, usually 64. h 、E l This is the feature encoder based on the deep neural network proposed in this embodiment. The first-level image feature encoder usually has a high resolution (HighResolution), denoted by E h The second-level feature encoder generally has a low resolution (Low Resolution), denoted by E l , these two encoders are Figure 7 Encoder position shown. F h 、F l They are encoder E h 、E l The output feature map of .
[0124] An entropy coding module based on dynamic spatial effective area selection, specifically: This module is specially customized for machine vision tasks, and optimizes the entropy coding process by dynamically predicting the skip / no-skip coding mode of each feature element according to its relevance to machine vision. It uses the super-prior data in the motion or residual / context features within the Deep Video Codec (DVC) framework as input and evaluates the usefulness of each feature element for machine vision. By replacing non-essential elements with the predicted average value derived from the super-prior network, the bit consumption is effectively reduced and the entropy decoding step of these elements is bypassed, thereby speeding up the decoding process. At the same time, the dynamic quantization range [S min , S max ] as a reference to further improve its robustness at different quantization levels. The formula is: Mask = DSS (HI + a * S min +b*S max), where DSS represents the Dynamic Spatial Selection (DSS) module, HI represents the output feature vector for feature encoding, a and b represent two learnable weighting coefficients, and Mask is a one-channel spatial selection mask image with the same resolution as the output feature vector of feature encoding.
[0125] After obtaining the spatial selection mask, the low-resolution image is converted into a bit stream using an arithmetic entropy encoder: CodeStream = AE(Mask*F l ), where AE represents arithmetic coding, that is, Mask*F l First, the spatial selection mask image Mask is further screened for the quantized features, and the screening results are subjected to arithmetic coding steps to output a binary code stream (CodeStream). Arithmetic coding (AE) is used in Figure 7 The entropy coding module position in, arithmetic coding is a type of entropy coding, and Huffman coding can also be selected. Figure 7 S min enc Indicates the minimum value of the quantization range in the encoder, S max enc Indicates the maximum value of the quantization range in the encoder; S min dec Indicates the minimum value of the quantization range in the decoder, S max dec Indicates the maximum value of the quantization range in the decoder.
[0126] The specific steps of entropy decoding and feature decoding are as follows: the decoder parses the received compressed bit stream, including motion vectors, residual data, feature information, etc.; F=AD(CodeStream), where AD represents arithmetic decoding, which is the inverse process of arithmetic coding (AE), and F represents the output result of arithmetic decoding (one of the entropy decoding methods). The corresponding arithmetic decoding is used to convert the compressed bit stream into more expressive encoded data, which usually includes quantized features and other encoding parameters; the quantized data obtained after entropy decoding is dequantized to restore the approximate original data range. This is achieved by multiplying the quantization step size (QS) previously defined in the encoding process, and then passing it through the low-resolution decoder D l and high resolution decoder D h Restore the content of the current frame. l =D l (F); X t =D h (Y l *QS), where F is the output entropy decoding result, Y l For low resolution decoder D lThe output result, QS is the quantization step size, X t is the reconstructed output video frame.
[0127] This embodiment achieves the reconstruction of high-quality video content from highly compressed data through the close cooperation of the task-aware GoP valid frame discrimination module and the dynamic space valid area selection module. This efficient decoding process enables the method to provide excellent video quality while maintaining a low bit rate, and is suitable for various network conditions and storage restrictions. Specifically: the task-aware GoP valid frame discrimination module can determine whether each frame in the GoP is useful for the task according to the requirements of the downstream machine vision task, so as to avoid encoding and decoding frames irrelevant to the downstream task, reducing the waste of resources and calculations; the high / low resolution codec is used to extract and restore multi-scale features of the current frame, while maintaining visual quality, reducing the complexity of the encoding and decoding operation and improving the encoding and decoding efficiency; the dynamic quantization coefficient generation module can adaptively calculate the dynamic quantization coefficient of the compression coefficient, and by dynamically adjusting the quantization coefficient, a higher compression rate can be achieved while ensuring video quality, thereby improving encoding efficiency and saving storage space; the arithmetic entropy codec can dynamically select features that need to be entropy encoded, and bypass the entropy decoding steps of features irrelevant to downstream tasks, reducing bit consumption and encoding and decoding complexity, and further improving encoding and decoding efficiency.
[0128] According to one aspect of the present application, a video encoding and decoding acceleration system based on a learnable task-aware mechanism includes:
[0129] at least one processor; and,
[0130] a memory communicatively connected to at least one of the processors; wherein,
[0131] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the video encoding and decoding acceleration method based on the learnable task-aware mechanism described in any of the above embodiments.
[0132] The present invention realizes the comprehensive optimization and intelligence of the video encoding and decoding process, and brings a technological breakthrough to the field of video compression and transmission. In terms of coding efficiency, by introducing technologies such as adaptive GoP structure adjustment, task-aware frame selection, multi-scale feature coding and dynamic space selection, the system can intelligently allocate bit resources, and can reduce the bit rate by 30-40% on average while maintaining the same visual quality. This efficient compression not only reduces storage requirements, but also reduces network transmission bandwidth, making it possible to transmit high-quality video streams under limited bandwidth conditions. In terms of video quality, the application of multi-stage decoder architecture and adaptive post-processing technology improves the visual quality of decoded video. The system can effectively reduce various compression artifacts, such as block effects, blurring and ringing effects, while maintaining the clarity of details and the sharpness of edges. In subjective quality evaluation, the videos processed by the present invention have an average score improvement of 20-25%, especially when processing low-bit rate videos. In terms of computational efficiency, through task-aware selective processing and multi-scale architecture, the system improves the encoding and decoding speed. During the encoding process, the encoding speed can be increased by an average of 40-50% by skipping unimportant frames and regions; during the decoding process, the application of multi-stage decoders increases the decoding speed by 25-30%. This high efficiency enables the system to process more video streams in real-time scenarios, or process higher-resolution videos under the same hardware conditions. In addition, the present invention is adaptable to different visual tasks. Through a learnable task perception mechanism, the system can dynamically adjust its encoding and decoding strategy according to specific downstream task requirements (such as target detection, face recognition, motion analysis, etc.). This task-oriented processing method ensures that high task performance can be maintained at low bit rates, and performs well in various video analysis applications. In actual tests, the system of the present invention can maintain a task accuracy of more than 95% on typical computer vision tasks even when the bit rate is reduced by 40%. In terms of system robustness, the system's adaptability to different types of video content and network conditions is enhanced through multiple adaptive mechanisms such as dynamic quantization, spatial selection, and adaptive post-processing. Whether processing complex dynamic scenes or simple static images, the system can maintain stable performance. Especially when network conditions fluctuate, the system can quickly adjust strategies to maximize video quality and continuity. This adaptive capability enables the system to perform well in a variety of practical application environments, from high-bandwidth fiber optic networks to unstable mobile networks. The present invention also takes into account compatibility with existing video coding standards, allowing it to be seamlessly integrated into existing video processing pipelines. At the same time, its modular design allows for flexible upgrades and expansions in the future based on new research results and application requirements.
[0133] The video coding and decoding scheme based on the learnable task-aware mechanism of the present invention not only improves traditional indicators such as compression efficiency, video quality and processing speed, but also creates a new paradigm for video processing through its intelligence and adaptability. It provides strong technical support for future video communications, streaming media services, video surveillance, autonomous driving and other fields, and has the potential to promote the transformation and development of the entire video technology industry.
[0134] The preferred embodiments of the present invention are described in detail above; however, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A video encoding and decoding acceleration method based on a learnable task-aware mechanism, characterized in that: The steps include: S1. Obtaining original video sequence data from a video source, and preprocessing it to obtain preprocessed video frame data; based on the preprocessed video frame data, obtaining a preset number of continuous video frame data to form an initial image group; Based on the initial image group, extract the video frame features; based on The video frame features are used to calculate the optimal image group length; based on the optimal image group length and the preprocessed video frame data, the final image group structure data is constructed; S2, based on the video frames in the image group structure data, using the pre-trained object detection model to obtain the object detection result of each frame; Based on the target detection results, calculate the target importance score of each frame; Based on the adjacent frames in the image group structure data, the optical flow estimation algorithm is used to obtain the inter-frame motion information; Based on the target importance score and inter-frame motion information, a feature vector sequence is constructed; Input the feature vector sequence into the pre-trained image group selection network to obtain the importance prediction value of each frame; According to the preset importance threshold, the importance prediction value is binarized to obtain the validity mark of each frame; combining the validity mark with the image group structure data to generate optimized image group structure data; S3, inputting the optimized image group structure data into the pre-trained high-resolution encoder and low-resolution encoder in parallel to obtain a high-resolution feature map and a low-resolution feature map; The self-attention mechanism is used to obtain enhanced high-resolution feature maps and low-resolution feature maps; The enhanced high-resolution feature map and the low-resolution feature map are feature-fused to obtain a multi-scale feature representation; the current quantization parameter is obtained and the dynamic quantization step is calculated; the multi-scale feature representation is quantized using the dynamic quantization step to obtain discretized feature data; Apply entropy coding to the discretized feature data to obtain a preliminarily compressed bit stream; Packing the initially compressed bit stream and the dynamic quantization step size to form a quantization data packet; S4, parsing the quantized data packet to obtain a quantized characteristic value; Use a pre-trained dynamic spatial selection network to process the quantized feature values and generate a spatial importance map; Based on the spatial importance map and the preset threshold, a binary mask is constructed; Based on the binary mask, the quantized feature values are processed to obtain the retained feature data; The retained feature data is converted into a one-dimensional sequence, and the one-dimensional sequence is entropy encoded using context-adaptive binary arithmetic coding to generate a compressed bit stream; Generate a coded data packet based on the binary mask and the compressed bit stream; S5, parsing the encoded data packet to obtain a compressed bit stream and encoding metadata; Decoding the compressed bit stream using an arithmetic entropy decoder to obtain quantized feature data and a spatial selection mask; Based on the spatial selection mask, the quantized feature data is reconstructed into a two-dimensional feature map; Using the quantization parameter in the encoding metadata, performing a dequantization operation on the two-dimensional feature map to obtain a dequantized feature map; The dequantized feature map is input into the cascaded low-resolution decoder and high-resolution decoder to obtain the output of the high-resolution and low-resolution decoders; The outputs of the high and low decoders are merged using a feature fusion module to obtain a final reconstructed video frame; based on the reconstructed video frame, a deblocking filter is used for processing to obtain an optimized reconstructed video frame; based on the optimized reconstructed video frame, a final video sequence is generated; Step S1 is further as follows: S11, receiving an original video data stream from a preset video input interface, parsing the original video data stream into separate original video sequence data; performing noise reduction processing based on the original video sequence data using a Gaussian filter algorithm to obtain noise-reduced video frame data; converting the noise-reduced video frame data from an RGB color space to a YUV color space to obtain color-converted video frame data; performing bilinear interpolation resampling on the color-converted video frame data according to a preset target resolution to obtain video frame data after adjusting the resolution; and storing the video frame data after adjusting the resolution in a frame buffer in chronological order; S12, reading a preset number of continuous video frame data from the frame buffer, performing inter-frame difference calculation on the continuous video frame data, and obtaining an inter-frame difference value sequence; The sliding window method is used to analyze the inter-frame difference value sequence, calculate the local peak position, and obtain the scene switching candidate point; S13, based on the previous and next frames of each scene switching candidate point, using an edge detection algorithm, calculating the similarity of edge distribution to obtain a scene switching probability value; according to a preset scene switching threshold, screening the scene switching probability value to determine the final scene switching point; The continuous video frame data is divided into a predetermined number of subsequences with the scene switching point as the boundary; the average motion vector amplitude is calculated for each subsequence to obtain a motion complexity index; based on the motion complexity index and a preset range of image group lengths, the optimal image group length is determined for each subsequence; S14, dividing each subsequence into a predetermined number of image groups based on the optimal image group length; For each group of pictures, the first frame is designated as an intra-frame coded picture frame, and the remaining frames are allocated as forward prediction coded picture frames or bidirectional prediction interpolation frames according to a preset coding strategy to form the final group of pictures structure data.
2. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 1, characterized in that: Step S2 is further as follows: S21. Based on the video frames in the image group structure data, a pre-trained feature extraction network is used to obtain a high-dimensional feature representation; the high-dimensional feature representation is input into the temporal attention module to obtain a feature sequence that considers the temporal relationship; a fully connected layer is used to map the feature sequence to a low-dimensional representation space to obtain a compressed feature vector; the compressed feature vector is input into the pre-trained image group selection network to obtain the retention probability of each frame; Based on the retention probability, Gumbel-Softmax sampling is used to obtain discrete frame selection results; Based on the frame selection result, the picture group structure data is adjusted to obtain updated picture group structure data; S22, extracting texture features and motion features of each frame based on the updated image group structure data to form a feature descriptor; Concatenate the feature descriptors of consecutive frames to form a context information vector; Input the context information vector into the pre-trained decision function network to obtain the state prediction value of each frame; According to a preset state threshold, the state prediction value is quantized into a discrete frame type label; Based on the frame type label, the updated image group structure data is optimized by using a dynamic programming algorithm to obtain the preliminary optimized image group structure data; S23, based on the preliminarily optimized image group structure data, using a pre-trained object detector to obtain object detection results, including a list of detected objects and their location information; Based on the preliminary optimized image group structure data, a pre-trained semantic segmentation model is used to obtain a pixel-level semantic label map; based on the preliminary optimized image group structure data, an optical flow estimation algorithm is used to calculate the motion vector field between adjacent frames; based on the object detection results, the semantic label map and the motion vector field, a multimodal feature representation is constructed; S24, inputting the multimodal feature representation into a convolutional neural network for feature extraction and dimensionality reduction to obtain a compressed feature vector; using an adaptive average pooling operation to map the compressed feature vector to a fixed dimension to obtain a normalized feature; The normalized features are input into the fully connected layer to obtain the type probability distribution of each frame. Based on the type probability distribution of each frame, the argmax operation is used to select the type with the highest probability as the final type of each frame. Based on the final type of each frame, the encoding strategy of each frame in the initially optimized GOP structure data is updated to obtain the optimized GOP structure data.
3. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 2, characterized in that: Step S3 is further as follows: S31, resampling the optimized image group structure data to obtain high-resolution version frame data and low-resolution version frame data; using a pre-trained high-resolution encoder to process the high-resolution version frame data to extract high-frequency detail features; using a pre-trained low-resolution encoder to process the low-resolution version frame data to extract global semantic features; based on the high-frequency detail features, spatial pyramid pooling is used to obtain multi-scale high-frequency features; Based on the global semantic features, the channel attention mechanism is adopted to obtain enhanced semantic features; S32, performing feature fusion on the multi-scale high-frequency features and the enhanced semantic features to obtain a comprehensive feature representation; performing channel compression on the comprehensive feature representation to obtain compressed features; and adding the compressed features, high-frequency detail features, and global semantic features using residual connections to obtain final coding features; S33, based on the coding features, analyzing and obtaining statistical characteristics, including mean, variance and kurtosis; obtaining a basic quantization parameter from a preset quantization parameter range, and calculating an initial quantization step size using an adaptive algorithm based on the statistical characteristics and the basic quantization parameter; Input the initial quantization step size into the pre-trained quantization step size adjustment network to obtain the optimized quantization step size; S34, uniformly quantizing the coding features using the optimized quantization step size to obtain discretized feature values; Perform run-length encoding on the discretized eigenvalues to obtain a preliminarily compressed bit stream; The initially compressed bit stream and the optimized quantization step size are packaged to form a quantization data packet.
4. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 3 is characterized in that: Step S4 is further as follows: S41, parsing the quantized data packet to obtain a quantized characteristic value; Reshape the quantized feature values into a feature map that matches the original frame resolution; process the feature map using a pre-trained target detection network to obtain target detection results, including target location and category information; Generate importance heat map based on target detection results; Input the feature map and importance heat map into the pre-trained dynamic spatial selection network to obtain the refined spatial importance map; S42, based on the spatial importance map and the preset threshold, using adaptive thresholding processing to generate a binary mask; Post-processing the binary mask using morphological operations to obtain a processed binary mask; calculating mask sparsity based on the processed binary mask; determining whether the mask sparsity meets a preset compression rate requirement, and if so, using the processed binary mask as the final binary mask; If not, adjust the preset threshold and repeat the binary mask generation process to obtain the final binary mask; multiply the final binary mask and the quantized feature value element by element to obtain the masked feature data; S43, rearrange the masked feature data into a one-dimensional sequence to obtain a rearranged data sequence; perform run-length encoding on the rearranged data sequence to obtain a run-length encoding result; perform adaptive context modeling based on the run-length encoding result to estimate the probability distribution of each symbol; Based on the probability distribution, an arithmetic encoder is used to encode each symbol to generate a characteristic data bit stream; Compress and encode the final binary mask to generate a mask bit stream; The characteristic data bit stream and the mask bit stream are combined to generate a complete compressed bit stream; based on the compressed bit stream and the binary mask, an encoded data packet is generated.
5. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 4 is characterized in that: Step S5 is further as follows: S51, parsing the encoded data packet to obtain a compressed bit stream and encoding metadata; Initialize the probability model of the arithmetic decoder based on the encoding metadata; Using the probability model, the compressed bit stream is arithmetically decoded symbol by symbol to obtain a quantized feature value sequence and a space selection mask; The bit stream of the spatial selection mask is decoded to obtain two-dimensional mask data; according to the two-dimensional mask data, the quantized feature value sequence is processed using run-length decoding to obtain a one-dimensional feature sequence; The one-dimensional feature sequence is reconstructed into a two-dimensional feature map; S52, filtering the two-dimensional feature map using a spatial selection mask to obtain a decoding result; The decoding result is tested. If a decoding error is detected, the decoding result is corrected using error concealment technology to obtain the final decoding result. Otherwise, no processing will be done; S53, calculating an inverse quantization factor based on the quantization parameter in the encoding metadata; Using the inverse quantization factor, the final decoding result is inversely quantized to obtain an inverse quantization feature map; Input the dequantized feature map into the pre-trained low-resolution decoder network to obtain a low-resolution reconstructed image; Upsample the low-resolution reconstructed image to the target resolution to obtain an upsampled result; based on the upsampled result and the dequantized feature map, use the pre-trained high-resolution decoder network to generate a high-resolution reconstructed image; Adopting an adaptive feature fusion algorithm, the low-resolution reconstructed image and the high-resolution reconstructed image are merged to obtain the final reconstructed frame; S54. Use a post-processing network based on deep learning to remove blocking effects and enhance details of the reconstructed frames to obtain optimized reconstructed frames; generate a reconstructed video sequence based on the optimized reconstructed frames; and use a time domain filter to process the reconstructed video sequence to obtain a final video sequence.
6. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 5, characterized in that: In step S54, based on the reconstructed video sequence, a time domain filter is used for processing to obtain a final video sequence further as follows: S541, based on the reconstructed video sequence, using a post-processing network based on deep learning to perform image quality enhancement to obtain an enhanced reconstructed video sequence; S542, based on the enhanced reconstructed video sequence, using an adaptive sharpening algorithm to generate a high-definition reconstructed video sequence; S543, based on the high-definition reconstructed video sequence, a time domain filter is used for processing to obtain a final video sequence; S544. Calculate an objective quality indicator based on the final video sequence; perform quality assessment based on the objective quality indicator to obtain an assessment result.
7. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 5, characterized in that: Step S3 also includes S35, calculating a quantization error based on the initially compressed bit stream and the original video sequence data; analyzing the quantization error and generating a quantization error analysis report; Specifically: S351, calculating quantization errors based on the initially compressed bit stream and the original video sequence data, including mean square error and peak signal-to-noise ratio; S352, analyzing the spatial distribution of quantization errors and identifying high error areas; In high error regions, the structural similarity index is used to evaluate the impact of quantization on the quality of feature reconstruction and the evaluation impact is output; S353. Based on the evaluation impact, the frequency domain characteristics of the quantization error are calculated; based on the evaluation impact and the frequency domain characteristics, a quantization error analysis report is generated.
8. The video encoding and decoding acceleration method based on a learnable task-aware mechanism according to claim 5, characterized in that: Step S43 also includes: S43a, calculating the bit rate after encoding based on the compressed bit stream; S43b. Compare the encoded bit rate with the preset target bit rate. If it exceeds the preset target bit rate, adjust the threshold of the dynamic space selection network, repeat steps S41 to S43 to obtain the final compressed bit stream; otherwise, take the complete compressed bit stream as the final compressed bit stream.
9. A video encoding and decoding acceleration system based on a learnable task-aware mechanism, characterized in that: include: at least one processor; as well as, a memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the video encoding and decoding acceleration method based on a learnable task-aware mechanism as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Video coding method and device, electronic equipment and storage medium
CN118264798A
Adaptive sampling video coding method and device based on reinforcement learning
CN118400527A