Semantic perception video compression method and system for man-machine mixed vision
By extracting dynamic semantics and generating finely reconstructed videos through the basic and auxiliary branches of a pre-trained video compression network, the problems of resource consumption and lack of versatility in existing technologies are solved, and video compression with high machine vision accuracy and good rate-distortion performance at low bitrates is achieved.
Patent Information
- Application Number
- CN202511067894.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing video compression methods, while meeting the needs of human-machine hybrid vision tasks, consume excessive computing resources and time, lack versatility, and struggle to maintain high accuracy in machine vision tasks at low bitrates while avoiding wasting bit resources.
A pre-trained video compression network, including a basic branch and an auxiliary branch, is used to generate regions of interest by extracting dynamic semantics, generate visually consistent focus frames using a neural rendering model, entropy model and conditional decoder to encode and decode feature probability distributions, and combine cross-attention mechanism to generate finely reconstructed videos.
It effectively preserves key information under low bitrate conditions, improves the accuracy of machine vision tasks, reduces the waste of redundant information, achieves better rate-distortion performance, and meets the needs of human-machine vision.
Smart Images

Figure CN120915949A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video redundancy information removal, in particular to a semantic perception video compression method and system for human-computer hybrid vision. BACKGROUND
[0002] With the explosive growth of video content, there are great challenges in storing and transmitting video data. Traditional video coding standards (such as H.264 / AVC, H.265 / HEVC, H.266 / VVC) mainly optimize signal fidelity for human vision, while deep learning-based video compression methods maintain signal fidelity while reducing bit rate. However, with the wide application of machine vision in intelligent transportation, autonomous driving and video surveillance, etc., existing compression methods often perform poorly in meeting the needs of machine vision tasks, because they fail to fully preserve semantic information. In addition, most machine vision-oriented compression methods are tightly coupled with specific tasks, lacking generality. Patent application CN115460415A discloses a video compression method, device and storage medium for human-computer hybrid vision. Although this method can support both human eye perception and machine analysis with the same code stream, it is limited by the constraints of a single machine analysis task, and needs to be retrained when dealing with multiple machine analysis tasks, thus consuming too much computational resources and time cost. SUMMARY
[0003] The purpose of the present application is to overcome the defects of the prior art and provide a semantic perception video compression method and system for human-computer hybrid vision. The present application effectively preserves key information, maintains high machine vision task accuracy under low code rate conditions, effectively avoids waste of bit resources on redundant information, and achieves better rate-distortion performance, achieving higher rate-accuracy performance in machine vision tasks.
[0004] The purpose of the present application can be achieved by the following technical solutions:
[0005] A semantic perception video compression method for human-computer hybrid vision, comprising the following steps:
[0006] Input the video sequence to be processed into a pre-trained video compression network, wherein the video compression network comprises a basic branch and an auxiliary branch;
[0007] Generate a semantic compression reconstructed video through the basic branch, comprising the following steps:
[0008] Extract the dynamic semantics of the video sequence to generate a region of interest;
[0009] Input the input frame of the video sequence and the corresponding region of interest mask into a neural rendering model to generate a visually consistent focused frame;
[0010] predicting a feature probability distribution of the focused frame through an entropy model, and compressing the feature probability distribution of the focused frame into a code stream through entropy coding;
[0011] decoding the code stream through a conditional decoder to obtain a semantic compression reconstructed video;
[0012] transforming the semantic compression reconstructed video into a final compression reconstructed video through the auxiliary branch, and the specific steps include:
[0013] aligning the decoded frames in the decoded frame buffer of the base branch and the auxiliary branch to generate predicted features;
[0014] inputting the predicted frames of the predicted features and the video sequence into an entropy model to obtain a feature probability distribution of the video sequence, and compressing the feature probability distribution of the video sequence into a code stream through entropy coding;
[0015] decoding the code stream through a conditional decoder to obtain reconstructed features;
[0016] transforming the reconstructed features into fine reconstructed features through a cross-attention mechanism to obtain a final compression reconstructed video.
[0017] Further, the dynamic semantics of the video sequence are extracted, and the specific steps of generating the region of interest include: in a task with clear semantic information, a Grounded-SAM-based prompt-driven extraction unit is used to extract a specific region of interest; in a task with ambiguous semantic information, a deep-based generalized extraction unit is used to generate a depth map through a depth estimation network to separate the foreground and the background to extract a generalized region of interest.
[0018] Further, the focused features of the focused frame are:
[0019]
[0020] wherein f roi is a focused feature, is an input feature, σ is an activation function, is a fusion unit, m t is a region of interest mask, g0, g1, g2, and g3 are general functions composed of convolution layer stacks, is a channel concatenation operation.
[0021] Further, the feature probability distribution is:
[0022]
[0023] wherein μ tis the mean of the feature probability distribution, σ t is the scale of the feature probability distribution, f pe is the parameter estimation network, f hd is the hyper-prior decoding network, f te is the temporal prior encoding network, f roin is the region of interest prior encoding network, is the hierarchical prior at time step t, C t is the contextual prior at time step t, m t is the region of interest mask, is the latent prior at time step t-1.
[0024] Further, the predicted feature is:
[0025]
[0026] wherein, is the predicted feature, l is the l-th scale level, and t is the time step, is a deformable convolution operation, is the reference decoding frame at time step t-1, is the focus frame at time step t, is a refinement layer composed of a stack of convolution layers.
[0027] Further, the fine reconstruction feature is:
[0028]
[0029] wherein, f refine is the fine reconstruction feature, g is a general function composed of a stack of convolution layers, is a channel concatenation operation, is the associated feature extraction, f t is the feature at time step t, is the feature of the n-th reference decoding frame, Softmax is a normalized exponential function, Q is an embedding layer of the query, K is an embedding layer of the key, V is an embedding layer of the value, and n is the number of reference decoding frames.
[0030] Further, the training step of the video compression network comprises a first stage training and a second stage training, the first stage training is used to enhance the video compression performance for machine analysis, and to reserve high-level semantic content; the second stage training is used to optimize the visual quality perceived by human beings, and to enhance the overall clarity and detail performance of the picture by improving the pixel-level accuracy.
[0031] Further, in the first stage training, the region of interest mask is integrated into a conventional video compression rate-distortion loss function, the reconstruction quality of the focus frame is maximized while minimizing the total bit cost, and the loss function of the first stage training is:
[0032]
[0033] wherein, is the loss function of the first stage training, N is the size of the image group, t is the time step, is the number of bits occupied by the latent motion representation, is the number of bits occupied by the latent context representation, λ is the Lagrange multiplier, is the weight parameter of the region of interest, is the distortion degree between the input frame and the reconstructed frame, x t is the input frame, is the reconstructed frame, is the weight parameter of the non-interest region, p is the penalty hyperparameter.
[0034] Further, in the second stage training, the bit cost estimation only considers the latent context representation, and the loss function of the second stage training is:
[0035]
[0036] wherein, is the loss function of the second stage training, N is the size of the image group, t is the time step, is the number of bits occupied by the latent context representation, λ is the Lagrange multiplier, is the distortion degree between the input frame and the reconstructed frame.
[0037] According to another aspect of the present application, a semantic perception video compression system for human-machine hybrid vision is provided, characterized in that it comprises:
[0038] a video sequence input module, configured to input a video sequence to be processed into a pre-trained video compression network, wherein the video compression network comprises a basic branch and an auxiliary branch;
[0039] a semantic compression reconstructed video generation module, configured to generate a semantic compression reconstructed video through the basic branch, and the specific steps comprise: extracting dynamic semantics of the video sequence to generate a region of interest; inputting an input frame of the video sequence and a corresponding region of interest mask into a neural rendering model to generate a visually consistent focus frame; predicting a feature probability distribution of the focus frame through an entropy model, and compressing the feature probability distribution of the focus frame into a code stream through entropy coding; decoding the code stream through a conditional decoder to obtain a semantic compression reconstructed video;
[0040] The final compression reconstructed video conversion module is used for converting the semantic compression reconstructed video into a final compression reconstructed video through the auxiliary branch, and the specific steps include: performing feature alignment on the decoded frames in the decoded frame buffer of the base branch and the auxiliary branch to generate a predicted feature; inputting the predicted frame of the predicted feature and a video sequence into an entropy model to obtain a video sequence feature probability distribution, and compressing the video sequence feature probability distribution into a code stream through entropy coding; decoding the code stream through a conditional decoder to obtain a reconstructed feature; and converting the reconstructed feature into a fine reconstructed feature through a cross-attention mechanism to obtain the final compression reconstructed video.
[0041] Compared with the prior art, the present application has the following beneficial effects:
[0042] 1. The dynamic region of interest coding of the present application accurately identifies and preferentially guarantees the coding quality of the region of interest, and concentrates the limited bit resources on the key region of the machine vision task, so that the intelligent bit allocation mechanism based on the region of interest effectively retains the key information, thereby maintaining a high machine vision task accuracy under low code rate conditions.
[0043] 2. The auxiliary branch refining of the present application focuses on the difference between the compression focus frame and the input frame, not only supplements the focus frame with more detailed information, but also improves the compression effect of the non-interest region, so that the mechanism effectively avoids the waste of bit resources on redundant information, thereby achieving better rate-distortion performance and higher rate-accuracy performance in the machine vision task, and meeting the needs of human vision and machine vision. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 A flowchart of a semantic perception video compression method for human-machine hybrid vision proposed by the present application;
[0045] Figure 2 A specific steps diagram for generating a semantic compression reconstructed video through a base branch;
[0046] Figure 3 A specific steps diagram for converting a semantic compression reconstructed video into a final compression reconstructed video through an auxiliary branch;
[0047] Figure 4 A BD-ACC performance comparison chart, wherein (4a) is a BD-METEOR performance comparison chart, (4b) is a BD-ROUGEL performance comparison chart, (4c) is a BD-CIDER performance comparison chart, (4d) is a BD-Top-1Accuracy performance comparison chart, (4e) is a BD-R Mean performance comparison chart, and (4f) is a BD-AP performance comparison chart;
[0048] Figure 5 Fig. 5 is a diagram of BD-PSNR performance comparison, wherein (5a) is a diagram of BD-PSNR performance comparison on HEVC Class B dataset, (5b) is a diagram of BD-PSNR performance comparison on HEVC Class C dataset, (5c) is a diagram of BD-PSNR performance comparison on HEVC Class D dataset, and (5d) is a diagram of BD-PSNR performance comparison on MCI JCV dataset;
[0049] Figure 6 Fig. 1 is a structural schematic diagram of a semantic perception video compression system for human-machine hybrid vision according to the present application. DETAILED DESCRIPTION
[0050] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. The present embodiment is implemented on the premise of the technical solution of the present application, and gives a detailed implementation manner and specific operation process, but the protection scope of the present application is not limited to the following embodiments.
[0051] Embodiment 1
[0052] The present embodiment provides a semantic perception video compression method for human-machine hybrid vision, as shown in Fig. 1, which comprises the following steps: Figure 1
[0053] S1, inputting a video sequence to be processed into a pre-trained video compression network.
[0054] The video compression network comprises a basic branch and an auxiliary branch. The training steps of the video compression network comprise a first stage training and a second stage training. The first stage training is used to enhance the video compression performance for machine analysis and to retain high-level semantic content. The second stage training is used to optimize the visual quality perceived by humans and to enhance the overall clarity and detail performance of the picture by improving the pixel-level accuracy.
[0055] In the first stage training, the region of interest mask is integrated into the conventional video compression rate-distortion loss function, the reconstruction quality of the focused frame is maximized while the total bit cost is minimized, and the loss function of the first stage training is:
[0056]
[0057] In the formula, L is the loss function of the first stage training, N is the size of the image group, t is the time step, is the number of bits occupied by the latent motion representation, is the number of bits occupied by the latent context representation, and λ is the Lagrange multiplier, is the weight parameter of the region of interest, where x is the distortion between the input frame and the reconstructed frame. t where x is the distortion between the input frame and the reconstructed frame. where x is the distortion between the input frame and the reconstructed frame. where p is the penalty hyper-parameter.
[0058] In the second stage training, the bit cost estimation only considers the latent context representation, and the loss function of the second stage training is:
[0059]
[0060] where x is the distortion between the input frame and the reconstructed frame. where x is the distortion between the input frame and the reconstructed frame. N is the size of the image group, t is the time step, where x is the distortion between the input frame and the reconstructed frame. N is the size of the image group, t is the time step, where x is the distortion between the input frame and the reconstructed frame. This structured loss function effectively balances the trade-off between different distortion measures, ultimately building a more robust and adaptive compression system.
[0061] S2, generating a semantic compression reconstructed video through the basic branch.
[0062] The specific steps of generating a semantic compression reconstructed video through the basic branch are as shown in Figure 2 The specific steps of generating a semantic compression reconstructed video through the basic branch are as shown in
[0063] S201, extracting the dynamic semantics of the video sequence to generate the region of interest.
[0064] The specific steps of extracting the dynamic semantics of the video sequence to generate the region of interest include: in a task with clear semantic information, a prompt-driven extraction unit based on Grounded-SAM is used to extract a specific region of interest, and the powerful semantic understanding ability and prompt-driven mechanism of Grounded-SAM can efficiently lock the target region to provide clear and accurate semantic guidance for neural rendering. In a task with ambiguous semantic information, a generalized extraction unit based on depth is used to generate a depth map through a depth estimation network to separate the foreground and background to extract a generalized region of interest. The depth estimation network generates a depth map containing rich spatial information such as the relative position and distance of objects in the scene. It is particularly important to note that depth information has a significant advantage in encoding the region of interest, which can effectively separate the foreground and background, and even improve the accuracy of machine analysis when the semantic information is not clear.
[0065] S202, inputting the input frame and the corresponding region of interest mask of the video sequence into the neural rendering model to generate a visually consistent focused frame.
[0066] The neural rendering model receives uncompressed frames and their corresponding semantic references, and generates focused frames for subsequent compression. The focused frames are input into the compression network. In this stage, the feature learning within the compression network is guided by the region-of-interest mask, ensuring that the compression process focuses on the relevant region-of-interest while weakening the attention to non-interesting regions. As a result, the compression network allocates more resources to the compression of key regions, while implementing a more aggressive compression strategy for secondary regions. The focused features of the focused frames are:
[0067]
[0068] where f roi is the focused feature, is the input feature, σ is an activation function, is a fusion unit, m t is a region-of-interest mask, g0, g1, g2, and g3 are general functions composed of convolutional layer stacks, is a channel concatenation operation.
[0069] By directing the allocation of computational resources, the representation ability of the region-of-interest is effectively enhanced, thereby improving the overall performance of the system.
[0070] S203, predict the feature probability distribution of the focused frame through the entropy model, and compress the feature probability distribution of the focused frame into a code stream through entropy coding.
[0071] The feature probability distribution is:
[0072]
[0073] where μ t is the mean of the feature probability distribution, σ t is the scale of the feature probability distribution, f pe is the parameter estimation network, f hd is the hyper-prior decoding network, f te is the temporal prior encoding network, f roin is the region-of-interest prior encoding network, is the hierarchical prior at time step t, C t is the contextual prior at time step t, m t is the region-of-interest mask, is the latent prior at time step t-1.
[0074] By introducing the region-of-interest prior, the system can allocate more bits to key regions, thereby improving the accuracy of the reconstructed features.
[0075] S204, decode the code stream through the conditional decoder to obtain a semantically compressed reconstructed video.
[0076] S3, reconstruct the video compressed by the semantic branch into the final compressed reconstructed video through the auxiliary branch.
[0077] The specific steps of reconstructing the video compressed by the semantic branch into the final compressed reconstructed video through the auxiliary branch are as shown in the figure, comprising the following steps: Figure 3
[0078] S301, align the decoded frames in the decoded frame buffer of the base branch and the auxiliary branch by features to generate predicted features.
[0079] The focus frame is fused with the reference decoded frame extracted from the previous reconstructed frame (the previous three high-quality decoded frames are selected as the reference decoded frame in this embodiment) to generate predicted features. The alignment of the focus frame and the reference decoded frame is realized by deformable convolution. Considering the advantage of the Transformer in capturing global context information, the Transformer is integrated into the temporal feature alignment process to provide multi-scale features containing richer and more accurate information. The multi-scale feature extraction process can be represented as:
[0080]
[0081] In the formula, l represents the lth level of scale, is a multi-scale feature extractor composed of Transformer blocks, each Transformer block contains block attention B attn and grid attention G attn The block attention focuses on maintaining fine-grained local details, and the grid attention enhances long-range dependencies, both of which cooperate to improve the feature representation ability.
[0082] The predicted feature is:
[0083]
[0084] In the formula, is the predicted feature, l is the lth level of scale, T is the time step, is a deformable convolution operation, is the reference decoded frame at time step T-1, is the focus frame at time step T, is a refinement layer composed of a stack of convolution layers.
[0085] S302, input the predicted frame of the predicted feature and the video sequence into the entropy model to obtain the video sequence feature probability distribution, and compress the video sequence feature probability distribution into a code stream through entropy coding.
[0086] S303, decode the code stream through the conditional decoder to obtain the reconstructed feature.
[0087] The code stream is converted into latent features, and the latent features are converted into reconstructed features through a conditional codec.
[0088] S304, the reconstructed features are converted into fine reconstruction features through a cross-attention mechanism to obtain a final compressed reconstruction video.
[0089] The fine reconstruction features are:
[0090]
[0091] In the formula, f refine is a fine reconstruction feature, g is a general function composed of a convolution layer stack, is a channel splicing operation, is a correlation feature extraction, f t is a feature of a time step t, is a feature of the nth reference decoded frame, Softmax is a normalized exponential function, Q is an embedding layer of a query, K is an embedding layer of a key, V is an embedding layer of a value, and n is the number of reference decoded frames.
[0092] In order to verify the performance of the above method, the following two experiments are designed: first, in order to evaluate the performance of the proposed method in machine analysis tasks, this paper compares the differences in rate-accuracy performance of different methods. Specifically, the bit per pixel (BPP) and the accuracy of each machine analysis task are used as evaluation indexes, and the traditional codec H.265 / HEVC is used as a benchmark to calculate the relative performance. As shown in Table 1, compared with traditional methods and learning-based methods, the proposed method shows the best performance in the BD-ACC index of four machine vision tasks.
[0093] Table 1 Comparison of BD-ACC performance relative to H.265 / HEVC
[0094]
[0095] In order to more intuitively show the rate-accuracy characteristics, Figure 4 the performance curves of different methods on each dataset are presented. As Figure 4 shown, Figure 4 (4a) of FIG. 4 is a comparison of BD-METEOR performance, Figure 4 (4b) of FIG. 4 is a comparison of BD-ROUGEL performance, Figure 4 (4c) of FIG. 4 is a comparison of BD-CIDER performance, Figure 4 (4d) of FIG. 4 is a comparison of BD-Top-1Accuracy performance, Figure 4 (4e) of FIG. 4 is a comparison of BD-R Mean performance, Figure 4Figure (4f) shows the performance comparison of BD-AP. The rate-accuracy curve of this method is consistently better than other methods, especially under low bitrate conditions. This phenomenon stems from the fact that traditional video compression methods often perform coarse-grained compression on the entire frame at low bitrates, resulting in severe image quality degradation and affecting the accuracy of machine vision tasks. In contrast, this method accurately identifies and prioritizes the encoding quality of regions of interest, concentrating limited bit resources on key areas of the machine vision task. This intelligent bit allocation mechanism based on regions of interest effectively preserves key information, thus maintaining high accuracy in machine vision tasks even under low bitrate conditions.
[0096] Secondly, to verify the compression performance of the proposed method in terms of signal fidelity, this paper compares the performance of different methods in terms of rate distortion. As shown in Table 2, the performance is compared using three methods: BD-Rate, BD-PSNR, and rate distortion curves.
[0097] Table 2 Comparison of BD-Rate and BD-PSNR performance relative to H.265 / HEVC
[0098]
[0099] like Figure 5 As shown, Figure 5 Figure (5a) shows the BD-PSNR performance comparison on the HEVC Class B dataset. Figure 5 Figure (5b) shows a comparison of BD-PSNR performance on the HEVC Class C dataset. Figure 5 Figure (5c) shows the BD-PSNR performance comparison on the HEVC Class D dataset. Figure 5 Figure (5d) shows the BD-PSNR performance comparison on the MCI_JCV dataset. Performance comparisons are performed using three methods: BD-Rate, BD-PSNR, and rate-distortion curves. The traditional codec H.265 / HEVC is used as the benchmark for relative performance calculation. Experimental results show that the video compression method proposed in this invention achieves an average BD-Rate reduction of 36.93% and a BD-PSNR improvement of 1.47dB on the benchmark dataset, significantly outperforming other comparative methods in terms of rate-distortion performance. Notably, the auxiliary branch plays a crucial role in improving the method's rate-distortion performance. This branch focuses on compressing the difference between the focused frame and the input frame, not only supplementing the focused frame with finer detail information but also improving the compression effect of non-interest areas. This mechanism effectively avoids wasting bit resources on redundant information, thus achieving superior rate-distortion performance.
[0100] Example 2
[0101] This embodiment provides a semantic-aware video compression system for human-computer hybrid vision, such as... Figure 6 As shown, it includes:
[0102] The video sequence input module is used to input the video sequence to be processed into a pre-trained video compression network, which includes a basic branch and an auxiliary branch.
[0103] The semantic compression and reconstruction video generation module is used to generate semantically compressed and reconstructed videos through basic branches. The specific steps include: extracting dynamic semantics from the video sequence to generate regions of interest; inputting the input frames of the video sequence and the corresponding region of interest masks into the neural rendering model to generate visually consistent focus frames; predicting the feature probability distribution of the focus frames through an entropy model and compressing the feature probability distribution of the focus frames into a bitstream through entropy encoding; and decoding the bitstream through a conditional decoder to obtain the semantically compressed and reconstructed video.
[0104] The final compressed and reconstructed video conversion module is used to convert the semantically compressed and reconstructed video into the final compressed and reconstructed video through auxiliary branches. The specific steps include: aligning the decoded frames in the decoding frame buffers of the basic branch and the auxiliary branch to generate predicted features; inputting the predicted frames of the predicted features and the video sequence into the entropy model to obtain the probability distribution of video sequence features, and compressing the probability distribution of video sequence features into a bitstream through entropy encoding; decoding the bitstream through a conditional decoder to obtain reconstructed features; and converting the reconstructed features into fine-grained reconstructed features through a cross-attention mechanism to obtain the final compressed and reconstructed video.
[0105] The rest is the same as in Example 1.
[0106] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A semantic-aware video compression method for human-machine hybrid vision, characterized in that, The method comprises the following steps: inputting a video sequence to be processed into a pre-trained video compression network, wherein the video compression network comprises a basic branch and an auxiliary branch; generating a semantic compression reconstructed video through the basic branch, and the specific steps comprise: extracting dynamic semantics of the video sequence to generate a region of interest; inputting an input frame of the video sequence and a corresponding region of interest mask into a neural rendering model to generate a visually consistent focused frame; predicting a feature probability distribution of the focused frame through an entropy model, and compressing the feature probability distribution of the focused frame into a code stream through entropy coding; decoding the code stream through a conditional decoder to obtain a semantic compression reconstructed video; transforming the semantic compression reconstructed video into a final compression reconstructed video through the auxiliary branch, and the specific steps comprise: aligning features of decoded frames in a decoded frame buffer of the basic branch and the auxiliary branch to generate a predicted feature; inputting a predicted frame of the predicted feature and a video sequence into an entropy model to obtain a feature probability distribution of the video sequence, and compressing the feature probability distribution of the video sequence into a code stream through entropy coding; decoding the code stream through a conditional decoder to obtain a reconstructed feature; transforming the reconstructed feature into a fine reconstructed feature through a cross-attention mechanism to obtain a final compression reconstructed video.
2. The semantic-aware video compression method for human-machine hybrid vision of claim 1, wherein, The specific steps for extracting dynamic semantics of the video sequence to generate a region of interest comprise: in a task with clear semantic information, a prompt-driven extraction unit based on Grounded-SAM is used to extract a specific region of interest; in a task with ambiguous semantic information, a deep-based generalized extraction unit is used to generate a depth map through a depth estimation network to separate foreground and background to extract a generalized region of interest.
3. The semantic-aware video compression method for human-machine hybrid vision of claim 1, wherein, The focused feature of the focused frame is: where f roi is a focusing feature, is an input feature, σ is an activation function, is a fusion unit, m t is a region of interest mask, g0, g1, g2, and g3 are general functions composed of a convolutional layer stack, is a channel concatenation operation.
4. The method of claim 1, wherein, The feature probability distribution is: where μ t is the mean of the characteristic probability distribution, σ t is the scale of the characteristic probability distribution, f pe is the parameter estimation network, f hd is the hyper-prior decoding network, f te is the temporal prior encoding network, f roin is the region of interest prior encoding network, is the hierarchical prior at time step t, C t is the contextual prior at time step t, m t is the region of interest mask, is the latent prior at time step t-1.
5. The semantic-aware video compression method for human-machine hybrid vision of claim 1, wherein, The predicted feature is: wherein, is a prediction feature, l is the l-th scale level, t is a time step, is a deformable convolution operation, is a reference decoded frame at time step t-1, is a focus frame at time step t, is a refinement layer composed of a stack of convolution layers.
6. The semantic-aware video compression method for human-machine hybrid vision of claim 1, wherein, The fine reconstructed feature is: where f rerine is a fine reconstruction feature, g is a general function composed of a convolutional layer stack, is a channel concatenation operation, is a correlation feature extraction, f t is a feature at time step t, is a feature of the n-th reference decoded frame, Softmax is a normalized exponential function, Q is an embedding layer of a query, K is an embedding layer of a key, V is an embedding layer of a value, and n is a number of reference decoded frames.
7. The method of claim 1, wherein, The training steps of the video compression network comprise a first stage training and a second stage training, the first stage training is used to enhance the video compression performance for machine analysis and retain high-level semantic content; the second stage training is used to optimize the visual quality perceived by humans, and enhance the overall clarity and detail performance of the picture by improving the pixel-level accuracy.
8. The semantic-aware video compression method for human-machine hybrid vision of claim 7, wherein, In the first stage training, the region of interest mask is integrated into the conventional video compression rate-distortion loss function, the reconstruction quality of the focused frame is maximized while the total bit cost is minimized, and the loss function of the first stage training is: wherein, is the loss function for the first stage training, N is the image group size, t is the time step, is the number of bits occupied by the latent motion representation, is the number of bits occupied by the latent context representation, λ is the Lagrange multiplier, is the weight parameter for the region of interest, is the distortion between the input frame and the reconstructed frame, x t is the input frame, is the reconstructed frame, is the weight parameter for the non-region of interest, p is the penalty hyperparameter.
9. The method of claim 7, wherein the semantic perception video compression method is human-machine hybrid vision-oriented, and In the second stage training, the bit cost estimation only considers the potential context representation, and the loss function of the second stage training is: In the formula, is the loss function of the second stage training, N is the size of the image group, t is the time step, is the number of bits occupied by the latent context representation, λ is the Lagrange multiplier, is the distortion degree between the input frame and the reconstructed frame.
10. A semantic-aware video compression system for human-in-the-loop vision, characterized in that, comprise: a video sequence input module, configured to input a video sequence to be processed into a pre-trained video compression network, wherein the video compression network comprises a basic branch and an auxiliary branch; The semantic compression reconstruction video generation module is configured to generate a semantic compression reconstruction video through the base branch, and specifically includes the following steps: extracting dynamic semantics of the video sequence, and generating a region of interest; inputting an input frame of the video sequence and a corresponding region of interest mask into a neural rendering model to generate a visually consistent focused frame; predicting a feature probability distribution of the focused frame through an entropy model, and compressing the feature probability distribution of the focused frame into a code stream through entropy coding; and decoding the code stream through a conditional decoder to obtain the semantic compression reconstruction video. The final compression reconstruction video conversion module is configured to convert the semantic compression reconstruction video into a final compression reconstruction video through the auxiliary branch, and specifically includes the following steps: aligning features of decoded frames in a decoded frame buffer of the base branch and the auxiliary branch to generate predicted features; inputting predicted frames of the predicted features and the video sequence into an entropy model to obtain a feature probability distribution of the video sequence, and compressing the feature probability distribution of the video sequence into a code stream through entropy coding; decoding the code stream through a conditional decoder to obtain reconstructed features; converting the reconstructed features into fine reconstruction features through a cross-attention mechanism to obtain the final compression reconstruction video.
Citation Information
Patent Citations
Video compression method for man-machine mixed vision
CN115460415A
Cited By
Compression coding processing method and system for observation data of low earth orbit satellite
CN121485791A